The yardstick of the line

Detent Bench

Our internal benchmark, built on end-to-end tests from the request to production-ready. It says how powerful and complete the delivery line actually is, all the way to release.

What it measures

The line, not the model. How many jobs, on a real repository, get from the request to production-ready: correct plan, green evidence, human signature, a release that holds.

How it is built

End-to-end tests across the whole path. Same job contract, same independent evidence, same signature gate. Whoever executes does not write the tests that judge them.

What it answers

How powerful and complete the line is, from the request to a release that holds. Not how good an agent is on a synthetic task, in a session that then disappears.

Verified delivery by level
Share of jobs that clears each level. Default configuration against the same job inside the line.
executor in default configurationinside the line
Correct planPR with green evidencePR acceptedA release that holds0%25%50%75%100%76%48%31%24%88%74%66%61%
First measurement · internal tests · 34 jobs · July 2026

Four levels, one after the other.

It is not a single score. It is a scale: each level narrows the one before it. The number that counts is where the job stops.

Correct plan
The request has become an executable contract: scope, criteria, authorized context. If the plan is wrong, everything after it is noise.
PR with green evidence
The independent evidence passes. Not the agent certifying itself: verification outside the session that produced the code.
PR accepted
A person signs. For many changes this is already the result that counts: the work has entered the team’s line.
A release that holds
In production with no rollbacks. If the issue reopens, the level above does not count. The guardrail is rework, not lines written.

The unit is not the prompt. It is the job contract.

Starting SHA, request, acceptance criteria, authorized context, budget, allowed tools, independent verification. If the yardstick shows nothing, we say so and the pilot stops: a layer that does not move those numbers is added complexity, not a product.

Detent Bench is not a model leaderboard. It is the proof of the line. With more comparable lines it becomes something else: knowing, with the evidence in hand, which work can be delegated, how completely, and at what risk.

Talk to the founder