Detent Bench
Our internal benchmark, built on end-to-end tests from the request to production-ready. It says how powerful and complete the delivery line actually is, all the way to release.
What it measures
The line, not the model. How many jobs, on a real repository, get from the request to production-ready: correct plan, green evidence, human signature, a release that holds.
How it is built
End-to-end tests across the whole path. Same job contract, same independent evidence, same signature gate. Whoever executes does not write the tests that judge them.
What it answers
How powerful and complete the line is, from the request to a release that holds. Not how good an agent is on a synthetic task, in a session that then disappears.
Four levels, one after the other.
It is not a single score. It is a scale: each level narrows the one before it. The number that counts is where the job stops.
The unit is not the prompt. It is the job contract.
Starting SHA, request, acceptance criteria, authorized context, budget, allowed tools, independent verification. If the yardstick shows nothing, we say so and the pilot stops: a layer that does not move those numbers is added complexity, not a product.
Detent Bench is not a model leaderboard. It is the proof of the line. With more comparable lines it becomes something else: knowing, with the evidence in hand, which work can be delegated, how completely, and at what risk.
Talk to the founder