← Blog

Spec Kit Writes the Spec. It Does Not Prove the Code Matches It.

Marco Masut

Spec Kit is an open-source toolkit from GitHub for spec-driven development with AI coding agents. It gives a team a fixed sequence of commands, constitution, specify, plan, tasks, implement, so the agent works from a written specification instead of a loose prompt, and the team keeps a shared structure across features. The repository is MIT licensed and, when we consulted it on 6 October 2026, showed more than 140,000 stars. What it standardises is the input: what the agent is asked to build, in what order, under which principles. What it does not produce is a record of what the agent actually did, which files changed, which checks ran, who reviewed the result and who signed it. A spec is an instruction written before the work. Evidence is what remains after it, and a team that ships for clients needs both.

What does Spec Kit actually standardise?

The problem is real. Without a shared structure, every feature starts from a different prompt, and the agent's interpretation of "build this" drifts from session to session. Spec Kit fixes the shape of the conversation. A constitution file holds the project's non-negotiable principles. The specify step turns a request into a written specification. Plan and tasks break that into technical decisions and ordered units of work, and implement executes them through whichever supported agent the team uses.

That is a genuine improvement over improvising, and it is why the toolkit spread. It also fits the broader move toward context engineering: the better the written context, the less the agent has to guess. Credit where it is due: a team using it will brief its agents more consistently than one that does not.

Specification, plan, tasks: what is the contract before the work?

Read the sequence as a contract. The specification says what, the plan says how, the tasks say in which order. Each step is a markdown file in the repository, versioned with the code, which is better than a chat history that disappears.

| Artifact | Written | Answers | Does not answer | | --- | --- | --- | --- | | Constitution | Once, up front | Which principles always apply | Whether this change respected them | | Specification | Before the work | What should be built | Whether what was built matches | | Plan | Before the work | How it should be built | Whether the agent followed it | | Tasks | Before the work | In which order | Which task actually changed which file |

Every artifact in the table describes intent. None of them is written after the agent has finished, so none can describe what happened.

Is there a gap between the spec and the diff that lands?

Yes, and it is the same gap that AGENTS.md leaves open. A specification is read at the start of a session. The agent may follow it faithfully, partly, or interpret an ambiguous line in a way nobody intended. The only way to know is to compare the specification against the actual diff, and that comparison is a separate step the toolkit's spec files do not perform on their own.

The toolkit also offers steps for refining and reconciling the written documents, which helps catch contradictions between them. They compare documents with documents. The question a client or an auditor asks is different: does the code in this pull request do what was agreed, and what proves it?

Why is first-pass success rate not an audit trail?

Spec-driven tools are usually sold on a metric: how often the agent gets it right the first time. It is a useful metric for the developer's afternoon. It is not evidence for anyone else. A high success rate says that most runs turned out fine, not which run produced this change, under which authorization, reviewed by whom. The METR randomised trial published on 10 July 2025 is a reminder that perceived speed and measured outcomes can diverge: experienced developers took 19% longer with AI tools while believing they were faster. Self-reported success is the weakest kind of signal, and an audit trail has to be independent of what the agent says about itself.

What does a spec-driven flow need to become evidence?

Keep the specification workflow. Then add what it was never designed to hold, as a checklist:

  • A verification step run by something other than the agent, with its output stored.
  • A scope check: the diff touches only the paths the task allowed.
  • A link from each task in the plan to the commit that implemented it.
  • A named human who reviewed the result and signed it before release.
  • A record that survives the session, so the answer to "what happened last Tuesday" does not depend on anyone's memory.

This is the layer Detent, the end-to-end delivery system adds on top of a spec: the system prepares, a person signs, and nothing ships without that signature. A software house can use Spec Kit for the briefing and still need that layer for the proof. For the wider picture of what a controlled flow has to cover, see what an AI governance framework for agent-written code needs, and for how these tools compare on what they leave behind, see the benchmark.