Prompt engineering is the practice of phrasing instructions for a model so the output matches what you wanted: role, examples, constraints, response format. For two years it was treated as the skill that separates people who get good results from people who don't. With coding agents the useful question has changed: not how to ask better, but how to verify what comes back.
Is prompt engineering still the bottleneck?
Largely no, for anyone using a coding agent on real code. A rough but clear request now almost always produces a plausible diff, one that compiles and often passes the existing tests. The risk is no longer the badly worded sentence; it is the diff that looks right and isn't: it solves a slightly different problem, widens a permission, or takes a shortcut that holds only on the cases anyone saw. Writing a better prompt lowers how often this happens, but it does not make the failures visible. What makes them visible is an acceptance criterion written beforehand, a verification command run outside the session, and a trace that remains afterward. For a team that delivers code to a client and answers for it months later, the bottleneck has become verification: what gets checked, who checks it, and what remains as proof.
Why did prompt engineering matter so much?
With the first conversational models, output quality depended very visibly on how you asked: the same task, phrased two ways, gave very different results. Guides, courses and reusable prompt templates grew out of that, and search results for this term are full of lists of techniques.
That phase left habits that still hold: state the context, say what is out of scope, ask for a precise response format. But they were remedies for a specific problem, a model that reacted badly to vague instructions, and that problem weighs less than it used to. Treating them as the hard part of the job means optimizing the stage where the least time is lost today.
What changed in the models, and what did not?
Changed: coding agents read files, run commands, correct their own mistakes in a loop, and tolerate imperfect instructions better than models did a couple of years ago. Part of what the prompt used to do, pointing at where to look, the agent now does itself. What remains is the question of what the model sees when it answers, covered in the piece on context engineering: it matters as much as what you ask.
Not changed: the model does not know whether it solved the right problem. It can declare the work finished with the same confidence when it is correct and when it is not, and no phrasing of the prompt gives it an external criterion to check itself against. That criterion has to be brought in from outside.
Where did the breaking point move?
Downstream, to three places a prompt does not cover.
| Breaking point | What the prompt does | What is actually needed |
|---|---|---|
| The diff looks right | Reduces gross errors | An acceptance criterion written beforehand, compared cold against the diff |
| The agent says it is done | Can ask it to "verify the work" | A verification command run by something that is not the agent |
| Months go by | Leaves nothing: the prompt lives in the chat | A trace tied to the change: request, authorized scope, check result |
The evidence that judgment inside the session is not enough is measured. The randomized controlled trial by METR, published in July 2025, had experienced open-source developers work on real tasks in their own repositories: with AI tools allowed they took 19% longer, despite having predicted beforehand that they would be 24% faster and believing afterward that they had been about 20% faster. METR has noted that the result applies to that setting and is not a permanent verdict on the tools. The gap between perceived and measured speed is still the right warning: the person inside the session misjudges their own work, and a better prompt does not change who is judging.
Why are instructions and proof two different things?
An instruction tells the agent what to do. Proof tells someone else what happened. Confusing the two is the most common mistake: a very long prompt gets written, full of rules and constraints, and the matter is considered closed because the rules were written down. But a rule in a prompt leaves no trace of having been followed. An instruction file like AGENTS.md is a good example: it steers the agent, and proves nothing about what the agent then did.
For anyone doing contract work the difference is concrete. Months later, you do not hand the client the prompt. You hand them what was requested, what the agent was allowed to touch, what it touched, and the result of a check the agent did not run itself. This is the level Detent, the end-to-end delivery system (detent-ai.com), works at: verification is run by the system, not declared by the model, and approving a release requires a human signature. The numbers on real tasks are on the /bench page.
What is still worth writing well?
The prompt is not useless; it has stopped being the place where you win or lose. Three things are worth writing well:
- acceptance criteria, before the session starts, in a form a command can check;
- scope: which files and folders the agent may touch and which it may not;
- the stable context of the project, written once and kept in a versioned file instead of repeated in every request.
Everything else, the magic formula, the role assigned to the model, the length of the prompt, matters far less than a verification someone actually runs and a trace someone can reread. If a team has time to invest, today it pays off more there than in polishing the fifth version of a prompt.
