← Blog

AI Software Development: Why Writing the Code Was Never the Hard Part

Marco Masut

"AI software development" now means building production software where an AI agent writes and edits a meaningful share of the code, inside a process a person still owns: what gets built, what "done" means, and who signs off before it ships. Most of the industry keeps score by counting code produced. That's the wrong unit.

Search the term today and you get two kinds of pages: a vendor selling an agent, or a think piece arguing whether AI replaces developers. Neither answers the question a delivery owner actually has, which is what changes about how software gets built when an agent writes a chunk of the code, and what doesn't change at all.

What does "AI software development" actually mean today?

The term covers a wide range in practice: a developer accepting agent-written functions inside an IDE, a headless agent running a whole ticket end to end overnight, and, at the far edge, someone with no engineering background shipping an app they never read the code for. All three get called AI software development, because all three route code through a model before a human sees it. What varies isn't whether AI wrote the code. It's how much of the surrounding process, deciding what to build, checking what came back, deciding it's safe to ship, still happens on purpose.

Building software has always required five things: understanding what to build (analysis), giving the builder accurate facts to work from (context), writing the code itself, checking that the result does what was asked (verification), and someone taking responsibility for shipping it (acceptance). AI coding agents have made real, measurable gains on exactly one of those five: producing code faster. The other four, deciding what to build, supplying the right context, verifying the result, and accepting responsibility for it, still run through a person, because none of them get solved by generating text more quickly. Judging AI software development purely by how much code an agent produces measures the one step of the five that was already the least scarce resource, and says nothing about the four that were always the actual work. That's the mistake the term keeps making.

Was writing code ever the real bottleneck?

Software engineering has treated this as settled since 1986, when Fred Brooks' essay "No Silver Bullet" split the difficulty of building software into two kinds. Essential complexity is the problem itself: the logic, the edge cases, the requirements that conflict with each other. Accidental complexity is the friction of expressing that problem in a given language and toolchain, syntax, boilerplate, the mechanics of getting an idea into a file the computer will run. Every advance since then, higher-level languages, frameworks, autocomplete, and now generative agents, has attacked accidental complexity, because it's the part that's tractable to automate. None of it has touched essential complexity, because deciding what the software should actually do was never a typing problem.

Even the narrower claim, that agents at least made the typing faster, doesn't hold up cleanly. A METR randomized controlled trial published in July 2025 had 16 experienced open-source developers complete 246 real tasks in codebases they knew well, half with AI tools allowed and half without. Before starting, they forecast AI would cut completion time by about 24%. The measured result went the other way: tasks took about 19% longer with AI allowed, and the developers still walked away believing they had been roughly 20% faster. If the one step AI was supposed to have already solved, producing code quickly, doesn't reliably get faster once someone measures it instead of guessing, the idea that AI software development reduces to a code-generation problem was wrong before the rest of the argument even starts.

Where does the time actually go, before and after AI agents?

The industry has enough survey data now to see where the freed-up time actually went, and it didn't evaporate. BairesDev's Q3 2026 Dev Barometer, a survey of 705 developers and 41 engineering leaders published September 15, 2026, found that 79% of developers now spend less than half their working week writing new code from scratch. That sounds like the productivity story AI software development is supposed to deliver. The same survey found that 67% of those developers spend more time reviewing AI-generated code than they did a year earlier, and 52% spend more time debugging problems that AI introduced. The hours didn't disappear. They moved one stage down the pipeline, from writing to checking.

Before AI agentsAfter AI agents
Where the hours goWriting and editing code by handReviewing, debugging, and verifying what the agent produced
The scarce skillWriting correct code fastSpecifying intent precisely, judging output fast
The bottleneckProducing enough codeEverything that happens once the code already exists
Who answers for a shipped changeWhoever wrote itWhoever approved it, whether or not they wrote it

What changes in the ratios, and what stays exactly the same?

The 2025 DORA report, published September 23, 2025, found something that looks contradictory until the two claims are separated: AI adoption is now positively correlated with software delivery throughput, a reversal from the previous year's finding, and it remains negatively correlated with delivery stability. Teams ship more, faster, and a larger share of what they ship destabilizes something downstream. DORA's own explanation is that AI accelerates development, but that acceleration exposes weaknesses that were already there: a team with weak verification and unclear acceptance criteria doesn't get those problems fixed by a faster agent, it gets them multiplied.

That's the ratio shift in one sentence: the cost of producing a line of code keeps falling, so a larger share of a team's total effort now sits in the four steps AI didn't touch, deciding what to build, supplying the agent accurate context, verifying what it did, and accepting responsibility for the result. None of those four got easier. Two of them got harder, in practice, simply because there's more code moving through them per week than there used to be.

How to run it: contract, verification, acceptance

For a team, the practical version of this is a discipline, not a mindset. It starts with treating agentic coding as a loop that needs a contract before it runs: what the agent is authorized to touch, what counts as done, and what evidence has to exist before anyone calls the task finished. It continues with verification that doesn't take the agent's own report as proof, a test suite, a build, a check a machine runs, not a summary the same model that wrote the code also wrote about itself. It ends with a human acceptance step that gets recorded somewhere durable, not implied by the fact that nobody complained.

None of that is a reason to slow down how a team picks its AI coding assistant: the tool question and the process question are separate, and a faster agent inside a disciplined process is strictly better than the same agent inside no process at all. What the discipline buys is the part most roundups of coding agents don't ask about: whether a team can still show, after the fact, what was asked, what the agent was allowed to do, what it actually did, and who accepted it.

What should you measure instead of lines of code?

A few numbers correlate better with whether a team actually has control over its own delivery than lines of code written per week ever did: the share of shipped changes with a recorded, checkable reason for existing; the time between an agent finishing a task and a human formally accepting it, not the time to a first draft; the rate of production incidents traced back to a change nobody reviewed closely; and, the blunt version of all three, whether a delivery lead could hand a client the full history of one specific change, six months later, without reconstructing it from memory or a chat log.

That's the boundary the layer missing above every coding agent is built to sit on: not writing better code, which agents are already doing, but tying together what a team meant to build, what it authorized, what happened, and who signed off, into something that survives the session and the person who ran it. Detent (detent-ai.com) is built around exactly that boundary, on the assumption that the industry's current unit of measurement, code produced, was never the one that mattered.