← Blog

What Is Agentic Coding? What Should Remain When the Loop Stops

Marco Masut

"Agentic coding" gets searched by people who've already used autocomplete and want to know what the next step actually is. The guides that answer the query, from cloud vendors to engineering blogs, do a decent job of defining the term and flagging the risks of letting a model run unsupervised. What almost none of them ask is a narrower, more practical question: once the loop finishes and the session closes, what is a team actually left holding.

That question matters more than it sounds. Agentic coding isn't a single feature you turn on. It's a loop a tool runs on your behalf, several times per task, mostly without anyone watching each step.

What agentic coding means, beyond autocomplete

Autocomplete finishes the line you're already typing. Agentic coding is a different job: a model is handed a goal, not a keystroke, and it works toward that goal across a repository, over multiple steps, deciding on its own what to read, what to change, and when it's done. Claude Code runs this from a terminal, headless if needed. Cursor's agent mode runs it inside an editor, with each edit visible as it happens. Both count as agentic because both do meaningfully more unsupervised work than a completion tool ever did.

That gap, more unsupervised work per request, is the whole reason the term needed to exist. It's also the reason what happens inside the loop matters more than it did when a human approved every single line.

The loop: plan, edit, run, verify

Strip away the branding, and agentic coding tools converge on the same four-step loop. The agent plans a change based on the goal and the context it can see. It edits one or more files to carry out that plan. It runs something, tests, a build, a linter, to check whether the edit actually works. Then it verifies the result against the goal and either stops, reports back, or loops again with what it learned.

That loop is what separates an agent from a smarter autocomplete: it doesn't just produce a change, it produces and checks its own change before handing it back. That's a real improvement over code that nobody, human or model, ran before you saw it. It's also easy to overstate. A test suite passing inside a session tells you the code did what the tests check for. It doesn't tell you whether the plan was the right one, whether the agent stayed inside what it was actually authorized to touch, or whether the reasoning behind the change survives past the session that produced it.

What existing guides measure

Most of what ranks for "agentic coding" today, from cloud provider explainers to engineering blogs and vendor content, does one of two things well: it defines the loop, or it warns about the risk of giving a model that much autonomy, unreviewed access to a codebase, runaway loops, actions nobody signed off on. Both are useful, and both stop at the same boundary: they treat the loop itself as the unit worth explaining. What the agent produces as output, once the loop is done and everyone has moved on, isn't the question they're answering.

What they leave out: what remains after the loop stops

A completed loop leaves two things behind by default: the changed code, and a log of the session, kept in whatever format the tool happens to use. Neither one is built to answer a question that comes up constantly on real projects, weeks or months later: why was this change made this way, who or what was authorized to make it, and how do we know the thing that ran matches what was actually asked for. What AI coding agents leave out once the session ends is the same gap from the angle of choosing a tool. Here it's a property of the practice itself, independent of which agent is running the loop.

A METR randomized controlled trial from 2025 is a useful check on how much to trust in-loop verification alone: experienced developers who used AI tools on real tasks took about 19% longer than the same developers working without them, yet still walked away believing they had been faster. If the people who ran the loop and watched it happen can misjudge what came out of it, a green test run inside the session is not, by itself, the kind of evidence that holds up to someone who wasn't there.

A definition of remaining proof for agentic coding

Remaining proof, for a loop that plans, edits, runs, and verifies, is what's still checkable after the session ends: the intent behind the task, the boundaries the agent was actually allowed to work inside, what it executed step by step, and an artifact tying those three together, specific enough to audit and portable enough to survive a change of tool. In-loop verification, the run and verify steps, checks that the output works. Remaining proof checks that the whole loop, plan included, matches what was asked and authorized. They answer different questions, and right now, almost nothing in the category produces the second on its own.

How this differs from vibe coding

Vibe coding and agentic coding get used loosely as if they were the same thing, and they're not. Vibe coding describes an attitude: accept what the model produces, keep the momentum, worry about the details later, often without running much of a loop at all. Agentic coding, done properly, already includes verification steps, plan, edit, run, verify, that vibe coding tends to skip. That makes it a real step up in rigor.

It doesn't close the gap this piece is about. A team can run the full loop, every step verified, and still have nothing beyond a tool-specific log to show a client or an auditor six months later. Verifying inside the session and producing proof that outlives the session are two different bars, and clearing the first one doesn't clear the second.

That second bar is the one Detent Bench measures agents against, on the line from request to production-ready, regardless of which agent or which loop produced the change.