← Blog

Agentic Engineering Is a Job, Not a Tool Choice

Marco Masut

Agentic engineering is the practice of building software by handing execution to a coding agent while a person keeps the goal, the constraints, and the call on what ships. It isn't a preference for one tool over another. It's what the job of writing software is turning into, one delegated task at a time, and it's the practitioner-level view of a bigger shift already changing software delivery.

Search the term today and the first page mixes IBM next to two personal blogs, Simon Willison and Addy Osmani, both writing about their own daily practice. That mix is the tell: on this question, a specific, lived answer beats a big domain. This is one.

What does the job look like now, concretely?

A software engineer doing agentic engineering spends less time typing syntax and more time on four things: three before the agent runs, one after. Before: scoping the task tightly enough that "done" has a checkable meaning, writing acceptance criteria the agent can be judged against instead of a vibe, and deciding how much context and which tools the agent is allowed to touch. After: reading the diff it hands back before anything merges. None of these existed as a named, hireable skill five years ago; together they used to be a small part of the job, handled almost automatically while typing the code yourself. All four now decide whether a session produces something a team can ship, or something that looks finished and isn't. The agent still writes most of the characters. It doesn't decide what "correct" means for this specific change, in this specific codebase, for this specific client. That decision, and the responsibility that comes with it, hasn't moved anywhere.

Which skills moved to the centre?

  • Scoping a task narrowly enough that an agent can actually finish it, instead of wandering
  • Writing acceptance criteria before a session starts, not inferring them from the result afterward
  • Reading a diff for intent and side effects, not just for syntax
  • Deciding when to stop an agent that's technically still making progress, but on the wrong path
  • Knowing which parts of a codebase an agent should never touch unsupervised

Autocomplete rewarded typing speed. Agentic engineering rewards judgment applied before and after the agent runs, not during it. The Stack Overflow Developer Survey 2025 found that 69% of developers using coding agents report higher individual productivity, but only 17% report better team collaboration. Speed at the keyboard was never the scarce skill; deciding what's worth building, and confirming it was built correctly, always was. Agentic engineering just moved that fact out into the open.

How do you read a diff you did not write?

Against the acceptance criteria written before the session started, not against how plausible the code looks. A diff from an agent compiles, often passes the tests it was told to pass, and can still do the wrong thing for a subtle reason: solve a slightly different problem than the one asked, quietly widen a permission, or take a shortcut that works today and breaks on a case nobody specified. METR's 2025 randomized controlled trial found that experienced developers using AI tools were 19% slower on real tasks, despite predicting beforehand they'd be 24% faster, and believing afterward they had been 20% faster. Self-assessment inside the session isn't reliable. The diff, read cold against a written bar, is what agentic engineering asks a team to trust instead.

How do you know when to stop the agent?

A session that's still producing output isn't the same as a session that's still on track. The signal worth watching is whether the agent is converging on the acceptance criteria or drifting into adjacent, plausible-looking work that was never asked for. Stopping early costs a few minutes of restarting; letting a drifted session run to "done" costs a diff that has to be read as carefully as if nothing had gone wrong, because from the outside it looks exactly the same. Most teams learn this the expensive way once, then start checking the signals below before letting a long session keep going unattended.

| Signal | Keep going | Stop and re-scope | | --- | --- | --- | | Diff size | Matches the task | Growing across unrelated files | | Test changes | New tests match the criteria | Existing tests weakened or skipped | | Explanation | Matches what was asked | Solves a different, adjacent problem | | Your confidence | You could explain the diff to a client | You'd need to ask the agent what it did |

What still has to be yours?

The scope, the acceptance criteria, the decision to stop, and the decision to ship: none of those are the agent's job, and none of them show up automatically as evidence once the session closes. A team that only tracks agentic engineering as a skill inside people's heads, passed on informally from one senior engineer to the next, has the same gap that subagent delegation and an instruction file like AGENTS.md already expose at the tool level: what a person decided and authorized during a session isn't the same record as what the session actually did.

Detent, the end-to-end delivery system (detent-ai.com), is built to hold that second record, so the judgment described here leaves something a team can point to once the engineer who ran the session has moved on to the next task — the same criterion Detent Bench measures along the line, from request to production-ready.