On August 14, 2026, Anthropic switched the default permission mode in Claude Code from step-by-step approval to auto mode, for every Pro, Max, and Team user who hadn't already pinned a different setting. The prompt that used to interrupt a session before almost every file edit, command, or git operation is gone by default. Claude Code now proceeds on its own, unless a classifier decides the action is irreversible, destructive, or aimed outside the environment it's running in.
That's a real change in how the tool behaves day to day, and it's worth being precise about what it does and doesn't touch, because the two get conflated fast in coverage of the move.
What changed on August 14
Before the switch, a new Claude Code session asked for confirmation before almost every consequential action: editing a file, running a shell command, pushing to git. According to Anthropic's own numbers, users approved 97% of those prompts and rejected only 3%, which is the kind of ratio that turns a safety check into a reflex click nobody is really evaluating anymore. Auto mode removes the prompt for everything the classifier doesn't flag, and only interrupts you for the action it judges risky enough to matter. You can still switch back to manual approval at any point, and an organization-managed default doesn't move at all unless an admin changes it.
Why Anthropic made the change
The argument isn't "trust the model more." It's that the step-by-step prompt had stopped functioning as a safety check for most users, and a classifier trained specifically on dangerous actions did the job better than a human clicking through a wall of routine approvals. In controlled testing with 1,053 paid testers, human reviewers caught just 13.6% of genuinely dangerous commands placed in front of them, while the auto mode classifier blocked 89%. Anthropic also reports that manually approved production sessions contained an unintended harmful action 2.6 times more often than auto mode sessions did (6.3% versus 2.4% at production severity). The engineering writeup describes a two-stage classifier that only sees user messages and the tool call being attempted, deliberately stripped of the model's own explanation, so the agent can't argue its way past the check with a persuasive rationalization.
What auto mode still blocks
The classifier's target list is specific: batch-deleting remote git branches from a vague instruction, searching environment variables for alternative API tokens, deleting jobs by name-similarity without confirmation, pushing scripts to a public Gist, retrying a deploy with verification flags skipped. When it blocks something, Claude Code usually finds a safer path on its own or asks you directly. If it hits three consecutive blocks, or twenty in one session, it drops back to manual approval for what's left; in headless mode, where there's no one to ask, it stops the session instead of guessing. Anthropic is explicit that this is not a substitute for human review on high-stakes infrastructure, and that framing matters more than the percentages: the classifier is scoped to catch the actions that were never supposed to happen, not to decide whether the actions that did happen were the right ones.
The real question: not fewer clicks, but what is kept as proof
Fewer approval prompts changes how a session feels. It doesn't change what a team is left holding once the session ends. A blocked destructive command and a signed-off release are two different kinds of safety: one stops an accident mid-session, the other confirms that what shipped was reviewed and authorized by a person willing to be accountable for it. Auto mode is entirely about the first kind. It has nothing to say about the second, and the classifier's own design confirms that: it evaluates one tool call at a time, against a narrow definition of danger, with no memory of what the task was actually supposed to accomplish or who authorized it. What a team should already be checking before it hands Claude Code more autonomy doesn't get easier because there are fewer prompts to click through, it gets more important, because there are fewer moments where a human was forced to look.
That's the same gap most roundups of coding agents skip when they judge agents purely on how well they write code inside a session: what happens after the agent finishes isn't part of the score. Auto mode makes the inside of the session faster and, by Anthropic's own numbers, safer against outright disasters. It does not produce, as output, a record of intent, authorized scope, and execution that a reviewer who wasn't there can check. That was already missing before auto mode; removing the manual prompts just means fewer moments where a person was in the loop by default, not more evidence about what happened once they weren't.
How a software house should adapt review, not abandon it
Turning auto mode off everywhere would trade a real safety improvement for the illusion of control that 97% rubber-stamping already wasn't providing. The better adaptation is to stop treating the approval prompt as where review happened and move that review to where it actually holds up: not "did I click yes fast enough to notice a problem," but "can I show, after the fact, what this session was authorized to touch, what it did, and who signed off before it reached production." A classifier that blocks 89% of dangerous commands earns its place in the loop. It doesn't replace the person who has to answer for a release six months later.
That's the same principle behind the layer that sits above the agent, not inside it: the system prepares, the person signs. Anthropic just moved where the automation sits inside a single session. It didn't move who's accountable for what leaves it.
