Search "claude agent sdk" today and the results are almost entirely how-to: the official repository, the docs, a handful of tutorials on wiring up a first agent. That's the right content for someone evaluating whether to build on it. It's not the content for someone who already decided to, and is now shipping that agent to a client.
What the Claude Agent SDK is, and who is building on it
The Claude Agent SDK is Anthropic's toolkit for building custom agents on the same infrastructure that runs Claude Code: the harness that manages context, calls tools, and loops on a task until it's done. It's available for Python and TypeScript, and it's the same engine Anthropic uses internally for Claude Code itself, exposed so other teams can build their own agent on top of it instead of just using the packaged one.
The teams reaching for it aren't hobbyists. They're software houses and internal platform teams that want an agent shaped around one specific workflow, one specific set of tools, one specific approval flow, rather than the general-purpose session Claude Code ships out of the box. Building on the SDK means the agent architecture becomes something the team owns and can shape, not something they configure around the edges.
Same engine as Claude Code, a different set of responsibilities
Claude Code and an SDK-based agent share the same core loop underneath: read context, decide on a tool call, execute it, observe the result, decide again. What differs is who is accountable for the parts around that loop. Claude Code, as a product, ships with permission prompts, session logs, and a support surface Anthropic maintains. An agent built on the SDK inherits the engine, not the product decisions wrapped around it: the team writing the agent decides what gets logged, what a human has to approve, and what happens when a tool call fails halfway through a task.
That's the trade a team makes on purpose, and it's usually the right one when the workflow is specific enough that the packaged tool doesn't fit. But it means every question a buyer would ask about Claude Code's own guardrails becomes a question the team has to answer about its own agent instead, because there's no vendor default to point to anymore.
What ships out of the box versus what a team has to add
The SDK gives you the mechanics: tool orchestration, context management, streaming, hooks to intercept actions before or after they run. It doesn't give you an audit trail by default. A hook can capture that a tool ran, but nothing forces the team to store that record anywhere durable, timestamp it, or tie it back to the request that triggered it. That has to be built, deliberately, on top of the hooks the SDK exposes.
The same gap shows up around authorization. The SDK lets a team wire in permission checks before a tool executes, but it doesn't ship an opinion on what should require a human signature versus what can run unattended. Claude Code's own product team made those calls for the packaged tool and keeps revising them; a team building on the SDK inherits none of that judgment automatically; it has to make the same calls itself, for its own agent, and be ready to defend them.
The proof gap that opens up when you build the agent yourself
This is where the responsibility genuinely shifts. With a packaged coding agent, a buyer can at least ask the vendor what the tool logs and how. With a custom agent built on the SDK, there's no vendor to ask: the log format, the retention, the link between a specific code change and the request that authorized it, all of that is a decision the building team made, or more often didn't make explicitly, somewhere in the codebase.
That gap doesn't show up during a demo. It shows up months later, when a client asks why a specific change was made, and the honest answer is that the agent's run logs were never designed to answer that question, because nobody treated "who can reconstruct this after the fact" as a requirement while the agent was being built. A custom agent that writes excellent code but leaves nothing structured behind about how it got there hasn't reduced the team's exposure, it's just moved it from "which tool" to "which internal decision, made by us, with no vendor to point to."
A short checklist before shipping a custom agent to a client
Before an SDK-based agent goes anywhere near a client's codebase, a few questions are worth answering explicitly, not left implicit in whatever the hooks happen to log:
- What triggered this session, and is that request recorded anywhere durable, not just in a chat history that can be edited or lost?
- What context was the agent authorized to read, and is that scope written down anywhere outside the prompt itself?
- Which actions ran unattended, and which required a human to approve, and is that distinction enforced in code or just in a policy document nobody checks?
- If a client asks, six months from now, why a specific file changed, can anyone answer from the log, or does the answer depend on someone remembering the session?
None of this is specific to the Claude Agent SDK. Any coding agent raises the same question once the session ends. What's different here is that with a packaged tool, a team can at least point to the vendor's own answer, imperfect as it is. Build the agent yourself, and the answer has to be one the team wrote, because the checklist a buyer runs before handing an agent to a team doesn't stop mattering just because the agent is custom, it gets harder to pass. Detent Bench measures exactly that line, from request to production-ready, whichever engine is doing the writing.
