← Blog

Context Engineering: What a Coding Agent Is Allowed to See, and What Proves It

Marco Masut

"Context engineering" is the term the market landed on for a question coding agents made unavoidable: not what you ask a model, but what it can see when it answers. The guides that rank for it, from infrastructure vendors to engineering blogs, are mostly good at explaining how to build that pipeline, retrieval, memory, tool definitions, codebase indexes. Almost none of them ask what a team can still show, once a session is over, about what was actually in the window that produced a specific change.

That gap matters more for a coding agent than for a chatbot. A chatbot's context shapes an answer someone reads once. A coding agent's context shapes a change that ships, and that someone else has to trust months later.

What context engineering means, beyond prompt engineering

The term traces to two tweets, six days apart, in June 2025: Shopify CEO Tobi Lütke first, calling it "the art of providing all the context for the task to be plausibly solvable by the LLM," then Andrej Karpathy, describing it as "the delicate art and science of filling the context window with just the right information for the next step." Both were drawing the same line: a prompt is the instruction you type. Context is everything else the model sees when it acts on that instruction, retrieved documents, file contents, tool outputs, prior turns, system instructions, most of which no one typed at all.

For a coding agent specifically, that "everything else" is not abstract. It's the files the agent read before editing, the tool definitions it had access to, the parts of the codebase its retrieval step decided were relevant, and the instructions sitting in a repository's own configuration. Prompt engineering asks how to phrase the request. Context engineering asks what the agent was working from when it answered it, a question that determines the ceiling on the answer regardless of how well the request was phrased.

What teams build today: retrieval, memory, tool context

In practice, context engineering for a coding agent is an assembly pipeline, not a single setting. A retrieval step decides which files or chunks of the repository are relevant to a task and pulls them into the window. A memory layer carries facts across turns or sessions that would otherwise be lost when the context resets. Tool definitions, MCP servers among them, tell the agent what it's allowed to call and what each call returns. Claude Code and similar tools add a further layer on top: repository-level instruction files that get loaded automatically, shaping how the agent behaves before a single task-specific token is written.

Each of these pieces solves a real problem: without them, an agent works from whatever fits in a single prompt, which is not enough for a task that spans a real codebase. Together, they're also the reason a coding agent's actual working context, on any given task, is something nobody typed and almost nobody can fully reconstruct after the fact.

Why more context is not automatically better context

The instinct, once a team has retrieval and memory in place, is to feed the agent more: more files, more history, more standing instructions. An arXiv study from ETH Zurich is a useful check on that instinct. Testing repository-level context files on SWE-bench and on real developer repositories, the authors found that adding them reduced task success rate and pushed inference cost up more than 20%, compared to giving the agent no repository context at all. More context, assembled without discipline, made the agent both worse and more expensive on the same tasks.

That result doesn't argue against context engineering. It argues against treating volume as the metric. A pipeline that retrieves more, remembers more, and instructs more isn't engineered, it's just larger, and a larger context window filled indiscriminately is closer to noise than to the "delicate art" Karpathy described.

The missing question: what context was actually used in this session

Even a well-built pipeline leaves a specific question unanswered, and current guides don't ask it: for a given change, what did the agent actually have in its window when it made that change, and can anyone check it afterward. A retrieval step might pull the wrong file quietly. A memory layer might carry over an instruction from an unrelated task. A tool definition might return stale data the agent had no way to flag as stale. None of that shows up in the diff. It shows up, if it shows up at all, in a log kept in whatever format the tool happens to use, the same gap what AI coding agents leave out once a session ends already describes for the loop as a whole.

Context engineering, done well, is upstream of that gap, not a fix for it. Getting the right documents into the window at the right time is a retrieval problem. Being able to show, after the fact, which documents actually made it in, and that the agent was authorized to see them, is a different problem, and it's the one every context engineering guide currently skips. It's the same distinction agentic coding's loop of plan, edit, run, verify draws between checking that an output works and checking that the whole process matches what was asked and authorized: verifying the pipeline's design isn't the same as verifying what one specific session actually drew from it.

From context engineering to a provable context ledger

A context pipeline that a team can only describe, never reconstruct for a specific past session, leaves the same hole every other layer of the agentic stack leaves: intent and output exist, but the middle, what the agent was actually authorized to see and did see, doesn't survive past the session that produced it. Closing that requires something closer to a ledger than a pipeline diagram: a record, tied to a specific change, of which sources were in scope, which ones were actually retrieved, and who or what authorized the scope in the first place.

That's a different bar than a good retrieval setup, and no context engineering guide on the market is built to clear it today. It's the same bar the layer above individual coding agents exists to hold, applied one step further upstream, to the context those agents worked from rather than only to the code they produced.