CI/CD is the practice of automating two separate phases of the software lifecycle: continuous integration (CI), which builds and runs tests on every change as soon as it's proposed, before it merges with the rest of the code, and continuous delivery or deployment (CD), which moves that change, once it clears the checks, toward a staging environment or all the way to production. A CI/CD pipeline is the tool, often a service like GitHub Actions, GitLab CI, or Jenkins, that lines up these steps and runs them on its own for every change, with no need to run them by hand or trust anyone's memory of what got checked.
People who search plainly for "CI/CD" are usually trying to understand what it means and how it applies to their own work, not looking for a specific tool. But the moment that definition actually matters is a different one: when the changes reaching the pipeline are no longer written by a person alone, but partly by an agent.
What does a CI/CD pipeline actually do?
Stripped down, a CI/CD pipeline does three things in sequence on every change: it takes the code, builds or prepares it, and runs it through a series of automated checks, unit tests, integration tests, linting, sometimes a security scan. If every check passes, CI marks the change as mergeable; CD, where it's configured, pushes it forward toward a staging environment or production, often with a human approval step before the last hop. What makes any of this useful isn't the automation itself, it's that every step produces a binary, tracked result: green or red, with a log of what ran and when. Nobody rewrites that result by hand afterward, because it was never written by hand in the first place.
What changes when an agent is the one proposing the change?
With a human developer, the pipeline is an independent check on work a person has already done and can, in theory, account for. With an agent writing code for a large share of the day, the same pipeline becomes something different: it's the first point in the flow where an independent check actually meets the agent's output, because everything upstream of it, the reasoning inside the session, the "I ran the tests and they pass" written into a reply, is verified by nobody except the agent itself. The more changes come out of an agentic session, the larger the volume of code the pipeline has to catch before it becomes a production problem, and the more it matters whether that automated check exists at all and what it actually does, not just that a green checkmark shows up somewhere.
Why does the check the system runs matter more than the one the model reports?
A coding agent can say it ran the tests. It can even "believe" it, in the limited sense that a model generates that sentence because it's consistent with the rest of the session, not because it holds a notion of truth independent of its own output. The reliability data on self-assessment backs this up: in METR's 2025 randomized controlled trial, experienced developers using AI tools on real tasks turned out to be 19% slower than working without them, despite predicting beforehand they'd be 24% faster and still believing, after the fact, that they had been 20% faster. If an expert's own read on their productivity can be off by forty percentage points, an agent's "the tests pass," reported inside the same session where it wrote the code, isn't a check: it's just another line of output, carrying the same reliability as all the others.
| Check run by the pipeline | Check reported by the model | |
|---|---|---|
| Who runs it | A third-party system, outside the session | The agent itself, inside the session |
| What it guarantees | That a specific command ran with a specific result | That the agent wrote a sentence consistent with the expected result |
| Who can tamper with it | Whoever has write access to the pipeline configuration | Anyone who rephrases the prompt, or the agent itself in a later loop |
| What's left afterward | A log with a command, a result, and a timestamp | A line of text inside a chat transcript |
What's worth keeping from every run, and for how long?
A green checkmark alone doesn't last, because almost no platform keeps it forever by default. GitHub, for instance, announced in late August 2026 that starting October 1, 2026, checks, workflow runs, and commit statuses, not just artifacts and logs as before, follow the same configurable retention window: up to 90 days for public repositories, up to 400 for private ones, 90 by default if nothing is changed. Before this change, according to GitHub's own changelog, those three items stayed visible for 400-plus days regardless of any retention setting. It's the kind of detail a software house finds out the hard way, months after promising a client it can reconstruct a specific delivery.
What's worth pulling out of the pipeline before it expires, into a separate record that doesn't depend on the CI platform's own retention settings, is small but precise: which commit was checked, which command ran, with what result, at what point in time, and linked to which request or task caused it to run. There's no need to export the entire raw log of every run: just that handful of fields, because those are exactly what someone will ask about six months after the green checkmark has already dropped off the dashboard.
How do you get from a green checkmark to a record someone can sign?
A green checkmark answers "this command ran successfully, at that point in time." It doesn't answer "who authorized this change" or "who takes responsibility for letting it through." Those are two different questions, and a CI/CD pipeline only answers the first, by design: it's a technical check, not an authorization process. An AI code review tool adds a second automated check on the quality of the diff, but it's still a check, not a signature: neither tool, on its own, produces the object that ties a technical result to a person who stands behind it.
The same holds for Claude Code hooks: they intercept an event the moment it happens, inside the agent's own session, one step before the code ever reaches the pipeline. They're a good place to capture what was attempted. The pipeline remains the best place to check, with an independent system, whether that attempt actually worked. Neither one, alone, is the final record: that record appears once the pipeline's technical result gets tied to the task that requested it, and to a person who signs off before the change counts as shipped. Detent, the end-to-end delivery system (detent-ai.com), holds these pieces together, from intent through automated verification to a human signature, so the pipeline's green checkmark stops being an isolated fact and becomes part of proof someone can show long after that run is over.
