← Blog

Vibe Coding: What Remains as Proof When the Vibe Wears Off

Marco Masut

"Vibe coding" is searched more than any other term in this space, by a wide margin. Most of what ranks defines it, again: where the phrase came from, what it feels like, whether it's a fad or the future of the job. That question is settled by now. The one that isn't: what does a software house have to show, months later, for a feature that got vibe-coded into a client's production system.

What vibe coding actually means today

Andrej Karpathy named the practice in a post in February 2025, and the definition that stuck is close to what he described: you describe what you want in natural language, the model writes the code, and you keep prompting instead of reading the diff line by line. The "vibe" is the part that matters, you're steering by outcome and feel, not by inspecting each change. That's different from agentic coding as a discipline, which also lets a model work across a repository unsupervised but wraps the work in an explicit loop, plan, edit, run, verify, at every step. Agentic coding, done properly, keeps that loop even when nobody's watching each edit. Vibe coding, by definition, tends to skip or shorten it: the whole appeal is not stopping to check.

Why teams love it, and why CTOs get nervous

The appeal is real and it isn't going away. A developer can go from idea to a working prototype in the time it used to take to set up the project. For a solo builder or an early-stage product, that speed is close to free: there's no client contract riding on it yet, and the cost of a wrong turn is an afternoon, not a production incident. The Stack Overflow Developer Survey 2025 shows why the habit spreads past prototypes: 69% of developers using AI coding tools report higher individual productivity. Momentum like that is hard to argue with in the moment.

It's also exactly what makes a CTO or a delivery lead nervous once the same habit reaches a codebase other people depend on. The same survey found only 17% report better team collaboration from the same tools, and the 2025 DORA report puts a number on the trust side of that gap: 30% of respondents report little or no trust in AI-generated code, even as adoption keeps climbing. Vibe coding didn't create that gap. It widens it, because the entire method is built around not stopping to build trust as you go.

The gap between shipping fast and shipping provable

Speed and provability aren't opposites, but they don't come from the same habit. Shipping fast means the feature works, today, for the case you tried. Shipping provable means someone who wasn't in the room, a reviewer, a new hire, an auditor six months from now, can reconstruct why a specific change happened the way it did, without asking the person who prompted it and hoping they remember.

Vibe coding optimizes hard for the first and does nothing by default for the second. That's not a flaw unique to the practice, it's the same gap that shows up across the whole category once a session ends. What AI coding agents leave behind once the session is over is usually a chat log in a proprietary format, tied to whichever tool happened to run that day. Vibe coding just gets there faster and with fewer intermediate checkpoints, because checkpoints are exactly what the workflow is designed to skip.

The risk isn't hypothetical, and it isn't really about whether the code runs. A METR randomized controlled trial from 2025 found that experienced developers using AI tools on real tasks took about 19% longer than working without them, while still believing they'd been faster. If the person who wrote the code, watching it happen in real time, can misjudge what came out of the session, a vibe-coded feature that "feels done" is not, by itself, evidence that it is.

What a software house should keep from a vibe-coded session

None of this means banning vibe coding from a serious engineering shop. It means treating it as a fast first draft, not a finished delivery, and deciding upfront what has to survive the session regardless of how it was written. At minimum, that's four things: the original request, in whatever form it came in, so intent isn't reconstructed from memory later. The boundaries the model was actually working inside, which files, which systems, which data it could touch. What it executed, step by step, not just the final diff. And a review step where a human actually reads the result against those three, before it reaches a client's environment, not after.

That last point is the one vibe coding is built to shortcut, and it's the one a software house can't afford to. The speed of the prompt-and-accept loop is a property of the drafting phase. It says nothing about what happens once the same code is running in front of a paying customer and something breaks at 2am.

From vibe coding to accountable delivery

Vibe coding is a real productivity gain at the point of writing code, and pretending otherwise doesn't help anyone choose a workflow. The problem isn't the speed, it's that speed alone doesn't leave a team anything to point to later. How a software house picks an AI coding assistant already has to weigh this: a tool that's great for a fast draft isn't automatically the same tool, or the same workflow, you want on the path to a client's production system.

The fix isn't a slower vibe. It's a line between drafting and shipping that doesn't disappear just because a model wrote most of the code. The layer that holds that line is what turns a fast, unreviewed session into a delivery a software house can actually stand behind, with intent, authorized scope, execution, and proof tied together, independent of how vibe the session that produced it felt at the time.