, ,

What is impacting quality of Agentic coding?

Previous post detailed specification, architecture, and verification become more valuable as coding gets cheaper, backed by GitClear/GitHub data. Natural next question — squeezed by what, mechanically? Ran a small experiment to find out. One task, one model, one run — treat as a worked example, not proof.

The experiment

A simple code migration, from one programming language to another. I chose migration over code refactoring, new code generation specifically since it has the clearest possible ground truth of any coding task — the source code already defines exact, unambiguous intent — so any quality gap that still shows up can’t be blamed on unclear requirements, only on the mechanics of production and verification.

Result

The model wasn’t incapable of finding any of this — it caught real issues, including in code it didn’t write, the moment it was asked to look. The gap isn’t capability. It’s default behavior — verification doesn’t happen unless someone directs it, scopes it, and checks the result.

What’s actually impacting quality here?

What happened in the experimentWhy it happensWhat it means for AI-assisted coding quality
First draft compiled and ran, but had real problems underneathAI optimizes for “this works,” not “this is good” — nobody asked for more than thatVerification is what catches this gap. A passing test isn’t the same as good code — someone still has to check for it
Nine issues surfaced the moment it was asked to self-review, zero before thatThe model can review code well — but only when someone tells it toThis is verification, and it’s now the valuable step — “AI checked it” only counts if a human explicitly asked for the check
The “trusted” human-written reference had a real, silent bug tooNobody — human or AI — had verified that code either, until askedVerification was never optional, even before AI. AI just made it cheap enough to actually go do it
Even the most detailed, fully-specified fix-everything prompt didn’t work on the first tryJudgment-heavy fixes take more than one pass, even with the full list of problems in handThis is where specification earns its value — the clearer the ask, the less iteration needed, but even a near-perfect spec still isn’t one-shot
This was one small, isolated function — about as easy as a coding task getsReal codebases are far bigger and harder to hold in view all at onceThis is squarely architecture — fitting new code into an existing system is a bigger task than writing one isolated function, and harder for AI to do alone
Nothing stopped the model from rewriting everything in one large passSpeed removes the friction that used to force smaller, safer changesWithout deliberate architecture and verification checkpoints, large fast changes accumulate defects as fast as they accumulate code

So what?

Specification, architecture, and verification aren’t abstract categories from a framework — they’re the exact three things this experiment shows breaking down by default.

None of it points to AI lacking the ability to write good code. It points to quality depending on someone explicitly supplying the specification, checking the architecture, and verifying the result — not something that happens automatically just because the code runs

Caveat: one task, one language pair, one model, one run. Worked illustration of a mechanism, not evidence it holds everywhere.

Then what?

  • If the review checklist is given upfront instead of after the fact, does the naive-vs-reviewed gap mostly disappear?
  • What is the level of automation achieved in specifications, architecture, verification with Agentic tools?
  • Are there tools that integrate with agentic coding that make it easier and faster to get to good quality output?