Previous post detailed specification, architecture, and verification become more valuable as coding gets cheaper, backed by GitClear/GitHub data. Natural next question — squeezed by what, mechanically? Ran a small experiment to find out. One task, one model, one run — treat as a worked example, not proof.
The experiment
A simple code migration, from one programming language to another. I chose migration over code refactoring, new code generation specifically since it has the clearest possible ground truth of any coding task — the source code already defines exact, unambiguous intent — so any quality gap that still shows up can’t be blamed on unclear requirements, only on the mechanics of production and verification.
- Took a working Python encode/decode (Run Length Encoding) function from Rosetta Code
- Asked Claude Code (Sonnet 5) to migrate to Scala, same functionality
- Asked it to critique its own output
- Compared against Rosetta Code’s own community Scala reference for the same task
- Asked it to critique that too
- Asked for a final version fixing everything found across both

Result
The model wasn’t incapable of finding any of this — it caught real issues, including in code it didn’t write, the moment it was asked to look. The gap isn’t capability. It’s default behavior — verification doesn’t happen unless someone directs it, scopes it, and checks the result.
What’s actually impacting quality here?
| What happened in the experiment | Why it happens | What it means for AI-assisted coding quality |
|---|---|---|
| First draft compiled and ran, but had real problems underneath | AI optimizes for “this works,” not “this is good” — nobody asked for more than that | Verification is what catches this gap. A passing test isn’t the same as good code — someone still has to check for it |
| Nine issues surfaced the moment it was asked to self-review, zero before that | The model can review code well — but only when someone tells it to | This is verification, and it’s now the valuable step — “AI checked it” only counts if a human explicitly asked for the check |
| The “trusted” human-written reference had a real, silent bug too | Nobody — human or AI — had verified that code either, until asked | Verification was never optional, even before AI. AI just made it cheap enough to actually go do it |
| Even the most detailed, fully-specified fix-everything prompt didn’t work on the first try | Judgment-heavy fixes take more than one pass, even with the full list of problems in hand | This is where specification earns its value — the clearer the ask, the less iteration needed, but even a near-perfect spec still isn’t one-shot |
| This was one small, isolated function — about as easy as a coding task gets | Real codebases are far bigger and harder to hold in view all at once | This is squarely architecture — fitting new code into an existing system is a bigger task than writing one isolated function, and harder for AI to do alone |
| Nothing stopped the model from rewriting everything in one large pass | Speed removes the friction that used to force smaller, safer changes | Without deliberate architecture and verification checkpoints, large fast changes accumulate defects as fast as they accumulate code |
So what?
Specification, architecture, and verification aren’t abstract categories from a framework — they’re the exact three things this experiment shows breaking down by default.
None of it points to AI lacking the ability to write good code. It points to quality depending on someone explicitly supplying the specification, checking the architecture, and verifying the result — not something that happens automatically just because the code runs
Caveat: one task, one language pair, one model, one run. Worked illustration of a mechanism, not evidence it holds everywhere.
Then what?
- If the review checklist is given upfront instead of after the fact, does the naive-vs-reviewed gap mostly disappear?
- What is the level of automation achieved in specifications, architecture, verification with Agentic tools?
- Are there tools that integrate with agentic coding that make it easier and faster to get to good quality output?
