The shift The bottleneck moved, and nobody moved with it
For twenty years the constraint on shipping was how fast people could write code. Every tool
worth buying attacked that constraint — frameworks, CI, platform teams, and finally agents that
write the change for you.
It worked. The constraint moved. A team that used to open five pull requests a week now opens
five a day, and the thing standing between those changes and production is the same thing it
was in 2015: a reviewer reading a diff, and somebody clicking through the app if there is time.
That is the shape of the problem. Production got industrialised. Proof did not.
The model Four layers, and the one that is missing
It helps to be precise about which layer does what, because the argument for a verification
layer is that the other three structurally cannot cover it.
| Layer | Question it answers | Status in most factories |
| Generation | Can we produce the change? | Solved, and then some |
| Review | Is the code sound? | Automated and human, working |
| Verification | Does the product do what the change claimed? | Manual, sampled, or absent |
| Release | Can we ship it safely? | Solved — flags, canaries, rollback |
Three of those scale with volume because they are machine work. Verification is the one still
priced in human attention, which is why it is the first thing dropped when the line speeds up.
Why the gap persists The two things that look like verification
Review inspects intent
A reviewer reads what the code says it will do. That catches bad abstractions and unsafe
queries. It cannot catch a flow that breaks in a browser, because the diff does not contain
that information — no amount of reviewer skill extracts it.
Tests encode foresight
A suite checks the properties somebody predicted mattered, at the time they wrote it down.
That is valuable and worth keeping. It is also a fixed asset depreciating against a
codebase that now changes faster than the suite is maintained.
The distinction that matters
Review and tests both work from a description of the software. Verification works from the
software. When output volume rises, only the second one scales, because it does not require
anyone to have anticipated the change.
Requirements What the layer has to be
Not every automated check qualifies. A verification layer that fails any of these degrades into
noise, and a noisy gate is worse than none — teams disable it and lose the signal entirely.
- Independent. Whatever produced the change cannot be what clears it, or a misread requirement survives the whole pipeline intact.
- Unprompted. It runs on every change, not when someone remembers. Coverage that depends on discipline is not a layer, it is a habit.
- Derived, not authored. If a human has to write the check, the layer inherits the bottleneck it was meant to remove.
- Evidentiary. The output has to outlive the run — a recording and repro context, not a claim in a log.
- Calibrated. It stops the line only on a contradiction, and reports everything else. Certainty is the price of being allowed to block.
Where Kery sits A pass that did not write the code
Kery occupies exactly that layer. It reads a pull request, works out what the change is
claiming to do, signs into the preview deployment built for it, walks those flows in a real
browser, and posts a verdict with video of what actually happened.
It does not need a suite, because the checks come from the diff. It does not need CI
configuration, because it runs off the GitHub App. And it fails a build only on a contradicted
check, so it stays switched on — which is the property that decides whether a verification
layer exists in six months or was quietly disabled in week three.
The practical starting point is one repository and one pull request. See
what a check posts back, or
how this compares to what you run today.