The Review Paradox

Developers who fully delegated to AI finished fastest and scored worst. The skills you need to supervise the model are the first ones it erodes, and you cannot feel it happening.

Share
Bloom 행사 현장

Developers who fully delegated to AI finished fastest and scored worst on evaluation.

The paradox

One observation has been circulating widely in developer communities. Engineers who handed tasks entirely to AI finished fastest and scored worst on evaluation. And the novices who gain the most from AI productivity are precisely the ones who need debugging skills to supervise it, which is the skill AI erodes first.

Call it the review paradox. The more you use AI, the less qualified you are to review it. The two are not separable.

Recognizing good work is not something you get from reading about good work. You get it from doing it badly, being corrected, and accumulating judgment over years.

It applies well outside engineering

This is discussed mostly among developers, but it lands the same way in every knowledge job.

Once AI handles most of the execution, what remains for us? Review, management, planning, strategy. That sounds appealing.

But consider how anyone learned those things. By doing the unglamorous work directly. By making mistakes and getting corrected. The judgment required to review output exists because you did the task hundreds of times yourself.

Take that away and the arrangement breaks. AI executes, you review, and you have no basis for knowing what good looks like. The worst part is not the erosion. It is that you cannot feel it happening.

Review the spec instead

Developer communities have already started solving this, and the principle is simple. Do not review the code. Review the spec and the architecture.

Write a real specification before any code exists. Define the problem precisely, understand the tradeoffs, translate business language into product requirements and product requirements into technical architecture. A human reads and reviews the spec, the architecture, and the verification plan, which means actually understanding what is being built and why.

Then AI writes the code, and the check is whether the code follows the spec. Verifying compliance is something models do well. Judging whether the spec is right is the human job.

Some teams now mandate this. Without a mandate nobody does it, because the path of least resistance is to work in flow and ship whatever comes out.

Typing was never the job

One line captures it:

"Software engineering was never just about typing code. It's defining the problem well, understanding the problem, translating the language from business to product to code, clarifying ambiguity, making tradeoffs, understanding what breaks when you change something."

Swap "software engineering" for any knowledge work and the sentence still holds.

Do not just take the answer

Claude has a Learning setting that trades back and forth with questions instead of handing over an answer. A few months ago that would have felt like friction. It reads differently now. If the thing you are protecting is judgment, then receiving finished answers is the worst possible input.

Intervening at the spec level does not slow you down. It moves your time to the part that decides the outcome.


Join Bloom

Bloom builds offline rooms where people and technology meet. We run them in Seoul, and now beyond it.

Stop formatting proposals. Start winning them. Try Contrl Free Join Beta