Business presentation in conference room

When polished work becomes abundant, judgment under questioning becomes the real signal.

getty

The first thing AI made abundant wasn’t knowledge. It was competent-looking work. And that changes what organizations should value in the people they hire.

I watched it happen in a single assignment.

Early in a business practicum I teach, students build an issue tree: they take a messy client problem, break it into a clean set of questions, and rough out some hypotheses. It is a foundational consulting skill, structured thinking made visible. Students could use AI, provided they were transparent about it. Roughly half did.

The AI-assisted trees were beautiful. Clean structures, tight logic, plausible hypotheses. Graded the old way, they earned high marks. The trees from students who worked without AI looked ordinary: a reasonable start, some muddled thinking, hypotheses that wobbled. Competent, not impressive.

Then the teams presented their findings. The pattern inverted. The groups with the polished issue trees were vague and off-target. The groups with the messy ones were tracking clearly. The difference traced straight back to that first assignment. Where AI had done the structuring, students had skipped the thinking. They had nothing to stand on when the work got hard. Where students had wrestled with the problem themselves, even clumsily, they had built something that carried them.

AI had improved the artifact without improving the understanding. The grade moved. The learning didn’t.

That gap is not a classroom curiosity. It is the central management problem of the next decade.

The Signal Just Broke

Research has already shown that generative AI can substantially improve the speed and apparent quality of knowledge work. For as long as knowledge work has existed, we have judged people by what they produce. A sharp memo, a clean model, a tight deck these were hard to make, so making one told you something true about the person who did. The output was a proxy. We rarely inspected the thinking directly because we didn’t have to; the work vouched for it.

AI severed that link. A good report is no longer evidence of a good thinker. It is evidence that someone had access to a good model. The proxy we have leaned on for a century, output as a stand-in for underlying capability, stopped being reliable almost overnight, and most hiring, promotion, and grading systems have not caught up.

This is subtler than a cheating problem, and more serious. Cheating is a violation of a rule that still works. This is the rule itself failing. When the deliverable no longer discloses the thinking behind it, every process built on that inference — the resume screen, the take-home assignment, the writing sample, the performance review anchored to “quality of work product” — is measuring something it can no longer see.

The organizations that notice this first will stop asking what did you produce and start asking what did you decide, and why.

What Stayed Scarce

Strip out what AI made abundant and look at what’s left.

A model can draft the stakeholder analysis. It cannot notice that the stakeholder in the room is the one who will quietly kill the project. It can generate three defensible options. It cannot tell you which one you should stake your name on when the data is incomplete and the client is impatient and both answers are partly right. It can produce the recommendation. It cannot be accountable for it.

What stayed scarce is judgment: the weighing of trade-offs no rubric anticipated, the read on a room, the decision held together under ambiguity and defended when someone pushes back. This was always the thing employers were paying for. The report was just how they found it. Now the report is free and the judgment is not, which means the price of one is collapsing while the price of the other climbs.

Here is the trap, though. Judgment cannot be produced by automation, and it cannot be developed by it either. It is built the way it has always been built, through friction, consequence, and the accumulated experience of being wrong in ways that cost something. An organization that routes all of its cognitive work through AI in pursuit of this quarter’s efficiency is quietly disinvesting in next quarter’s judgment. The productivity shows up immediately. The risk of cognitive atrophy shows up later, a concern reflected in research examining how generative AI affects critical thinking.

Why Experiential Work Suddenly Matters More

The fix is not less AI. It is putting people in situations where AI can’t do the part that counts — and where the failure to think is immediately visible.

That is exactly what the issue tree exposed. AI produced a flawless structure and left the comprehension hollow, and only a live, high-stakes moment — a real presentation, with real questions — revealed the hollowness. When a person has to defend a recommendation to an actual client, present findings in real time, or revise a plan after a stakeholder pushes back, the question shifts. “Did AI write this?” stops mattering. “Walk me through your reasoning” starts mattering. The live context is the assessment. A fluent draft cannot game it.

I have watched AI used well in exactly these settings, not as a shortcut, but as a sparring partner. Teams interviewed AI personas built around difficult stakeholders and learned that the same recommendation lands differently depending on who is hearing it. A coaching AI pushed teams to surface their working assumptions before conflict made those assumptions personal. In each case the machine raised the difficulty rather than removing it. That is the tell. AI that eliminates the struggle erodes judgment. AI that sharpens the struggle builds it.

This is also where real AI literacy develops, not platform fluency, but the harder discrimination of knowing when the tool is useful, when it is misleading, when to check it, and when to shut it off. That judgment is learned in places where shallow thinking has visible consequences. A generic analysis fails the moment the client knows her own industry. The feedback is immediate, contextual, and impossible to fake; everything a rubric is not.

What This Means for Anyone Who Hires

The organizations that win the next decade will not be the ones that adopt AI fastest. Adoption is becoming universal and therefore worthless as an advantage. The winners will be the ones that rebuild their screenings for judgment while everyone else keeps reading outputs that no longer mean what they used to.

That has concrete implications.

Stop treating output as evidence of capability. A polished deliverable used to signal a capable person. It now signals access to a capable model. Hire and promote on demonstrated reasoning, how someone weighed the trade-offs, not on the artifact alone.Interview for the decision, not the deliverable. Put candidates in front of an ambiguous problem and watch them reason out loud. What they defend under questioning tells you what a take-home never will again.Protect the productive struggle. Offload cognitive work where speed matters more than understanding. Guard the work where the struggle is the point. Make that a deliberate choice rather than a default.Value the experiences that build judgment. The people worth premium compensation are the ones who have been tested against real ambiguity, real pushback, real consequences; not the ones with the cleanest portfolio a model could have produced.The Bottom Line

AI did not make expertise obsolete. It made it easy to replicatecheap, and in doing so it revealed what expertise was always standing in for: judgment under uncertainty, exercised by someone willing to be accountable for the outcome.

The work no longer speaks for the worker. Learn to hear the difference, or keep paying a premium for a signal that stopped meaning anything.