For more than a decade, executives have been promised that the next wave of AI would finally deliver transformative productivity gains. Harang Ju, who co-directs the AI Agent Lab and studies the organizational dynamics of human-AI collaboration, thinks most companies are about to repeat a mistake industrial firms made a century ago. 

Ju’s central argument is simple: dropping a more capable AI agent into an unchanged job is neither collaboration nor transformation. It is decoration. Real gains, he says, come only when organizations rethink who holds the working state of a task, who verifies outcomes, and where human judgment is deliberately preserved as agents take on more of the execution.

The Electric Motor Problem

Ju reaches for a 100-year-old analogy to explain why so many companies are struggling to capture value from AI agents today. When factories first replaced steam engines with electric motors, he notes, most kept their existing layout of belts, pulleys, and shafts running off a single central power source. The result was that productivity barely moved for decades. It was only when managers rebuilt their factories from scratch, distributing smaller motors directly to individual machines and rethinking the entire floor plan, that electricity’s real advantages showed up in the numbers.

“I see the same pattern today,” Ju, who will be speaking at AIRF 2026, says. Organizations are buying powerful agentic tools and fastening them onto processes designed for a pre-AI world: data still trapped in scanned PDFs, agents forced to click through interfaces built for human eyes rather than working through APIs, and job descriptions that haven’t changed even as the nature of the work has. “It is easy to buy the motors,” he says. “It takes real work to redesign the factory.”

That redesign has a clear shape. Agents should run the day-to-day execution of a process. People should own the junctions that matter, the moments where goals are set, outcomes are verified, and work is pulled back for correction. Without that explicit division of labor, he warns, companies get an impressive demo and little else. “The ultimate goal is not maximum autonomy, but a deliberate redesign of our workflows as agents take on more of the process.”

Four Tests for What Stays Human

If agents are going to run more of the process, leaders need a disciplined way to decide which decisions they can safely hand over and which must stay with people. Ju offers four tests that he uses to draw that line.

The first is checkability. If an agent can draft something in seconds, but it takes a human days of tedious effort to verify that the output is actually correct, the decision should remain human-led. Speed of generation, in other words, is worthless if verification can’t keep pace.

The second is the stakes. Low-risk, reversible work is a natural candidate for agent autonomy; high-risk, hard-to-reverse decisions are not. This sounds obvious, but Ju argues it is routinely violated in practice.

The third is judgment. Wherever a decision involves conflicting priorities, ethical trade-offs, or organizational values, the kind of call that doesn’t have a single correct answer derivable from data, that judgment has to remain with people.

The fourth is duration. The longer an agent operates without a human check-in, the more likely its behavior is to drift from its original intent, making periodic human touchpoints essential even for otherwise well-functioning agents.

Ju says many organizations get this sequencing backward. “Many firms automate the wrong tasks first by targeting high-stakes analysis that is incredibly difficult to verify, simply because the initial demo looks impressive.” 

A flashy pilot on a hard, high-consequence problem may win executive attention, but it is precisely the kind of task that fails the checkability and stakes tests simultaneously. His guidance is close to a rule of thumb for the C-suite: Never hand AI a decision you cannot afford to verify.