Two years into deploying agentic AI on the plant floor, manufacturers have mostly settled the engineering question.
That is, AI systems now detect a degrading bearing, generate a work order, requisition the replacement part, and re-sequence the production schedule with no operator input. The models work.
What is still unresolved is a systems-design problem: setting the control logic that governs when the agent (1) proceeds on its own and (2) hands control back to a person.
According to Deloitte’s 2026 State of AI in the Enterprise report, roughly three-quarters of manufacturers intend to deploy agentic AI within two years, yet only about one in five currently has a model reliable enough to run unsupervised.
That spread between intent and readiness isn’t a modeling gap. It’s a control-system gap, and it behaves like any unmanaged process variable. Left unaddressed, it drifts toward one of two failure modes.
Set the handback threshold too conservatively, and operators start waving through alerts without reading them, the same way an operator ignores a control chart that flags every batch. Set it too loosely, and the system executes on a judgment call as conditions shift out of the envelope it was trained on, and the deviation compounds silently until it shows up downstream as scrap, downtime or a missed order.
Neither failure announces itself at the moment it occurs. Both show up later, in the numbers.
Deployments
Most current deployments still treat human oversight as a binary control: Either the agent runs open-loop or every action queues for sign-off.
That’s a poor model of the underlying variable. Risk and novelty on a production line are continuous, not discrete, yet a routine consumable reorder and a line stoppage triggered by an out-of-spec sensor reading are frequently routed through the identical approval gate.
A better design treats oversight the way process engineers treat control limits; instead of a single approval gate applied uniformly, there are three zones to map decision risk:
Proceed covers low-risk actions that match patterns the system has executed successfully at volume, like routine preventive-maintenance scheduling and part reorders within established points.
Pause covers moderate-risk or moderately novel cases, where the agent keeps collecting data, logs its reasoning, and holds briefly for a quick confirmation before acting.
Escalate covers high-risk or high-novelty cases, where a person reviews before anything executes on the floor.
This tracks with findings from the Stanford University Human-Centered AI 2026 AI Index Report, where report co-chair Raymond Perrault noted that organizations generally lack a working measure of how reliably a system needs to perform in a specific operating context. Absent that measure, oversight policy defaults to a blanket rule instead of a calibrated response to risk and novelty — precisely the gap a three-zone model is built to close.
Field reporting backs this up. Institute of Electrical and Electronics Engineers (IEEE) senior member Ramakrishna Garine has characterized current deployments as running in a semiautomatic, human-in-the-loop mode, where full-function agentic systems exist, but unplanned scenarios still route to a person.
A three-zone model doesn’t invent that behavior; it gives it defined limits instead of leaving the boundary to be redrawn case by case.
Over-Inspection
It’s intuitive to assume that adding checkpoints always adds safety margin. In practice, however, oversight has a saturation curve, much like 100-percent inspection on a line: Past a certain checkpoint density, the marginal safety return becomes negative.
When operators are asked to sign off on every reorder, schedule shift and minor deviation, they stop evaluating each request on its merits and start approving by reflex. Checkpoint control erodes into a formality — control on paper, not in practice.
The data supports treating this as measurable drift rather than anecdote. Analysis from Digital Applied’s 2026 enterprise agent research found that the oversight rate functions as a production trust metric in its own right. A workflow with a low escalation rate and strong adoption behaves nothing like one with a high escalation rate and weak adoption, even though both are technically in production.
On the floor, an escalation queue that grows without a corresponding rise in decision quality isn’t a sign of caution. It’s a sign the control limits are out of calibration and need to be reset, the same way you’d re-baseline a control chart that’s flagging good parts as defects.
Threshold Ownership Belongs With Operations
Deciding where the line sits between routine and escalation-worthy is usually delegated to the group that implemented the system — the software supplier or IT integration team. Neither has the process knowledge to set that boundary well.
Suppliers optimize for a deployment that generalizes across many customers. IT optimizes for uptime and security. Neither group carries the operating knowledge of what a false escalation actually costs on a specific line, or what a missed one actually risks in scrap, downtime or safety exposure.
Setting an escalation threshold is closer to setting a process tolerance than configuring a software parameter. It requires the same plant-specific knowledge that not only tells you when packaging and automotive assembly lines don’t share a tolerance stack-up, but also that plants can differ meaningfully in what “routine” means for their equipment and failure-mode histories.
Operations leaders, the people who oversee the process, should set and own these thresholds, with IT and the supplier supporting implementation rather than dictating policy.
What Metrics Reveal
A human-in-the-loop setup can look fully instrumented while providing no real signal, the equivalent of a gauge that’s technically installed but never calibrated.
A handful of metrics separate functioning governance from a rubber stamp. Track escalation rate as a trend, not a snapshot: A rate that stays flat as volume scales suggests the model isn’t learning to discriminate routine cases from unusual ones, the same red flag as a control chart with no variance at all.
Track time-to-resolution on escalated items, which reveals whether people are actually engaging with flagged cases or just clearing a queue. Also, track the calibration of the system’s own uncertainty estimate, checking that the cases it flags as uncertain are those that actually need correction downstream. That last metric is close to the clearest available test of whether the escalation logic is calibrated at all.
Governance maturity gaps across manufacturing still have room to close. Broader enterprise research compiled this year found that only about 20 percent of organizations have a mature governance model for autonomous AI agents.
That gap is consistent with what shows up on the floor. Closing it depends less on further model improvements and more on manufacturers building the operational discipline to set thresholds deliberately — decision by decision, the way they’d set and maintain any other process control.
(Photo credit: Getty Images/MF3d)