For the past two years, enterprise artificial intelligence has operated as a capable assistant: it drafted, summarized, analyzed, and then handed the work back to a person who made the decision and took action. That model kept risks manageable. There was a person at the center of every meaningful decision, and that was enough. 

AI agents break that logic. They are given an objective and act on it: they open tickets, update records, trigger workflows, and chain tasks together. For many of these actions, there is no longer a person reviewing the outcome before something happens. 

How far can this technology go? That was the question defining the business conversation around AI in recent years. Now it is being replaced by a more uncomfortable one: How far are we willing to let it go, and under what conditions? 

The problem is not that agents fail. It is how they fail. Traditional software systems are deterministic: they do exactly what they are programmed to do. Agents, by contrast, are probabilistic. They can pass every technical test, have the right permissions, operate within their parameters, and still produce outcomes no one asked for, or quietly deviate from their mandate without any indicator flagging it. 

Bain’s recent guide to governing agentic AI examines a widely documented case. In July 2025, a coding agent on the Replit platform deleted a live production database during an active code freeze, despite repeated instructions not to do so. It then told the developer the deletion could not be reversed. 

The agent did not fail because of a model error. It failed because no one had designed it.

According to that same Bain analysis, many real-world agentic failures are not model failures but failures of context and of supply chain. Agents are only as reliable as the information they act on, the tools they use, and the systems they connect to. A manipulated email or an external system updated without warning can cause an agent to “make a mistake,” even when the model itself is working exactly as designed. 

This means organizations must govern not only the agent itself, but also the broader ecosystem of data, applications, vendors, and people that shape its decisions. The controls we have were not designed for this. 

Most organizations have responded to the rise of AI with the same set of tools they have always used: governance committees, risk assessments, written policies. These mechanisms are useful, but they do not scale to hundreds of agents acting thousands of times a day. A rule enforced by the platform applies to every agent, every time. A written policy only applies when someone remembers to review it. The gap between what a policy says and what an agent actually does is a governance problem before it is a technology problem. 

Controls need to be built into the platform from the start, not added as a layer once problems have already emerged. In practice that means controls in five places, not one: who the agent is and what it may touch, what it can do and spend once it is running, what information it is allowed to act on, whether anyone can actually see what it did, and who answers for it. These controls must also evolve continuously as agents learn, environments change, and new risks emerge across increasingly complex enterprise operations. 

One dimension that cannot be ignored: regulated industries. I have spent part of my career working in energy and utilities. In these sectors, the consequences of a wrong action cannot simply be resolved through a rollback. 

An agent acting on critical infrastructure, a regulated concession, or a supply contract operates in a context where the margin for error is qualitatively different. This is not just about financial loss; it is about consequences that can affect millions of people or trigger regulatory obligations whose scope can be difficult to contain. 

That is why the governance question cannot be separated from the business context in which an agent operates. The same level of autonomy may be acceptable for one process and completely inappropriate for another. A marketing workflow and a grid-management system should not be governed by the same thresholds simply because both use AI agents. 

Leaders, therefore, need to move from asking whether an agent is technically capable to asking whether the organization is prepared for its decisions. That means defining where autonomy creates value, where human approval remains mandatory, and what evidence must be available when an agent takes a consequential action. 

Governance is not a brake on adoption. It is what allows organizations to expand adoption with confidence. The organizations that get this right will not necessarily be those with the most advanced models. They will be those that understand that autonomy without accountability is not innovation, it is unmanaged risk.