Businessman stacking ROI blocks with coins and calculator

A Businessman stacking ROI blocks with coins and calculator

getty

Most executives don’t fear AI.
They fear being embarrassed by AI in a board meeting, an audit, or a budget review, when someone asks a simple question:

“Is this creating value… or just creating activity?”

Now imagine sitting in the room, when the team walks in with charts: copilots deployed, prompts written, “hours saved.” And then the CFO leans forward:

“Show me the proof I can defend.”

That’s the moment many AI programs lose credibility, not because the technology failed, but because the measurement system rewarded the wrong behavior.

In jazz, you don’t judge a bassist by how many notes he plays. You judge him by whether the band can trust the groove. AI ROI works the same way: counting activity is easy; proving impact is the hard part.

According to Arvind Narayanan & Sayash Kapoor in their book AI Snake Oil: What Artificial Intelligence Can Do, What It Can’t, and How to Tell the Difference, state that “AI reflects its training data. It learns patterns about the people who make up the data, and the decisions made by AI reflect these patterns. But when the decision subjects come from a population with different characteristics than those in the training data, the model’s decisions are likely to be wrong.”

Here’s the good news: 90 days is enough to produce decision-grade proof, if you stop measuring AI like a novelty and start measuring it like an operating system.

The “what is” problem: AI dashboards invite metric theater

Right now, many organizations are “winning” on the dashboard while losing in reality.

That’s not because people are dishonest. It’s because metrics don’t just measure performance, they shape it. “Overemphasizing metrics leads to… manipulation, gaming, and a myopic focus on short-term qualities and inadequate proxies.” (ScienceDirect)

When careers, budgets, and narratives depend on a number, teams will find a way to make the number look better, sometimes while the business quietly gets worse.

And AI makes this easier to mess up because teams often measure what’s available:

Tool usageContent volumeSelf-reported “time saved.”A demo-set accuracy score

Those are often measures of activity, not measures of value.

So, the mandate isn’t “get better metrics.”
It’s: build proof that resists gaming.

The “what could be” alternative: re-constructible proof in 90 days

If you want AI ROI that survives a CFO cross-examination, you need a standard that doesn’t rely on belief.

Here’s the board-ready test:

Can Finance reconstruct the result?
Not “Does the story sound plausible?”
Not “Is adoption trending up?”
But: Can a skeptical reviewer follow the evidence from baseline → method → outcome → tradeoffs → economics → decision?

That’s what a Proof Pack is for: it turns AI ROI into an evidence case, not a vibe.

The PROOF-90 method: a 90-day Proof Pack boards can trust

I use a simple operating method: P.R.O.O.F. 90, a cadence designed to make metric manipulation harder than real improvement.

P — Pick one unit of value (don’t measure “the model”)

AI ROI becomes defensible when you can point to one unit of value:

One workflow (e.g., contract review, customer support triage, underwriting, procurement exceptions)One decision owner (someone accountable who can validate the outcome)One measurable outcome (cycle time, error rate, cost-to-serve, conversion, risk reduction)

According to Eric Siegel, in his book, The AI Playbook: Mastering the Rare Art of Machine Learning Deployment, he states, “I say that my definition of project success is when a model has been developed and deployed such that it has created—note the past tense—business value for the organization that paid for it. When you impose that criterion, man, it’s quiet out there.”

If you can’t name the decision, you can’t prove the ROI.

R — Register the baseline (and your “doesn’t count” rules)

Before the pilot begins, register three things:

Baseline performance (what is true today)Definition of success (what must improve)What doesn’t count (so the metric can’t be inflated later)

“Goal setting [should be] a prescription-strength medication that requires careful dosing, consideration of harmful side effects, and close supervision.” (Harvard Business School)

This one move kills most gaming, because gaming thrives in ambiguity.

O — Observe behavior change in the workflow (not just “usage”)

ROI is not the number of people who tried the tool.
It’s whether the workflow changed:

Are decisions faster and correct?Are exceptions decreasing?Are escalations dropping?Are humans relying on AI in the moments that matter, or only when it’s convenient?

Ethan Mollick, in his book, Co-Intelligence: The Definitive Guide to Living and Working with AI, states, “AI adoption is happening much more quickly, and much more broadly, than previous waves of technology. And we are still unclear as to what the limits, and possibilities, of this new technology are, how quickly they will continue to grow, and how ahistorical and strange the effects might be.”

Usage can be mandated. Workflow improvement must be earned.

O — Offset with counter-metrics (every win needs a bodyguard)

Any success metric that can improve while the business gets worse is not an ROI metric. It’s a gaming invitation.

So, every “win metric” needs bodyguard metrics, signals that protect quality, risk, rework, compliance, and trust.

Examples:

Faster cycle time → rework rate/defect rateLower cost → quality score/customer impactMore throughput → escalations/overridesMore automation → exception volume/compliance flagsF — Finance + forensics (translate value and preserve the evidence trail)

Two things turn AI ROI into CFO-grade proof:

Finance translation: unit economics, assumptions, sensitivity ranges, cost-to-deliver, and time-to-valueForensics: an evidence archive (baseline data, change log, limitations, monitoring plan, governance posture)

The goal isn’t to “win the pilot.”
The goal is to produce enough clean evidence to make one decision: scale, hold, or kill.

The one-page board view: PROOF-90 executive scoreboard

If you want the board to trust your AI results, keep the “board view” brutally simple. Use six lines:

Unit of Value — What workflow decision did AI improve?Baseline — What was true before AI?Outcome Improvement — What got better?Counter-Metric Stability — What did not get worse?Financial Translation — What is the economic value (and assumptions)?Governance Posture — Can we defend and monitor it?

This makes the conversation executive-ready: What changed? What didn’t get worse? What decision follows?

A practical 90-day operating timeline

Here’s a cadence you can run immediately:

Days 1–10: Choose the workflow, decision owner, baseline, and counter-metricsDays 11–30: Instrument the workflow and capture baseline realityDays 31–60: Run the pilot and review weekly evidence (not stories)Days 61–90: Translate results into unit economics and decide scale/hold/kill

One rule: treat ROI as a causal question (“compared to what?”), A/B, staggered rollout, matched controls, or another quasi-experimental design, so the story can’t be rewritten after results appear.

Outcomes over hype

The fastest way to kill an AI program is to reward theater.

AI ROI is not proven by AI activity. It is proven when one important workflow decision improves relative to a clear baseline, while counter-metrics show the business did not get worse elsewhere. A 90-day pilot should not try to prove enterprise transformation. It should produce enough clear evidence for Finance and the board to make one honest decision: scale, hold, or kill.

So, lead differently:

Reward outcomes, not activityReward learning, not dashboardsReward proof, not hype

Like a great bassist, you don’t accelerate when the room gets loud. You lock the groove so everyone else can play faster with confidence.