AI Pilot Projects and SAP

MIT’s NANDA initiative published a figure last year that ought to worry anyone selling AI transformation: across three hundred enterprise deployments, roughly ninety-five percent failed to deliver a measurable financial return.

RAND‘s own research from the year before, using a different method entirely, arrived at a similarly grim number: eighty percent of enterprise AI projects fail to produce the business value promised at the outset.

These are the closest thing the industry has to independent evidence, and they should reframe how every SAP consultancy talks to clients about AI.

Inside the SAP world, SAPinsider’s most recent transformation research found that intent to adopt Joule grew more than forty percent year on year, while only thirty-four percent of SAP customers report having fully completed their move to S/4HANA or SAP’s cloud offerings. 

Discussion of AI in the SAP ecosystem is running well ahead of the platform work required to support it. This article sets out the specific gates a business has to pass through before an AI pilot can become something that runs a business process.

The 2027 Collision

Most conversations about AI readiness treat data quality, governance, and skills as separate hurdles a business clears one at a time, at its own pace. SAP customers don’t get that luxury.

Mainstream maintenance on ECC ends in 2027 (2030 for customers paying for Extended Maintenance), and every organization still running it is trying to complete a migration to S/4HANA at the same time as its board is asking why Joule hasn’t delivered anything yet. (Joule’s most valuable capabilities depend on the harmonized S/4HANA data model, which a business still running ECC doesn’t have.)

This changes what “AI readiness” actually means for a business under this pressure.

A data model that needs harmonizing for S/4HANA is the same data model an AI agent will need to reason over. A process redesign done properly for migration removes the ambiguity that later trips up an autonomous agent trying to follow that process without constant human oversight. 

Every gate discussed in the rest of this article competes directly against a mandatory deadline that has nothing to do with AI at all, and any consultancy proposing an AI initiative without acknowledging that order of precedence, with all the necessary time and investment, is proposing something unrealistic.

Data Readiness

Gartner’s prediction, repeated consistently through 2025 and 2026, is that sixty percent of AI projects lacking properly governed, use-case-specific data will be abandoned. Gartner’s broader research consistently names poor data quality as the single most common root cause of AI project failure, ahead of model performance or talent gaps. These numbers point at the same underlying problem from two angles, and SAP’s own leadership has been unusually candid about the scale of it.

Christian Klein told the audience at Sapphire this year that harmonizing SAP’s internal data model exposed over 7.5 million data fields requiring semantic context before an agent could use them reliably: a harmonization project SAP began five years ago and still describes as ongoing. That’s the company selling the AI platform admitting its own data still isn’t fully prepared after five years of dedicated work.

For a typical SAP customer, this problem becomes apparent as duplicated master data surviving from old acquisitions, chart of accounts structures that differ business unit to business unit, and years of custom ABAP code holding business logic not documented anywhere else.

An AI agent, unlike an experienced finance clerk, has no instinct for working around an inconsistency that has not been defined. SAP’s answer is Business Data Cloud, paired with the Knowledge Graph. No independent, publicly documented customer case study yet quantifies how much time or cost that combination saves in practice, which is worth remembering before anyone repeats the marketing claim as settled fact.

Process Alignment

Klein offered a useful anecdote at the same event, describing a pricing agent that needed to pull material data from S/4, find pricing information that might sit in Salesforce rather than SAP, and follow a quoting sequence that varies from one customer to the next.

A technically well-built agent fails here because the process it’s meant to execute was never named consistently enough for a machine to follow without a human filling the gaps.

SAP’s own CMO for Finance and Spend Management, Etosha Thurman, made a related point about her company’s event structure: procurement, finance, and supply chain conversations have historically run in separate rooms, addressing separate audiences, even at SAP’s own conferences. 

An AI pilot scoped by one department inherits that department’s boundaries, and those boundaries are where cross-functional processes break down. MIT’s research found the highest financial returns from AI sitting in back-office functions: the same functions built from processes that cross department lines constantly.

Governance and Compliance

Multiple 2025 industry analyses point to the same governance gap: a large majority of executives say their organization has a formal AI governance framework, while a much smaller share say that framework is actually fully operational.

For SAP customers running HCM or finance processes through an AI agent, that gap represents financial exposure. The EU AI Act classifies AI used in recruitment screening, performance evaluation, and workforce management as high-risk, which brings mandatory conformity assessments, documentation requirements, and human oversight obligations with it.

Finance-adjacent uses (anything resembling automated credit decisions or fraud scoring) can fall into the same category depending on the specific use case.

Penalties for violations involving high-risk systems such as HCM or credit scoring can run to fifteen million euros or three percent of global turnover, and the most severe category of violation under the Act can reach thirty-five million euros or seven percent: either of which should be enough to get governance funded properly.

SAP’s response to this is scattered across several separate initiatives.

The API policy changes introduced this year, the organizational memory work previewed at Sapphire, and the hardened API gateways built with Cloudflare all address pieces of the governance and security picture, but no single SAP product yet covers the full compliance lifecycle a customer building agents under the EU AI Act actually needs.

That gap is one a good consultancy can fill with a service offering.

Skills and Change Management

SAP changed its certification model this year, moving toward open-book, scenario-based exams that test whether a candidate can actually solve a problem with the resources in front of them.

That tells you something about where the skills gap really is. The current formal AI credential, the Generative AI Developer certification, covers the Generative AI Hub, AI Core, and BTP integration, and it’s built for developers.

No certification yet tests whether a consultant can use governance tools like SAP’s AI Agent Hub to make sound trust decisions inside a live process: which is a different skill from configuring the tool itself. The scarce skill is judgement about where automation belongs and where a human still needs to sign off.

Google’s DORA research, published in 2025, attributed seventy percent of AI transformation value to people, organization, and process work rather than to the technology itself. Deloitte’s own 2026 survey found only thirty-seven percent of organizations had invested properly in the change management, training, and incentives that make that seventy percent land.

The gap between those two numbers explains a lot of what is seen on the ground: staff using their own ChatGPT accounts because the sanctioned tool doesn’t fit their workflow, and finance teams refusing to sign off on agentic output inside anything touching the close process until someone can show them exactly how a number was produced.

Planning and Scoping Faults

A pilot that never gets abandoned yet never gets scaled is its own particular form of failure, distinct from a pilot that simply doesn’t work. Three separate faults produce it.

The first is a missing success metric defined before the pilot starts, which means nobody has grounds to call it finished either way.

The second is sponsorship spread across a committee rather than held by one accountable person, since a committee can extend a pilot indefinitely without anyone individually having to defend that choice.

The third is subtler: a board sees an AI keynote, mandates an initiative by a fixed date, and hands it to a team with no real business problem underneath it. Half of respondents in a 2025 DSAG, ASUG, and UKISUG survey said they think AI’s current potential is overrated, and this pattern of leadership-mandated AI without a defined business problem is one plausible driver of that skepticism.

The fix for all three starts with picking the right metric before anything else gets built.

The number can then be measured against from the beginning, which removes the ambiguity that keeps a pilot alive by default rather than by merit.

An Example

Picture a mid-sized manufacturing client, mid-migration to S/4HANA, that launches a Joule-based agent meant to handle first-line supplier queries inside procurement.

Six months in, the agent has handled straightforward status checks well and has fallen over on anything involving contract terms, because pricing data sat partly in a legacy spend system nobody had migrated yet.

Nobody had defined what “working” meant before launch, so the pilot kept running on the strength of the queries it handled well, while the failures accumulated in an unreviewed ticket queue.

Diagnosis here is mostly in one category: an integration gap dressed up as a technology problem. The fix was finishing the data migration for that one system, then relaunching with an exception rate target the business had agreed to in advance.

The Intervention Framework

Every stalled SAP AI pilot sorts reasonably well into one of five categories: a technology limitation, a data gap, an integration gap, an adoption failure, or a governance gap.

Naming which one applies changes the corrective action, and most consultancies skip that defining step entirely, treating every stall as a technology problem because that’s the easiest assumption to make.

Governance and data work need to happen before the technical build starts, or run alongside it, never after. A pilot built first and governed later is a pilot that will need rebuilding.

A fixed-fee readiness assessment works well as the entry point for this kind of work, giving a client a defined, low-risk way to find out which of the five categories applies to them before committing to a larger program.

Klein has said publicly that he wants large systems integrators to compress their own delivery timelines rather than stretch them, which tells you even SAP recognizes the old time-and-materials migration model works against the speed clients now expect.

MIT’s research found that specialist partnerships outperform internal AI builds roughly two to one. SAP’s own Joule for Consultants tool, grounded in the company’s non-public enablement content, is already narrowing the knowledge gap between an experienced consultant and a capable client-side team, which changes what a consultancy can charge for and what it needs to be exceptional at instead.

What This Means for You, and What This Means for Your Firm

For the individual consultant, the direct response to a stalling pilot is running the five-category diagnosis yourself before a client asks for one.

A consultant who can walk into a stuck Joule deployment and explain whether the block sits in data, process, governance, adoption, or the technology itself is more valuable than one who can only build. 

That diagnostic skill is what actually ends the limbo for pilot AI projects, which is why it’s the skill the market will pay for. Demonstrating it against a real outcome such as a data cleanup that unblocked a specific agent, or a governance framework that got a finance pilot signed off, is more impressive on a CV than another certificate.

For the consultancy, the same logic scales into three concrete offers.

Sell the readiness assessment as its own engagement, built explicitly around the five-category framework, so a client walks away with a specific answer instead of a vague health check.

Staff engagements so a process specialist and a governance-literate consultant work the same account from day one, since the research shows those gaps are influenced by each other.

Bid against the large SIs on the speed of diagnosis: Klein’s own public pressure on SI timelines suggests even SAP wants speed to matter more than scale now.

A firm that can point to three pilots it moved from stalled to production, and name which gate it cleared each time, has a sales case that is difficult to match. That evidence is what is used to justify the rate premium and lead to the revenue growth.

If you are an SAP professional looking for a new role in the SAP ecosystem our team of dedicated recruitment consultants can match you with your ideal employer and negotiate a competitive compensation package for your extremely valuable skills, so join our exclusive community at IgniteSAP.