Enterprise AI is a confounding place to be right now. The frontier model economy brings the unwelcome era of tokenomics, as the diminishing returns of LLM scale loom.
Meanwhile, poorly implemented “AI First” mandates from Meta (and so many others) threaten a dystopian workplace.
But the markets aren’t all-knowing. Your vendor’s shiny new agent doesn’t get the sensational headlines of the Anthropic Fable 5 ban, but no matter: enterprise vendors think they have figured out a few things about AI.
Have they?
Perhaps, but as I argued on a recent DisrupTV appearance:
I don’t see any enterprise vendors that have gotten the context layer right yet, nor do they seem to really grasp the essence of the issues, and I’ve had the opportunity to ask most of them at this point. They are mistaking having solved some of the data components to having solved the context layer problem.
How did we arrive at agentic context? Enterprises need more than AGI fever dreams
The limitations of LLMs have spawned a flurry of advancements, albeit not the kind the frontier models envisioned when they promised AGI (Artificial General Intelligence) to anyone who would listen. Not to mention a trillion dollar ‘autonomous economy’ where humans would be freed up to scramble into a deep existential reckoning of our own.
But enterprise vendors couldn’t wait for AGI fever dreams to get real – and RAG caught on (Retrieval Augmented Generation). What once felt like a desperate band-aid, via RAG-like attempts to influence LLM output, now points to viable use cases across industries. What makes these use cases different?
(more affordable) rightsized models
compound architectures (including RPA/deterministic workflows and tool calls) where LLMs play a key role, but not an all-consuming one, and:
deeper/specific data that constrains/informs language models to an enterprise purpose.
This is nothing like the AGI ambition on which the frontier models’ eye-watering valuations depend – though frontier models have certainly taken note, with the latest flurry of coding tools looped into verification systems that some have argued are closer to neuro-symbolic AI, where the “symbolic” provides a structured version of the truth for a better cross-check… And that’s the kind of architectural discipline enterprises need.
We can awkwardly lump these advancements into the “bounded autonomy” of narrower LLM agents, informed by specialized data. However, these agents don’t necessarily need to be fully autonomous to have useful applications, such as legal document review, aka document intelligence. (I call this granular autonomy, where customers automate at their own pace, according to their own risk profile).
This is all about mitigating the notoriously jagged intelligence of out-of-the-box LLMs. But some enterprise vendors want to blast past data conundrums and right into “agent orchestration.” When analyst Thomas Wieberneit noted that Google Cloud Next was about Gemini for the orchestration layer, I blew a small gasket:
It’s a good summary, and a good play by Google. I care more about getting the context/data layer right, because without context, all you will be doing is orchestrating overhyped pseudo-automated mediocrity with unwanted downstream impacts (which you can observe and audit but not fundamentally fix lol).
We’ve already been through two disappointing generative AI cycles in the enterprise:
Winging it with off the shelf LLMs (including “shadow AI” IP exposure)
Underwhelming generic productivity “copilots,” versus cost outlay (e.g. email composition)
Most of the reports documenting enterprise AI underperformance/failures are about these two phases.
This third phase, contextual/industry AI, is far more promising, though we’re not out of the woods on tokenomics yet, especially if we overdo it with frontier model pricing dependency. The margin for error for successful projects is still small, especially at scale in compliance-heavy settings. No one has expressed this new opportunity better than Christopher Lovejoy:
Our bet really is that when it comes to vertical AI applications, the system that you build for incorporating your domain Insights is far more important than the sophistication of your models and your pipelines. So the limitation these days is not how powerful your model is, and whether it can reason to the level you need it to. It’s more: can your model understand the context in that industry for that particular customer, and perform the reason that it needs to.
Enter context engineering – and the pros and cons of agentic context
This drops us right into the context frenzy. But why now?
Obviously, “context” has now expanded, into the buzzword-laden fields of context engineering, and now, harness engineering, which are both different (but related) ways of talking about how LLM agents can be constrained, informed, and perhaps, “harnessed” into self-improving agentic loops. I like this brief context engineering definition by Langchain I overheard in a lecture, which has been slightly paraphrased:
Context engineering is building dynamic systems to provide the right information and tools in the right format such that the LLM can accomplish the task.
“Harness engineering” adds the twist of keeping the agent on its (guard)rails, and: potentially “looping” it through for task completion, sub-agent handoffs, and, if needed, correction. This definition via Google captures the essence:
While prompt engineering focuses on a single instruction and context engineering manages what the AI sees, harness engineering builds the entire system the agent lives inside.
It would take whole articles to detail this, so let’s run with it for now, and bear down on “context,” because context is what I’m here to debate today – both its limitations, and its overlooked possibilities.
The discipline of LLM context is tied directly to what enterprises care about: the relevance and consistency of the output (no, I don’t consider that “intelligence,” but vendors seem to feel differently under the lights of the keynote stage).
My own research runs smack into the pros and cons of LLM context. The pros? All the useful ways you can make model output relevant (I call this the ‘fascimile of intelligence” – it can be convincing enough to feel pretty darn smart, until that moment where it’s not).
On the other side: the unimaginative ways vendors use the “context” phrase – when in fact their vision of LLM context tends to land solely (and conveniently) in the data of their own products, or, if the vendor in question has a data harmonization or “agentic orchestration layer,” that aspiration might include the context of other enterprise applications as well.
Foundation Capital scorches SaaS with its “decision traces” critique – why does this matter?
That’s a limited view of context, which is why Foundation Capital’s AI’s trillion-dollar opportunity: Context graphs jolted the SaaS industry. Put aside the focus on context graphs. As Constellation’s Esteban Kolsky rightly insisted, that’s just one architectural possibility. Foundation Capital made a much more potent argument than the tool alone.
For all the ridiculous “SaaS is dead” critiques that have emerged from Silicon Valley, Foundation Capital’s Jaya Gupta and Ashu Garg actually have a coherent one. They responded to Jamin Ball’s defense of systems of record:
Ball’s framing assumes the data agents need already lives somewhere, and agents just need better access to it plus better governance, semantic contracts, and explicit rules about which definition wins for which purpose.
That’s half the picture. The other half is the missing layer that actually runs enterprises: the decision traces – the exceptions, overrides, precedents, and cross-system context that currently live in Slack threads, deal desk conversations, escalation calls, and people’s heads.
This is the distinction that matters: Rules tell an agent what should happen in general (“use official ARR for reporting”) Decision traces capture what happened in this specific case (“we used X definition, under policy v3.2, with a VP exception, based on precedent Z, and here’s what we changed”).
Agents don’t just need rules. They need access to the decision traces that show how rules were applied in the past, where exceptions were granted, how conflicts were resolved, who approved what, and which precedents actually govern reality.
Foundation Capital was not necessarily saying SaaS is dead. But they were certainly arguing that SaaS vendors lacked what they believed to be the most important data for organizational decisions. (Yes, the agentic startups they work with are building the architecture to capture those decision traces).
How valid is this position? We can debate it – and I have. But as I see it, all of the proponents of context, from enterprise vendors to agentic startups, are missing the scope of what context could be – and what it actually needs. (If you want a full run down, check the audio replay, AI’s Missing Layer: Why ‘Organizational Truth’ Is the Next Battleground).
At its best, “context” is a white board opportunity, a generational chance to rethink how organizational decisions are made, and how workflows are informed. One the downside: it doesn’t matter how spiffy your frontier model is, if it’s missing (or misunderstands) a crucial piece of context, it looks pretty stupid pretty fast. So, for that matter, does your vendor’s spiffy new ‘autonomous agent,’ happily compounding context adherence problems with each sub-agent handoff.
Here’s the kicker: the crucial pieces of context that make so-called “intelligent agents” look stupid can come from many places, including real-time sources agents may be blind to. Consider this exasperating/mundane example I unfurled on DistrupTV: my computer mouse replacement went terribly awry, when Google Gemini told me the red light indicated a battery failure (in my new model, a red light indicated pairing mode). As I ranted to hosts Vala Afshar and Ray Wang:
So the AI gave me the sh*t information, and this is Google Gemini, which is trained with massive quantities of information, but it was lacking the real-time piece that the design of my mouse had changed. So it misled me, and it was essentially stupid and wrong.
That’s why agents need… “Real-time organizational truth”
Put Foundation Capital’s decision traces on one end of the continuum, and a SaaS vendor’s structured data on the other. Neither have a corner on what agents ultimately need. And that, dear reader, is what led me to my aspirational catchphrase “real-time organizational truth.”
If you want agents to be deeply relevant across use cases, you need to be able to provide that truth, at some level of scale. Is all the data truly real-time? Of course not; that’s just the downside of catchy catchphrases. For each role/workflow/decision, the data requirements are a bit different. Some must be real-time, some not.
That’s the “decision intelligence” white board opportunity: for each company/department/role to ask: what are all the data elements I need for this type of workflow or decision?
But: I have yet to see a vendor – startup or behemoth – put up a comprehensive list of real-time decision ingredients. So let’s sketch out a working list, including that crucial external data:
Agentic memory (short term and long term – the heart of the so-called ‘decision traces’)
Collaborative decisions and policy modifications in Slack, Teams, email and other unstructured channels
Excel spreadsheets and PDFs (ugh)
Structured data in SaaS environments (including deep/localized compliance and regulatory data)
Processes that are poorly documented, or even in people’s heads (what one vendor calls “organizational memory” – one of the least developed, but most promising options on this list)
Modern (cloud) databases/lakehouses, but also a spaghetti maze of legacy databases and siloed systems (heavy lifting alert).
External data (weather data, social media signals, demographic, supply chain data, etc.)
Social media brand signals, and relevant breaking news (e.g. tariff decisions, oil price shifts)
Industrial data, including shop floor and device signals
Sentiment data (why is this never on the context list? AI is pretty good at measuring a brand’s real-time shifts in sentiment)
It’s not a comprehensive list. But: can we claim to be anywhere near real-time organization truth without all these contextual building blocks? And: how compromised will agentic “intelligence” be when it’s missing some of these crucial sources? See: the red light on my supposedly defective mouse.
On a technical level, the ingredients for this context layer could take many forms, and even be stored in different places. They could be delivered via tool call output, or via RAG “chunking,” which is optimization of the top contextual results for a particular query or task. The needed context will shift from agent to sub-agent and so on, but the point stands. For a geekier tech view, here is a live view of an LLM context window, via Vizuara’s Context Engineering course:
LLM context window – via Vizuara’s Context Engineering course on YouTube
How do we get to better AI context? Practical views
If you take “real-time organizational truth” as your frame, as your burning obsession to support agents (and humans), then you must build/integrate/compile the architecture to support it also. Don’t allow vendors get away with a narrow definition of context. Narrow agents might still need broad contextual data in their topic area.
Of course, real-time organizational truth is always, to some degree, aspirational – and it’s always a moving target. LLMs are notoriously stateless – that’s good for some things, but for coherent enterprise workflows, we need stateful data.
The pursuit of real-time organizational truth surfaces a new problem: which stateful context do we include in the ‘context window’ for a particular query? (Example: there are three different discussions of client pricing adjustments in Slack; we need to pull the best one, the definitive stateful information for our agentic needs).
I hear the objection: what is he smoking? Well, I don’t, but – I’m not that far out on a limb here. I’m simply arguing for parallel tracks:
Enterprises should pursue solutions with the minimal viable context needed today, but on a parallel track, they should push for broader. What is that minimum context? Will a customer service AI agent be effective without real-time inventory visibility? (Walmart dropped OpenAI’s “Instant checkout” integration for that very shortcoming). Or: do you wait, and roll out when you have that minimal integrated context? It’ a crucial question.
My spring event travels didn’t lead to many breakthroughs here. Enterprise software vendors were way too focused on the narrower context of their own structured data. But there were exceptions – consider this Shipping Agent example from Sage:
Rob Sinfield, SVP ERP at Sage, walked me through one standout example, where Sage built a shipping intelligence agent that pulled in tariff data, matched it to shipments at sea, and pulled in container tracking data – and then pro-actively recommended actions to the customer. To me, that’s the exciting edge of where AI agents are realistically headed: combining quality system of record data with impactful external sources, and delivering new actions/decision points in real-time. [What separates a good finance agent from a weekend project? Inside Sage’s AI architecture with CTO Aaron Harris]
Yep – start small, notch customer wins, but with that context layer in mind. That’s a different cost model than frontier model pricing hikes. Oh, and a different agentic reality than off the shelf. At that same event, Sage’s Aaron Harris talked about the economic virtues smaller models, finance-specific models, and what they call the “arbiter,” which would fit in the category of harness/context engineering:
Every prompt that goes into the agents and the model goes through this firewall, and every response from the model back to the user goes back through the firewall on the way out. Part of its responsibility is to apply semantics to the conversation that are tuned to the finance and accounting world – that are tuned to the industry that the customer operates in.
That’s the path to smarter enterprise agents, but there are still missing pieces – including the crucial undocumented know-that that Foundation Capital called out. It was refreshing to hear SAP talk about “organizational memory” in its Sapphire 2026 keynote in Orlando. I’ve had early talks with the team developing this; SAP CTO Philipp Herzig calls this “tribal knowledge.” For organizational memory to be a viable part of the context layer, it must be dynamically updated, with the right authoritative data point for that particular agent or query. Vendors that make strides here will be worth watching.
My take – the frontier model economy is flawed. The pursuit of context is a better path
diginomica research has confirmed the garbage-in, garbage-out problem with AI data quality, and the problem is not small: diginomica independent research – the enterprise data health study.
But quality data doesn’t automatically lead to great AI; thus the context conundrum. You’ll need that harness; you’ll need agentic evaluation.
I see big clues in the cloud ERP midmarket, where countless customers I’ve documented swear by the value of living on a “single source of truth.” (That’s a nice head start to better AI, by the way). For most large enterprises, a single source of truth – even on the transactional level – is a fanciful notion. But better context for better AI workflows? That’s an achievable thing.
Proper AI context Is an imperfect pursuit; even the best-engineered context is more fragile that it appears. One of the most useful metrics in agentic evaluation: agents selecting the right tool (Tool selection errors), and then we have context adherence (did the LLM adhere to the context?) and completeness (was the full/necessary context provided?).
A major topic the context engineering course I am taking is the getting the ingredients of context right, including the notorious problem that the middle part of the context is often the least prioritized by the LLM. And, even with expanding context windows, providing AI with too much context can be detrimental; and don’t get me started on context rot. And no, this is not “grounding.” How can you “ground” an LLM in context it is free to ignore? (hear my takedown on grounding in this podcast with Esteban Kolsky).
I don’t consider this the path to cognitive systems. And yes, most enterprises will need some guidance here; this is not easy to build (and maintain) on your own. But I do think it’s the path to better enterprise AI in 2026, and that counts too.
In pursuit of breakthroughs beyond LLMs, some of the bigger luminaries in the AI field have moved on to world models. World models have many potential industrial uses as well; providing LLMs with a deeper real world sensibility could be one of them.
Set AI futures aside: who is best positioned to provide this expanded definition of context, aka my “real-time organizational truth.”? Is it agentic startups? Enterprise software behemoths? Hyperscalers? Data platform vendors? SaaS incumbents? Will enterprises bring this challenge in-house? We are about to find out.
My SaaSpocalypse rethink? Value is being redistributed
Will SaaS vendors win the context layer? Hmm… I’ve had a rethink on the SaaSpocalypse. II still think “SaaS is dead” hyperbole underestimates the embedded value of deep, end-to-end compliance built around automated (and much less expensive) deterministic workflows. Some SaaS vendors fit that description better than others…
This context layer discussion shows one way value is shifting. As my colleague Derek du Preez has noted several times lately, maybe your app is an agent. Maybe “data primacy” wins over applications alone. I don’t think the SaaS layer goes away, but how much will SaaS value be re-distributed? Will SaaS vendors claim some of the other areas of value? Offhand, we can think of:
Context layer (which might include all range of context graphs, and other forms of agent-ready data).
Orchestration, governance and integration layer (including agentic evaluation and observability). This might also include the so-called “harness” that requires each agent to be built (and run) according to secure, built-in governance.
SaaS/software layer – where screens are now less about step-by-step workflow, and more about process validation and exploration, and may even be generated on the fly.
Agent building layer (particularly for customers and partners; may the best ecosystem win)
User experience layer – not to be underestimated. Not all prompt-based environments are created equal. AI First mandates only get you so far – or not far at all. In the end, you need enthusiastic user adoption.
With all those layers of competing value, why am I obsessed with context? Well, because I love seeing how smart LLMs can appear to be, when they are provided with the proper specialized info for a particular user (or industry). If I’m doomed to interact with service bots, how about one that can source the data scope of my relationship?
To me, this is about building cool things – reliably and affordably, with smaller/cheaper models that bring a bit more sanity into the sustainability mix (and a few less tokens). Yes, it still requires use case design, weighing the cost of outliers, and sticking closely to granular autonomy controls.
Which brings us back to where we started this post. Can enterprises counter the economic overreach of the frontier model economy? I’m not sure. The enterprise margin for AI success is still narrow; the architectures required to get context right are complex. But context advancements show us that customers don’t necessarily need frontier model dependence.
That brings us to that fork in the road: customers can go for the AI First pressure cooker, and max up their tokens, or they can architect a different, more sustainable AI – one intended to make the lives of their employees better, and therefore their customers? Alas, these examples are outliers, but they do exist.
I’d be a fool to claim that’s where we’re headed collectively right now, but am I a fool to say that customers have another choice, another way to get to AI success? That’s the real lesson of getting context right. I hope to document much more of that on these pages.