{"id":109360,"date":"2026-07-17T09:51:11","date_gmt":"2026-07-17T09:51:11","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/109360\/"},"modified":"2026-07-17T09:51:11","modified_gmt":"2026-07-17T09:51:11","slug":"qcon-ai-boston-production-ai-moves-beyond-prompts-to-platforms-harnesses-and-evals","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/109360\/","title":{"rendered":"QCon AI Boston: Production AI Moves Beyond Prompts to Platforms, Harnesses, and Evals"},"content":{"rendered":"<p><a href=\"https:\/\/boston.qcon.ai\" rel=\"nofollow noopener\" target=\"_blank\">QCon AI Boston 2026<\/a> marked a turning point. We have spent the last couple of years learning to build AI agents. Now the question is how to run them, safely and reliably, once they are live. Almost every talk came back to the same theme: agents are forcing teams to build real production infrastructure around them.<\/p>\n<p>OpenAI\u2019s Martin Spier set the tone in the opening keynote. His talk was about performance, but not in the narrow &#8220;make inference faster&#8221; sense. There is a quiet stretch before inference where the product has to make the conversation usable for the model: enough context to help, enough trimming to keep it fast. In other words, there is still plenty of work to make the product fast, even when the model is fast.<\/p>\n<p><img decoding=\"async\" alt=\"\" class=\"zoom-image\" src=\"https:\/\/www.infoq.com\/news\/2026\/07\/production-ai-platforms-evals\/news\/2026\/07\/production-ai-platforms-evals\/en\/resources\/226figure-1-1784187384003.jpg\" style=\"width: 2048px; height: 1133px;\" rel=\"share\"\/><\/p>\n<p style=\"text-align:center\">&#8220;The basics became more important.&#8221;<br \/>&#13;<br \/>\nMartin Spier\u2019s &#8220;<a href=\"https:\/\/boston.qcon.ai\/keynote\/boston2026\/keeping-chatgpt-fast-ai-development-accelerates\" rel=\"nofollow noopener\" target=\"_blank\">Keeping ChatGPT Fast as AI Development Accelerates<\/a>&#8220;<\/p>\n<p>It turned out to be a good lens for the rest of the conference. The work around agents is getting less shiny. It is becoming the boring infrastructure work that decides whether a system survives contact with real users. The first recurring trend was context and agent infrastructure rising into a platform layer of its own. Teams are moving beyond single-purpose applications and toward shared systems for context, tool access, identity, and state. This is where ideas like context engineering, MCP gateways, and semantic tool catalogs start to look like core infrastructure. And as core building blocks, they need owners and contracts.<\/p>\n<p><img decoding=\"async\" alt=\"\" class=\"zoom-image\" src=\"https:\/\/www.infoq.com\/news\/2026\/07\/production-ai-platforms-evals\/news\/2026\/07\/production-ai-platforms-evals\/en\/resources\/169figure-2-1784187384004.jpg\" style=\"width: 2048px; height: 1152px;\" rel=\"share\"\/><\/p>\n<p style=\"text-align:center\">&#8220;Precision + Security + Cost&#8221;<br \/>&#13;<br \/>\nFabiane Nardon\u2019s &#8220;<a href=\"https:\/\/boston.qcon.ai\/presentation\/boston2026\/architecting-data-layer-ai-agents-transactional-systems-mcp-and-semantic\" rel=\"nofollow noopener\" target=\"_blank\">Architecting the Data Layer for AI Agents: From Transactional Systems to MCP and Semantic Models<\/a>&#8220;<\/p>\n<p><img decoding=\"async\" alt=\"\" class=\"zoom-image\" src=\"https:\/\/www.infoq.com\/news\/2026\/07\/production-ai-platforms-evals\/news\/2026\/07\/production-ai-platforms-evals\/en\/resources\/142figure-3-1784187384004.jpg\" style=\"width: 2048px; height: 1152px;\" rel=\"share\"\/><\/p>\n<p style=\"text-align:center\">&#8220;Context engineering isn\u2019t a feature, it\u2019s architecture. Get this right and everything else gets easier&#8221;<br \/>&#13;<br \/>\nRicardo Ferreira\u2019s &#8220;<a href=\"https:\/\/boston.qcon.ai\/presentation\/boston2026\/beyond-prompting-context-engineering-production-grade-ai\" rel=\"nofollow noopener\" target=\"_blank\">Beyond Prompting: Context Engineering for Production-Grade AI<\/a>&#8220;<\/p>\n<p><img decoding=\"async\" alt=\"\" class=\"zoom-image\" src=\"https:\/\/www.infoq.com\/news\/2026\/07\/production-ai-platforms-evals\/news\/2026\/07\/production-ai-platforms-evals\/en\/resources\/100figure-4-1784187384004.jpg\" style=\"width: 2048px; height: 1152px;\" rel=\"share\"\/><\/p>\n<p style=\"text-align:center\">&#8220;Own the state. Order the mutation. Prove the action&#8221;<br \/>&#13;<br \/>\nVinoth Govindarajan\u2019s &#8220;<a href=\"https:\/\/boston.qcon.ai\/presentation\/boston2026\/agent-harness-control-planes-invariants-and-approval-boundaries-production\" rel=\"nofollow noopener\" target=\"_blank\">The Agent Harness: Control Planes, Invariants, and Approval Boundaries for Production AI Agents<\/a>&#8220;<\/p>\n<p>The second trend was trust &#8211; a shift from prompt-level guardrails toward trustworthy execution, a harness. As agents gain access to tools and files, security can no longer depend on instructions in a prompt. The harness is the system that sits around the model. A tool can run while the user sees nothing, so production systems need clear ownership of state, ordered writes, approval boundaries, and a real audit trail. The problem is no longer whether an agent gives a good answer, but whether the system can prove what action was taken, by which component, and under which constraints and privileges.<\/p>\n<p>&#8220;The most effective orgs do two things:<\/p>\n<p>&#13;<br \/>\n\tThoroughly improve AI usage across SDLC&#13;<br \/>\n\tResolve the bottlenecks that limit outcomes&#8221;&#13;<\/p>\n<p><img decoding=\"async\" alt=\"\" class=\"zoom-image\" src=\"https:\/\/www.infoq.com\/news\/2026\/07\/production-ai-platforms-evals\/news\/2026\/07\/production-ai-platforms-evals\/en\/resources\/73figure-5-1784187384004.jpg\" style=\"width: 2048px; height: 1152px;\" rel=\"share\"\/><\/p>\n<p style=\"text-align:center\">Lizzie Matusov\u2019s &#8220;<a href=\"https:\/\/boston.qcon.ai\/keynote\/boston2026\/five-stages-ai-maturity-engineering-organizations-where-and-why-teams-get-stuck\" rel=\"nofollow noopener\" target=\"_blank\">The Five Stages of AI Maturity in Engineering Organizations \u2014 Where and Why Teams Get Stuck<\/a>&#8220;<\/p>\n<p><img decoding=\"async\" alt=\"\" class=\"zoom-image\" src=\"https:\/\/www.infoq.com\/news\/2026\/07\/production-ai-platforms-evals\/news\/2026\/07\/production-ai-platforms-evals\/en\/resources\/53figure-6-1784187384003.jpg\" style=\"width: 2048px; height: 1152px;\" rel=\"share\"\/><\/p>\n<p style=\"text-align:center\">&#8220;Write strategy early. Build around customers. Own company-fit surfaces&#8221;<br \/>&#13;<br \/>\nSiddharth Kodwani and Swaroop Chitlur\u2019s &#8220;<a href=\"https:\/\/boston.qcon.ai\/presentation\/boston2026\/building-genai-platform-doordash\" rel=\"nofollow noopener\" target=\"_blank\">Building GenAI Platform at DoorDash<\/a>&#8220;<\/p>\n<p>The third trend was AI adoption itself becoming an engineering operating model. Once usage spreads, the boring questions arrive quickly: who pays for this, who can call which tools, where do failures show up, and how do teams learn from them? Exposing a model through an API or handing engineers a chatbot is not enough. Teams need paved paths, shared policy surfaces, evaluation loops, observability, cost attribution, and feedback mechanisms that make the right behavior easier than the quick, risky one.<\/p>\n<p>A standout topic is how engineering organizations should think about evaluation. A one-shot test can catch obvious failures, but agents do not always fail on the first turn. Single-turn tests and static benchmarks are a weak fit for systems that use tools, maintain state, carry context, and behave differently across turns. So the testing has to get closer to the shape of the product: conversations, traces, simulations, production feedback. Without that, the tests may report success while users hit failures the benchmark never exercised.<\/p>\n<p>Taken together, <a href=\"https:\/\/boston.qcon.ai\" rel=\"nofollow noopener\" target=\"_blank\">QCon AI Boston 2026<\/a> suggested that production AI is becoming less about prompt engineering and more about a systems problem. The hard problems are shifting toward context, data contracts, LLM and MCP gateways, state, evaluation, latency, cost, observability, and, ultimately, security and trust. The harness around the model now matters as much as the model inside it. Agents may talk like coworkers, but they fail like software, and operating them well depends on old lessons from platform engineering and distributed systems.<\/p>\n","protected":false},"excerpt":{"rendered":"QCon AI Boston 2026 marked a turning point. We have spent the last couple of years learning to&hellip;\n","protected":false},"author":2,"featured_media":109361,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[6],"tags":[4077,24,405,7537,205,634,527,5095,56361,1690,56362,56363,314],"class_list":["post-109360","post","type-post","status-publish","format-standard","has-post-thumbnail","category-agentic-ai","tag-agents","tag-ai","tag-ai-agents","tag-artificial-intelligence-agents","tag-infrastructure","tag-ml-data-engineering","tag-model","tag-platforms","tag-production-ai-platforms-evals","tag-productivity","tag-qcon-ai-boston-2026","tag-qcon-software-development-conference","tag-security"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/109360","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=109360"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/109360\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/109361"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=109360"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=109360"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=109360"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}