{"id":50604,"date":"2026-05-25T16:55:17","date_gmt":"2026-05-25T16:55:17","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/50604\/"},"modified":"2026-05-25T16:55:17","modified_gmt":"2026-05-25T16:55:17","slug":"llms-corrupt-the-documents-they-work-on-does-agentic-ai-make-it-worse","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/50604\/","title":{"rendered":"LLMs corrupt the documents they work on. Does agentic AI make it worse?"},"content":{"rendered":"<p>First, the researchers curated <a href=\"https:\/\/github.com\/microsoft\/delegate52\/tree\/main\/domain_viewer\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">DELEGATE-52<\/a>, a dataset of documents across 52 different professional domains, from accounting ledgers to aviation bulletins, calendars to crystal structures. Then they designed prompts for performing relevant edit tasks for each document. More specifically, they designed pairs of edit tasks: a \u201cforward\u201d instruction to change the document and a \u201cbackward\u201d instruction that reverses the change. \u00a0Example: for a piece of <a href=\"https:\/\/github.com\/microsoft\/delegate52\/blob\/main\/domain_viewer\/musicsheet.md\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">sheet music<\/a> in G major, the forward instruction \u201ctranspose this up a perfect fourth to C major\u201d is followed by \u201ctranspose this down a perfect fourth to G major.\u201d\u00a0 In other words, pitch the music up, then pitch it down by the same interval. In theory, each \u201cround trip\u201d of forward and backward edit should yield a document identical to the original. But it doesn\u2019t.<\/p>\n<p>The researchers quantified corruption by comparing the original document to the state of the document after each round trip. On average, after just two simulated LLM interactions (or one round trip), 18% of a document\u2019s content no longer matched the original.\u00a0 After six interactions, a third of document content was corrupted. After 20 interactions, the documents were, on average, over 50% corrupted.<\/p>\n<p>The severity of the problem depended on the type of document being worked on. In general, LLMs corrupted documents less when they were repetitive, numerical and structurally dense\u2014they almost perfectly preserved Python code\u2014and corrupted them more when they contained mostly natural language prose such as creative writing or recipes.<\/p>\n<p>Some LLMs <a href=\"https:\/\/arxiv.org\/pdf\/2604.15597#section.4\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">fared better than others<\/a>, but by the end of the simulation even the top three models\u2014Gemini 3.1 Pro, Claude 4.6 Opus and GPT-5.4\u2014had degraded documents by 25% on average. The Microsoft paper makes it clear that LLMs alone, acting outside of any <a href=\"https:\/\/www.ibm.com\/think\/topics\/agentic-architecture\" target=\"_self\" rel=\"noopener noreferrer nofollow\">agentic architecture<\/a>, shouldn\u2019t be trusted with complex documents.<\/p>\n<p>\u201cNone of this is surprising or shocking,\u201d said Mihai Criveti, a Distinguished Engineer at IBM and Chief Architect of <a href=\"https:\/\/www.ibm.com\/products\/watsonx-orchestrate\" target=\"_self\" rel=\"noopener noreferrer nofollow\">watsonx Orchestrate<\/a>, in an interview with IBM Think. \u201cLarge language models are unreliable narrators.\u201d<\/p>\n<p>This, one might think, is a problem we can solve with <a href=\"https:\/\/www.ibm.com\/think\/topics\/agentic-ai\" target=\"_self\" rel=\"noopener noreferrer nofollow\">agentic AI<\/a>. An agentic harness ostensibly equips an LLM with tools, instructions, guidelines and guardrails that help it to execute tasks more reliably and autonomously. So the researchers\u2019 next observation might seem surprising: basic agentic tool use actually made things (slightly) worse.<\/p>\n","protected":false},"excerpt":{"rendered":"First, the researchers curated DELEGATE-52, a dataset of documents across 52 different professional domains, from accounting ledgers to&hellip;\n","protected":false},"author":2,"featured_media":50605,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[6],"tags":[179,30048,7493,10004,6179,4594],"class_list":["post-50604","post","type-post","status-publish","format-standard","has-post-thumbnail","category-agentic-ai","tag-agentic-ai","tag-agentic-architecture","tag-agentic-artificial-intelligence","tag-ai-orchestration","tag-document-ai","tag-natural-language-processing"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/50604","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=50604"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/50604\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/50605"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=50604"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=50604"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=50604"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}