{"id":8427,"date":"2026-04-20T15:56:21","date_gmt":"2026-04-20T15:56:21","guid":{"rendered":"https:\/\/www.europesays.com\/ai\/8427\/"},"modified":"2026-04-20T15:56:21","modified_gmt":"2026-04-20T15:56:21","slug":"when-llms-get-lost-in-multi-turn-chats-and-how-oci-stm-helps","status":"publish","type":"post","link":"https:\/\/www.europesays.com\/ai\/8427\/","title":{"rendered":"When LLMs Get Lost in Multi\u2011Turn Chats \u2013 and How OCI\u2011STM Helps"},"content":{"rendered":"<p>Chat is the most natural interface for working with large language models (LLMs). But it also exposes a common weakness: as conversations unfold across multiple turns, models can lose track of what matters,\u00a0especially when requirements, constraints, and corrections arrive gradually.\u00a0\u00a0<\/p>\n<p>That\u00a0resembles\u00a0how people use chat. They\u00a0don\u2019t\u00a0provide a perfectly specified prompt up front. They iterate.\u00a0<\/p>\n<p>The multi\u2011turn problem (it shows up sooner than you think)\u00a0<\/p>\n<p>A common pattern in chat applications is to\u00a0append the entire chat history to every new request.\u00a0It\u2019s\u00a0straightforward and often works initially. But even in\u00a0relatively short\u00a0conversations, it can introduce three practical issues:\u00a0<\/p>\n<p>Quality drift:\u00a0Fragmented context and early ambiguity can cause the model to miss constraints, carry forward outdated assumptions, or get distracted by irrelevant dialog.\u00a0\u00a0<\/p>\n<p>Latency and cost growth:\u00a0As prompts expand turn by turn, time\u2011to\u2011first\u2011token and inference cost typically rise.\u00a0\u00a0<\/p>\n<p>Context window pressure:\u00a0Prompt length grows\u00a0roughly linearly\u00a0and can quickly consume the model\u2019s input budget,\u00a0eventually forcing truncation and loss of earlier context.\u00a0<\/p>\n<p>In other words: as the conversation gets longer (and messier), it gets more expensive\u00a0and more likely to drift.\u00a0[1][2][3][4]<\/p>\n<p>Importantly, while some\u00a0state\u2011of\u2011the\u2011art hosted models include\u00a0built\u2011in conversation compaction\u00a0(often tuned for\u00a0very long\u00a0chats approaching or exceeding the context window), that capability is typically\u00a0model-specific\u00a0and may apply only to a limited set of models. Many\u00a0open\u2011weight models\u00a0also do not provide this behavior out of the box.\u00a0That\u2019s\u00a0why we built\u00a0OCI\u2011STM: a short\u2011term memory capability designed to work consistently across the models we support through the\u00a0OCI GenAI Enterprise AI Responses API.\u00a0<\/p>\n<p>OCI\u2011STM: compact memory, updated as the chat evolves\u00a0<\/p>\n<p>OCI\u2011STM (OCI Short\u2011Term Memory)\u00a0compresses\u00a0multi-turn chat history by periodically replacing older\u00a0conversation\u00a0turns with a compact, structured\u00a0\u201cmemory state\u201d\u00a0that aims to preserve the user\u2019s key requirements, decisions, and constraints\u00a0while removing redundancy and irrelevant dialog,\u00a0and we keep the most recent turns\u00a0verbatim to\u00a0maintain\u00a0recency and fidelity.\u00a0<\/p>\n<p>A key design point:\u00a0condensation runs asynchronously in the background.\u00a0In our reference implementation, this is designed to keep memory-updates off the critical path of the user-facing response, so it typically does not add user\u2011perceived latency,\u00a0while the main model\u00a0benefits\u00a0from\u00a0prompts\u00a0with fewer tokens. [1]\u00a0<\/p>\n<p><img fetchpriority=\"high\" decoding=\"async\" width=\"856\" height=\"492\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/04\/1776700579_663_image-3.png\" alt=\"On the left side, it shows turn by turn of a chat and how prior turns is added to the context of every new turn. This is the baseline for how chat history is usually included. On the right it shows OCI-STM flow, where after 4 turns are accumulated in history context, they are passed through a Condenser operation resulting in shortened history C1. Then, C1 is passed instead of raw first 4 turns, until 4 new turns are accumulated. Then the process repeats to produce C2 which is used in place of raw turns moving forward.\" class=\"wp-image-3335\"  \/><\/p>\n<p class=\"has-text-align-center\">Figure 1: OCI-STM operation overview.\u00a0<\/p>\n<p>In our evaluations, OCI\u2011STM:\u00a0<\/p>\n<p>produces\u00a0fewer\u00a0tokens\u00a0per turn, with larger benefits as conversations grow.\u00a0(Figure\u00a02)\u00a0<\/p>\n<p>can lead to\u00a0net token savings over the lifetime of a conversation, even after accounting for the background condensation\u00a0work.\u00a0(Figure 3)\u00a0<\/p>\n<p>is model-agnostic and works with all models across model families supported by Oracle Cloud.\u00a0<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" width=\"438\" height=\"220\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/04\/1776700580_41_image.png\" alt=\"A chart showing linear increase in chat history tokens for baseline as no. of turns increase in chat going up to 1475 tokens for 10-turn chats. For OCI-STM, the growth is constrained below 750 tokens after 4 turns.\" class=\"wp-image-3332\" style=\"width:427px;height:auto\"  \/><\/p>\n<p class=\"has-text-align-center\">Figure 2: Chat-history token reduction in the main user chat prompt<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" width=\"435\" height=\"217\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/04\/image-1.png\" alt=\"A chart showing linear increase in total token consumption to reach a turn for baseline as no. of turns increase in chat going up to 19000 tokens for 16-turn chats. For OCI-STM, the growth is slowed after turn 7 reaching about 12.5k tokens for 16 turn chats. This demonstrated overall token consumption saving for background condenser process combined with main chat token consumption.\" class=\"wp-image-3334\" style=\"width:405px;height:auto\"  \/><\/p>\n<p class=\"has-text-align-center\">Figure 2: Chat-history token reduction in total token consumption to reach a turn (main chat + background OCI-STM processes token consumption)\u00a0<\/p>\n<p>We also\u00a0observe\u00a0that response quality is\u00a0maintained,\u00a0and in some cases\u00a0improves, relative\u00a0to sending the full raw\u00a0transcript\u00a0in\u00a0each turn in multi\u2011turn scenarios. (see\u00a0Figure 3)\u00a0<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" width=\"346\" height=\"336\" src=\"https:\/\/www.europesays.com\/ai\/wp-content\/uploads\/2026\/04\/image-2.png\" alt=\"Figure showing 3 bars. First bar is for baseline that uses raw chat history scoring a success rate of 78.98%. Second bar is 76.20% for using simple summary, and third bar is 82.25% for OCI-STM.\" class=\"wp-image-3333\"  \/><\/p>\n<p class=\"has-text-align-center\">Figure 3: Task success accuracy sending raw turns, using summarization with our one-off-sequential implementation, versus using\u00a0our\u00a0OCI-STM condenser.\u00a0<\/p>\n<p>We\u00a0validate\u00a0across broad\u00a0domain and dataset coverage\u00a0by using multi-turn benchmarks spanning\u00a0math, code, text-to-SQL, and tool\/function-calling, plus\u00a0robustness\u00a0variants that inject realistic conversational noise and distractors. The core suite includes\u00a0sharded\u00a0versions of standard datasets\u2014GSM8K\u00a0(incremental multi-step math constraints),\u00a0HumanEval\u00a0(requirements revealed over turns),\u00a0Spider\u00a0(schema\/constraint clarification over turns), and\u00a0BFCL\u00a0(multi-turn tool-use specification)\u2014alongside episodic multi-turn categories (recollection, refinement, expansion, follow-up) to test memory and iterative instruction adherence. For\u00a0evaluation method coverage, we measure both\u00a0token efficiency\u00a0and\u00a0task quality\u00a0in multi-turn settings, reporting dataset-appropriate metrics such as\u00a0accuracy\/exact match\u00a0for math\/code\/SQL\/tool calls,\u00a0BLEU or task-specific structured-generation scores, and\u00a0LLM-judge\/rubric-based ratings\u00a0for open-ended refinement and follow-up when automated metrics are insufficient. [3]<\/p>\n<p>How it differs from \u201cjust summarization\u201d\u00a0<\/p>\n<p>Generic summarization may drop key instructions or constraints. Our approach aims to\u00a0maintain\u00a0a structured \u201cmemory state\u201d that is updated over time, with emphasis\u00a0on\u00a0identifying\u00a0and\u00a0retaining\u00a0key user instructions, constraints, and decisions, while condensing supporting or repetitive context. Recent turns\u00a0remain\u00a0verbatim, and older turns are compacted repeatedly into this memory to help the\u00a0model\u00a0can\u00a0stay consistent without carrying the full\u00a0transcript\u00a0every turn.\u00a0<\/p>\n<p>How is it different from \u201cRAG\u201d\u00a0<\/p>\n<p>Retrieval (for example, storing chat history in a memory store and fetching relevant snippets with RAG) can help control prompt length, but it typically adds\u00a0runtime moving parts\u00a0(embedding\/indexing, query formulation, retrieval calls, and re-ranking)\u00a0which can introduce\u00a0additional\u00a0latency\u00a0and operational complexity. It can also miss important multi\u2011turn dependencies when the \u201cright\u201d context is distributed across several turns or depends on conversational ordering. OCI\u2011STM takes a complementary approach by\u00a0maintaining\u00a0an updated, compact conversation state directly.\u00a0<\/p>\n<p>Multi\u2011turn chat is quickly becoming a default interface for working with AI. OCI\u2011STM is designed to help multi\u2011turn experiences scale more smoothly by keeping conversations more\u00a0on\u2011track and efficient\u00a0as they grow, while aiming to preserve the key context the model needs across turns.\u00a0<\/p>\n<p>\u2014\u2014\u2013<\/p>\n<p>Get started today to try out this feature and more.<\/p>\n<p>\u2014\u2014\u2013<\/p>\n<p>[1] Singh, Jyotika, et al. \u2018<a href=\"https:\/\/arxiv.org\/abs\/2604.08782v1\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">MT-OSC: Path for LLMs That Get Lost in Multi-Turn Conversation<\/a>\u2019.\u00a0arXiv [Cs.CL], 2026, https:\/\/doi.org\/10.48550\/ARXIV.2604.08782.<\/p>\n<p>[2] Levy, Mosh, et al. \u2018<a href=\"https:\/\/aclanthology.org\/2024.acl-long.818\/\" data-type=\"link\" data-id=\"https:\/\/aclanthology.org\/2024.acl-long.818\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Same Task, More Tokens: The Impact of Input Length on the Reasoning Performance of Large Language Models<\/a>\u2019.\u00a0Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Association for Computational Linguistics, 2024, pp. 15339\u201315353, https:\/\/doi.org\/10.18653\/v1\/2024.acl-long.818.<\/p>\n<p>[3] Laban, Philippe, et al. \u2018<a href=\"https:\/\/arxiv.org\/abs\/2505.06120\" data-type=\"link\" data-id=\"https:\/\/arxiv.org\/abs\/2505.06120\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">LLMs Get Lost In Multi-Turn Conversation\u2019<\/a>.\u00a0arXiv [Cs.CL], 9 May 2025, https:\/\/doi.org\/10.48550\/arXiv.2505.06120. arXiv.<\/p>\n<p>[4] Liu, Nelson F., et al. <a href=\"https:\/\/aclanthology.org\/2024.tacl-1.9\/\" data-type=\"link\" data-id=\"https:\/\/aclanthology.org\/2024.tacl-1.9\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">\u2018Lost in the Middle: How Language Models Use Long Contexts\u2019<\/a>.\u00a0Transactions of the Association for Computational Linguistics, vol. 12, MIT Press, Feb. 2024, pp. 157\u2013173, https:\/\/doi.org\/10.1162\/tacl_a_00638.<\/p>\n","protected":false},"excerpt":{"rendered":"Chat is the most natural interface for working with large language models (LLMs). But it also exposes a&hellip;\n","protected":false},"author":2,"featured_media":8428,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2],"tags":[24,25,7342,7343,7344],"class_list":["post-8427","post","type-post","status-publish","format-standard","has-post-thumbnail","category-ai","tag-ai","tag-artificial-intelligence","tag-oracle-ai","tag-oracle-cloud-infrastructure-oci","tag-technical-solutions"],"_links":{"self":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/8427","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/comments?post=8427"}],"version-history":[{"count":0,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/posts\/8427\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media\/8428"}],"wp:attachment":[{"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/media?parent=8427"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/categories?post=8427"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.europesays.com\/ai\/wp-json\/wp\/v2\/tags?post=8427"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}