# AI Agents Need Memory Harnesses for Long Tasks _Stefania Druga of Sakana AI discusses memory harnesses for AI agents, addressing context bloat and the benefits of local models for long-running tasks._ **Updated:** 2026-08-22 **Published:** 2026-08-12 **Source:** https://www.startuphub.ai/ai-news/ai-research/2026/ai-agents-need-memory-harnesses-for-long-tasks --- Stefania Druga, a research scientist at Sakana AI in Tokyo, recently addressed a critical challenge facing advanced AI agents: the problem of "context bloat" in long-running tasks. Speaking at the AI Engineer World's Fair, Druga highlighted how AI models can start to contradict themselves, repeat work, or drift from their objectives when they struggle to retain information over extended periods. Context BloatDriver AI models contradict, repeat work, or drift from objectives in long tasksFrom the article 6 mentionsStefania Druga, a research scientist at Sakana AI in Tokyo, recently addressed a critical challenge facing advanced AI agents: the problem of "context bloat" in long-running tasks.amplified byLong-Horizon TasksDriverincreasingly complex and extended tasks amplify context rot issues for AI agentsFrom the article 8 mentionsA second set of experiments used the X-Bench benchmark for long-horizon tasks.drives need forLocal Models RiseContextbeneficial for long-running tasks, but need better memory managementFrom the article 2 mentionsDruga noted the increasing viability of local models, such as GLM and Deep Seek V4 Flash, for agentic tasks and tool use, even with the ongoing RAM bottlenecks.requireMemory HarnessCoreStefania Druga's solution for AI agents to manage information retentionFrom the article 7 mentionsDrugo outlined her harness design, conceptualizing memory as a "write-manage-read loop." This system aims to provide durable memory for research agents that otherwise have none.usesWrite-Manage-Read LoopCorecore design of the memory harness for effective information handlingFrom the articleDrugo outlined her harness design, conceptualizing memory as a "write-manage-read loop." This system aims to provide durable memory for research agents that otherwise have none.tested viaTest EffectivenessContextliterature review and long-horizon tasks used to validate memory solutionsFrom the articleTo evaluate the effectiveness of memory harnesses, Druga conducted experiments on two key task types.informsRecall PolicyOutcomekey takeaway for efficient and cost-effective memory management in agentsFrom the article 7 mentionsDruga's findings indicated that a "ranked recall" policy, which prioritizes retrieved memories, performed best.leads toCost EfficiencyOutcomeoptimized memory reduces computational expenses for AI agent operationsFrom the article 2 mentionsShe found that when all relevant information fit within the context window, memory didn't significantly improve performance and only added cost. ## The Growing Challenge of Context Bloat Drugo explained that as AI models are tasked with increasingly complex and longer-horizon tasks, the issue of context bloat becomes more pronounced. This "context rot" manifests in several ways: contradictions, where the model revises earlier decisions; redundant work, as the model forgets it has already completed a task; and drift, where the agent loses focus on the original objective. This problem is amplified by a trend towards fewer, but more capable, model releases, meaning a greater reliance on existing models for more demanding tasks. ## The Rise of Local Models and Their Memory Needs The presentation also touched upon the growing trend of utilizing local AI models, citing Coinbase's success in reducing AI spend while increasing usage through a transition to local models and improved practices like better routing, caching, and context management. Druga noted the increasing viability of local models, such as GLM and Deep Seek V4 Flash, for agentic tasks and tool use, even with the ongoing RAM bottlenecks. She shared her personal setup, a Mac Studio with an M3 Ultra chip, running two such models: Qwen 27B (4-bit quantized) and Deep Seek V4 Flash. ## Designing the Memory Harness: A Write-Manage-Read Loop Drugo outlined her harness design, conceptualizing memory as a "write-manage-read loop." This system aims to provide durable memory for research agents that otherwise have none. The harness includes a core component that is always shown to the agent, a recall block for testing different memory retrieval modes, and an archival block for cross-session information persistence. She tested various recall modes, from no memory at all to vector-based retrieval, a "decisions ledger" that tracks decisions made at each turn, and an "oracle" that provides the ground truth for memory retrieval. ## Testing Memory Effectiveness: Literature Review and Long-Horizon Tasks To evaluate the effectiveness of memory harnesses, Druga conducted experiments on two key task types. The first involved a literature review where the agent needed to identify a retracted scientific claim from a large corpus of papers. She found that when all relevant information fit within the context window, memory didn't significantly improve performance and only added cost. However, for tasks where the necessary context exceeded the context window, a good memory harness proved invaluable. A second set of experiments used the X-Bench benchmark for long-horizon tasks. Druga's findings indicated that a "ranked recall" policy, which prioritizes retrieved memories, performed best. She explained that even with an "oracle" providing the correct information, the model might still fail to retrieve it accurately if not properly guided. The ranked policy consistently outperformed other methods across different models and benchmarks, including Spider V2. ## Key Takeaways: Recall Policy and Cost Efficiency Drugo emphasized that "bad memory is expensive" because it wastes tokens and can lead the agent astray, while a well-structured recall policy can save both tokens and budget. She encouraged the AI community to consider the recall policy as a "first-class metric," urging developers to think critically about the types of memories they store, how they rank them, and how to design effective recall functions that persist across multiple runs and sessions. The research demonstrated that better recall not only improves accuracy but also reduces costs, making it a crucial area for future AI development. Druga also highlighted the importance of sovereign capability in AI, emphasizing the control gained by running and evaluating models locally, which she sees as a key aspect of her work at Sakana AI. --- Original analysis from [startuphub.ai](https://www.startuphub.ai), the #1 AI startup directory.