PRO-LONG: Memory for LLM Agents

PRO-LONG revolutionizes LLM agent capabilities for long-horizon tasks, achieving SOTA performance with drastic token efficiency and cost reduction.

Diagram illustrating the PRO-LONG framework for LLM agents with a structured interaction log and search mechanism.
The PRO-LONG framework enables LLM agents to effectively manage context over long-horizon tasks through structured logging and efficient search.
Visual TL;DR
LLM Agents StruggleDriver
handling long-horizon tasks with sustained perception, reasoning, and exploration
From the article 3 mentionsThis is a critical bottleneck for developing robust LLM agents long horizon tasks.
Context Window BottleneckDriver
difficulty retrieving relevant details from extensive environmental observations
From the articleThis programmatic memory approach circumvents the limitations of traditional context windows by offering a searchable, organized record of past interactions.
PRO-LONG IntroducedCore
From the article 3 mentionsA new minimal context management framework, PRO-LONG, addresses this tradeoff directly.
Programmatic MemoryContext
maintains a complete, structured interaction log for efficient searching
From the articleThis programmatic memory approach circumvents the limitations of traditional context windows by offering a searchable, organized record of past interactions.
SOTA PerformanceOutcome
achieves state-of-the-art results on benchmarks like ARC-AGI-3
From the article 2 mentionsFor instance, with Fable 5, PRO-LONG achieves 97.4% best@2 at a total cost of $1,750, highlighting a significant leap in both performance and cost-effectiveness for LLM agents long horizon tasks.
Leverages Coding AgentsContext
efficiently searches the comprehensive history of past interactions
From the article 2 mentionsIt maintains a complete, structured interaction log and leverages advancements in coding agents to efficiently search this comprehensive history.
Token EfficiencyOutcome
drastic reduction in token usage and associated operational costs
From the articleCrucially, it matches or exceeds state-of-the-art specialized harnesses, reaching up to 76.1% pass@1, all while utilizing 4.2-5.8x fewer tokens.

The persistent challenge of enabling Large Language Model (LLM) agents to handle long-horizon tasks, characterized by sustained perception, reasoning, and exploration, has limited their out-of-the-box performance on benchmarks like ARC-AGI-3. Existing agent harnesses grapple with a fundamental tradeoff: preserving extensive environmental observations for context increases the difficulty of retrieving relevant details. This is a critical bottleneck for developing robust LLM agents long horizon tasks.

Programmatic Memory: Bridging the Context Gap

A new minimal context management framework, PRO-LONG, addresses this tradeoff directly. It maintains a complete, structured interaction log and leverages advancements in coding agents to efficiently search this comprehensive history. This programmatic memory approach circumvents the limitations of traditional context windows by offering a searchable, organized record of past interactions.

State-of-the-Art Performance with Unprecedented Efficiency

On the full ARC-AGI-3 public game set, PRO-LONG demonstrates substantial gains. It improves upon a base coding agent by an average of 18.0 percentage points across frontier models. Crucially, it matches or exceeds state-of-the-art specialized harnesses, reaching up to 76.1% pass@1, all while utilizing 4.2-5.8x fewer tokens. For instance, with Fable 5, PRO-LONG achieves 97.4% best@2 at a total cost of $1,750, highlighting a significant leap in both performance and cost-effectiveness for LLM agents long horizon tasks.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.

More from Daniel Singer