PRO-LONG: Memory for LLM Agents

PRO-LONG revolutionizes LLM agent capabilities for long-horizon tasks, achieving SOTA performance with drastic token efficiency and cost reduction.

5 min read
Diagram illustrating the PRO-LONG framework for LLM agents with a structured interaction log and search mechanism.
The PRO-LONG framework enables LLM agents to effectively manage context over long-horizon tasks through structured logging and efficient search.

Visual TL;DR. LLM Agents Struggle due to Context Window Bottleneck. Context Window Bottleneck solves PRO-LONG Introduced. PRO-LONG Introduced uses Programmatic Memory. Programmatic Memory enabled by Leverages Coding Agents. PRO-LONG Introduced leads to SOTA Performance. SOTA Performance with Token Efficiency.

  1. LLM Agents Struggle: handling long-horizon tasks with sustained perception, reasoning, and exploration
  2. Context Window Bottleneck: difficulty retrieving relevant details from extensive environmental observations
  3. PRO-LONG Introduced: a new minimal context management framework addressing the context tradeoff
  4. Programmatic Memory: maintains a complete, structured interaction log for efficient searching
  5. Leverages Coding Agents: efficiently searches the comprehensive history of past interactions
  6. SOTA Performance: achieves state-of-the-art results on benchmarks like ARC-AGI-3
  7. Token Efficiency: drastic reduction in token usage and associated operational costs
Visual TL;DR
Visual TL;DR, startuphub.ai PRO-LONG Introduced uses Programmatic Memory. PRO-LONG Introduced leads to SOTA Performance uses leads to LLM Agents Struggle PRO-LONG Introduced Programmatic Memory SOTA Performance From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai PRO-LONG Introduced uses Programmatic Memory. PRO-LONG Introduced leads to SOTA Performance uses leads to LLM AgentsStruggle PRO-LONGIntroduced ProgrammaticMemory SOTA Performance From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai PRO-LONG Introduced uses Programmatic Memory. PRO-LONG Introduced leads to SOTA Performance uses leads to LLM Agents Struggle handling long-horizon tasks with sustainedperception, reasoning, and exploration PRO-LONG Introduced a new minimal context management frameworkaddressing the context tradeoff Programmatic Memory maintains a complete, structuredinteraction log for efficient searching SOTA Performance achieves state-of-the-art results onbenchmarks like ARC-AGI-3 From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai PRO-LONG Introduced uses Programmatic Memory. PRO-LONG Introduced leads to SOTA Performance uses leads to LLM AgentsStruggle handlinglong-horizon taskswith sustained… PRO-LONGIntroduced a new minimalcontext managementframework… ProgrammaticMemory maintains acomplete,structured… SOTA Performance achievesstate-of-the-artresults on… From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai LLM Agents Struggle due to Context Window Bottleneck. Context Window Bottleneck solves PRO-LONG Introduced. PRO-LONG Introduced uses Programmatic Memory. Programmatic Memory enabled by Leverages Coding Agents. PRO-LONG Introduced leads to SOTA Performance. SOTA Performance with Token Efficiency due to solves uses enabled by leads to with LLM Agents Struggle handling long-horizon tasks with sustainedperception, reasoning, and exploration Context Window Bottleneck difficulty retrieving relevant detailsfrom extensive environmental observations PRO-LONG Introduced a new minimal context management frameworkaddressing the context tradeoff Programmatic Memory maintains a complete, structuredinteraction log for efficient searching Leverages Coding Agents efficiently searches the comprehensivehistory of past interactions SOTA Performance achieves state-of-the-art results onbenchmarks like ARC-AGI-3 Token Efficiency drastic reduction in token usage andassociated operational costs From startuphub.ai · The publishers behind this format
Visual TL;DR, startuphub.ai LLM Agents Struggle due to Context Window Bottleneck. Context Window Bottleneck solves PRO-LONG Introduced. PRO-LONG Introduced uses Programmatic Memory. Programmatic Memory enabled by Leverages Coding Agents. PRO-LONG Introduced leads to SOTA Performance. SOTA Performance with Token Efficiency due to solves uses enabled by leads to with LLM AgentsStruggle handlinglong-horizon taskswith sustained… Context WindowBottleneck difficultyretrieving relevantdetails from… PRO-LONGIntroduced a new minimalcontext managementframework… ProgrammaticMemory maintains acomplete,structured… Leverages CodingAgents efficientlysearches thecomprehensive… SOTA Performance achievesstate-of-the-artresults on… Token Efficiency drastic reductionin token usage andassociated… From startuphub.ai · The publishers behind this format

The persistent challenge of enabling Large Language Model (LLM) agents to handle long-horizon tasks, characterized by sustained perception, reasoning, and exploration, has limited their out-of-the-box performance on benchmarks like ARC-AGI-3. Existing agent harnesses grapple with a fundamental tradeoff: preserving extensive environmental observations for context increases the difficulty of retrieving relevant details. This is a critical bottleneck for developing robust LLM agents long horizon tasks.

Programmatic Memory: Bridging the Context Gap

A new minimal context management framework, PRO-LONG, addresses this tradeoff directly. It maintains a complete, structured interaction log and leverages advancements in coding agents to efficiently search this comprehensive history. This programmatic memory approach circumvents the limitations of traditional context windows by offering a searchable, organized record of past interactions.

State-of-the-Art Performance with Unprecedented Efficiency

On the full ARC-AGI-3 public game set, PRO-LONG demonstrates substantial gains. It improves upon a base coding agent by an average of 18.0 percentage points across frontier models. Crucially, it matches or exceeds state-of-the-art specialized harnesses, reaching up to 76.1% pass@1, all while utilizing 4.2-5.8x fewer tokens. For instance, with Fable 5, PRO-LONG achieves 97.4% best@2 at a total cost of $1,750, highlighting a significant leap in both performance and cost-effectiveness for LLM agents long horizon tasks.

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.