Visual TL;DR. OpenAI models struggle due to Generic harness design. Generic harness design caused Memory is Key. Memory is Key addressed by Tweaked API settings. Tweaked API settings led to Tripled benchmark scores. Tripled benchmark scores demonstrates Harness design impacts. Generic harness design related to Compaction Cuts Waste. Harness design impacts illustrates Benchmark Realities.
- OpenAI models struggle: GPT-5.6 Sol scored 7.8%, GPT-5.5 only 0.4% on ARC-AGI-3 benchmark
- Generic harness design: benchmark's simple harness discarded private reasoning after each action
- Memory is Key: AI forced to restart problem-solving from scratch every turn
- Tweaked API settings: researchers adjusted two simple API settings for model interaction
- Tripled benchmark scores: models achieved significantly higher scores on the ARC-AGI-3 benchmark
- Harness design impacts: proving how benchmark harness design significantly affects AI evaluation
- Compaction Cuts Waste: optimized memory usage by compacting redundant information
- Benchmark Realities: highlights the difference between academic and commercial benchmark optimization
Visual TL;DR
