OpenAI Triples Benchmark Scores With Simple Settings
OpenAI found simple API setting tweaks tripled scores on the ARC-AGI-3 benchmark, proving harness design significantly impacts AI evaluation.

Visual TL;DR
GPT-5.6 Sol scored 7.8%, GPT-5.5 only 0.4% on ARC-AGI-3 benchmark
From the article 2 mentionsOpenAI's advanced AI models struggled on the ARC-AGI-3 benchmark until researchers tweaked two API settings, according to a new OpenAI News publication.
benchmark's simple harness discarded private reasoning after each action
From the article 2 mentionsThe issue wasn't inherent model weakness, but rather the benchmark's generic harness design.
AI forced to restart problem-solving from scratch every turn
optimized memory usage by compacting redundant information
researchers adjusted two simple API settings for model interaction
From the article 2 mentionsOpenAI's advanced AI models struggled on the ARC-AGI-3 benchmark until researchers tweaked two API settings, according to a new OpenAI News publication.
models achieved significantly higher scores on the ARC-AGI-3 benchmark
From the article 2 mentionsThe combination of retained reasoning and compaction tripled scores on the public task set, boosting GPT-5.6 Sol's performance from 13.3% to 38.3% against the official harness.
proving how benchmark harness design significantly affects AI evaluation
From the article 2 mentionsThey are also influenced by API settings and harness design.
highlights the difference between academic and commercial benchmark optimization
From the article 9 mentionsThis performance on the 2D puzzle game benchmark was perplexing, given the models' success on complex math problems and other games like Pokémon FireRed.
Contents(3)
© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Written by
Daniel SingerEditor, StartupHub.ai
Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.
More from Daniel Singer