API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces

API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces

API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces, an audit by Jennifer Wang, Joachim Baumann, Daniel E. Ho and Sanmi Koyejo, finds that API measurements do not faithfully predict chatbot behavior according to arXiv. The study evaluates ChatGPT, Claude and Gemini across seven systems and nine benchmarks spanning general capability, social bias and sycophancy.

On average, API evaluations scored 3.4 percentage points higher in accuracy and 2.1 percentage points higher in test-retest agreement than corresponding interface evaluations.

For ChatGPT the API versus interface gap exceeded the API-only difference between GPT 5.3 and GPT 5.4. Put differently, switching access surfaces can degrade performance as much as downgrading a full model generation.

The authors varied system prompts, sampling parameters and reasoning settings to see if exposed API controls could reproduce interface behavior, and found they shifted behavior in some cases but did not reliably eliminate the gap.

That documents what the authors call a context-validity gap: measurements obtained through APIs do not necessarily generalize to deployed interfaces, which complicates using API evaluations as proxies for deployed systems. The divergence is familiar to practitioners outside the study, for example OpenAI’s GPT-4o supports up to 128K tokens in the latest API though the consumer ChatGPT interface usually has a smaller limit.

The result does not explain the mechanism, only that visible knobs are insufficient, so a benchmark cited from an API should not be read as a guarantee of chat product performance.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.