API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces

S
StartupHub.ai Staff
2 min read
API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces

API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces, an audit by Jennifer Wang, Joachim Baumann, Daniel E. Ho and Sanmi Koyejo, finds that API measurements do not faithfully predict chatbot behavior according to arXiv. The study evaluates ChatGPT, Claude and Gemini across seven systems and nine benchmarks spanning general capability, social bias and sycophancy.

On average, API evaluations scored 3.4 percentage points higher in accuracy and 2.1 percentage points higher in test-retest agreement than corresponding interface evaluations.

For ChatGPT the API versus interface gap exceeded the API-only difference between GPT 5.3 and GPT 5.4. Put differently, switching access surfaces can degrade performance as much as downgrading a full model generation.

The authors varied system prompts, sampling parameters and reasoning settings to see if exposed API controls could reproduce interface behavior, and found they shifted behavior in some cases but did not reliably eliminate the gap.

That documents what the authors call a context-validity gap: measurements obtained through APIs do not necessarily generalize to deployed interfaces, which complicates using API evaluations as proxies for deployed systems. The divergence is familiar to practitioners outside the study, for example OpenAI’s GPT-4o supports up to 128K tokens in the latest API though the consumer ChatGPT interface usually has a smaller limit.

The result does not explain the mechanism, only that visible knobs are insufficient, so a benchmark cited from an API should not be read as a guarantee of chat product performance.

StartupHub data

OpenAI and OpenAI

Artificial intelligence research and deployment company focused on developing advanced AI models like GPT-5.6 and GPT-Live, offering products such as ChatGPT...

Founded
2015
Location
San Francisco, United States
Funding
$122.0B

An AI research and deployment company building safe and beneficial artificial general intelligence.

Founded
2015
Location
San Francisco, United States
Funding
$190.6B

OpenAI is an AI research and deployment company dedicated to ensuring that artificial general intelligence benefits all of humanity.

Founded
2015
Location
San Francisco, United States
Valuation
Private / $100B+ est
© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
S

Written by

StartupHub.ai Staff

Editorial team

The staff writers of StartupHub.ai, ranging from investment analysts to avid AI tool users, early adopters and critical enthusiasts. Backgrounds span engineering, business and the arts. We hold every piece to rigorous standards of research and review.