AI Vision Test Reveals Model Hallucinations

A new benchmark, PerceptionBench, reveals that even advanced multimodal AI models struggle with basic visual perception, often guessing answers instead of truly seeing.

4 min read
Conceptual image representing AI vision and data analysis with abstract visual elements.
PerceptionBench aims to isolate and test atomic visual perception capabilities in multimodal AI models.
Visual TL;DR
AI models struggleDriver
advanced multimodal AI models often guess answers instead of truly seeing
From the article 7 mentionsMoonshot AI has launched PerceptionBench, a new AI benchmark designed to rigorously test the visual perception skills of multimodal large language models (MLLMs).
PerceptionBench launchedCore
From the article 5 mentionsMoonshot AI has launched PerceptionBench, a new AI benchmark designed to rigorously test the visual perception skills of multimodal large language models (MLLMs).
Focus: Atomic CapabilitiesContext
From the articleThis AI benchmark aims to dissect how well these models truly 'see' by focusing on atomic visual capabilities, rather than overall reasoning ability.
Kimi Team developedCore
From the article 2 mentionsDeveloped by the Kimi Team, PerceptionBench stems from analyzing where current frontier models falter.
10 Perceptual CapabilitiesContext
From the article 7 mentionsIt isolates 10 distinct perceptual capabilities, including visual relation, counting, attribute recognition, depth perception, localization, comparison, fine-grained recognition, context integration, OCR, and hallucination detection.
3,000+ Verified QuestionsContext
From the articleThe benchmark comprises over 3,000 verified questions, each crafted to be answerable by simple observation, requiring no complex inference or external knowledge.
Pure Perceptual AccuracyOutcome
From the articleThis approach ensures the evaluation focuses purely on perceptual accuracy.

Moonshot AI has launched PerceptionBench, a new AI benchmark designed to rigorously test the visual perception skills of multimodal large language models (MLLMs). This AI benchmark aims to dissect how well these models truly 'see' by focusing on atomic visual capabilities, rather than overall reasoning ability. You can learn more about the work on kimi.com.

Developed by the Kimi Team, PerceptionBench stems from analyzing where current frontier models falter. It isolates 10 distinct perceptual capabilities, including visual relation, counting, attribute recognition, depth perception, localization, comparison, fine-grained recognition, context integration, OCR, and hallucination detection.

The benchmark comprises over 3,000 verified questions, each crafted to be answerable by simple observation, requiring no complex inference or external knowledge. This approach ensures the evaluation focuses purely on perceptual accuracy.

Atomic Capabilities Under Scrutiny

PerceptionBench categorizes failures into distinct perceptual tasks, such as identifying the number of red dots in an image or determining the length of a line segment. For instance, it asks questions like "What is the length of side AC?" with an answer derived directly from the visual data.

Other examples include counting objects like plates on a table or identifying the number of faces of cubes in contact with the ground. The benchmark also tests OCR capabilities with questions like "What is the blue number?".

Crucially, the benchmark highlights a significant issue: many correct answers fail to be reproduced upon repeated questioning. This suggests that current MLLMs may be guessing answers rather than consistently perceiving visual information accurately.

Model Performance Falls Short

The results from evaluating several models are stark. No model tested managed to exceed 60% accuracy on PerceptionBench. Furthermore, models with similar overall scores often displayed vastly different strengths and weaknesses across the 10 perceptual categories.

The Kimi Team's initiative, which also includes advancements like Kimi K2.6, aims to provide a sharp diagnostic tool. This is in contrast to benchmarks that offer a single score, potentially masking underlying perceptual deficits.

The benchmark's design, driven by observed model failures across over 40 existing benchmarks, ensures it targets genuine weaknesses. This failure-driven taxonomy, combined with rigorous verification, makes PerceptionBench a critical tool for advancing faithful and consistent visual perception in multimodal AI.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.