Beyond RGB: Grounding Vision-Language on Raw Sensor Data

PRISM-VL advances vision-language models by grounding them in raw camera measurements, not just RGB, significantly improving performance on challenging visual tasks.

Diagram illustrating PRISM-VL framework with RAW sensor data input and RGB proxy.
PRISM-VL architecture showing the integration of raw sensor measurements.
Visual TL;DR
VLMs use RGBDriver
From the articleVision-language models (VLMs) typically operate on post-image signal processing (ISP) RGB images.
RGB loses dataDriver
preprocessing discards crucial sensor evidence through clipping, suppression, or quantization
From the articleThis technique effectively transfers supervision signals from readily available RGB proxies to the more granular, raw measurement-domain observations, addressing a fundamental challenge in training on sensor data.
PRISM-VL approachCore
grounds vision-language models in raw camera measurements, not just RGB
From the article 2 mentionsA new approach, PRISM-VL, investigates whether grounding performance improves when the visual interface is moved closer to the original camera measurement.
RAW-derived Meas.-XYZContext
directly incorporates raw sensor data inputs for improved grounding
From the articleThis framework directly incorporates RAW-derived Meas.-XYZ inputs.
Camera-conditioned groundingCore
a key innovation for better understanding of sensor data
From the article 3 mentionsA key innovation is its camera-conditioned grounding mechanism and Exposure-Bracketed Supervision Aggregation.
Improved performanceOutcome
significantly improves performance on challenging visual tasks
From the articleA new approach, PRISM-VL, investigates whether grounding performance improves when the visual interface is moved closer to the original camera measurement.
Exposure-Bracketed SupervisionCore
transfers supervision from RGB proxies to raw measurement domain observations
From the article 2 mentionsA key innovation is its camera-conditioned grounding mechanism and Exposure-Bracketed Supervision Aggregation.
Quantifiable gainsOutcome
demonstrates measurable improvements in challenging scenarios
From the articleThis represents a substantial leap over the RGB-based Qwen3-VL-8B baseline, with gains of +0.1074 BLEU, +0.1071 ROUGE-L, and +4.46 percentage points in LLM-Judge accuracy.

Vision-language models (VLMs) typically operate on post-image signal processing (ISP) RGB images. This preprocessing pipeline often discards crucial sensor evidence through clipping, suppression, or quantization, thereby limiting the model's ability to accurately ground its understanding. A new approach, PRISM-VL, investigates whether grounding performance improves when the visual interface is moved closer to the original camera measurement.

Bridging the Measurement-to-RGB Gap

The researchers introduce measurement-grounded vision-language learning, instantiated as PRISM-VL. This framework directly incorporates RAW-derived Meas.-XYZ inputs. A key innovation is its camera-conditioned grounding mechanism and Exposure-Bracketed Supervision Aggregation. This technique effectively transfers supervision signals from readily available RGB proxies to the more granular, raw measurement-domain observations, addressing a fundamental challenge in training on sensor data.

Quantifiable Gains in Challenging Scenarios

PRISM-VL-8B, trained on a 150K instruction-tuning set and evaluated on a benchmark targeting low-light, HDR, visibility-sensitive, and hallucination-sensitive cases, achieved significant improvements. It reached 0.6120 BLEU and 0.4571 ROUGE-L scores, alongside an 82.66% LLM-Judge accuracy. This represents a substantial leap over the RGB-based Qwen3-VL-8B baseline, with gains of +0.1074 BLEU, +0.1071 ROUGE-L, and +4.46 percentage points in LLM-Judge accuracy. These results strongly suggest that a portion of VLM grounding errors stems directly from information lost during standard RGB rendering, underscoring the value of preserving measurement-domain evidence for enhanced multimodal reasoning.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.