# Beyond RGB: Grounding Vision-Language on Raw Sensor Data _PRISM-VL advances vision-language models by grounding them in raw camera measurements, not just RGB, significantly improving performance on challenging visual tasks._ **Updated:** 2026-08-22 **Published:** 2026-05-13 **Source:** https://www.startuphub.ai/ai-news/ai-research/2026/beyond-rgb-grounding-vision-language-on-raw-sensor-data --- Vision-language models (VLMs) typically operate on post-image signal processing (ISP) RGB images. This preprocessing pipeline often discards crucial sensor evidence through clipping, suppression, or quantization, thereby limiting the model's ability to accurately ground its understanding. A new approach, [PRISM-VL](https://arxiv.org/abs/2605.11727v1), investigates whether grounding performance improves when the visual interface is moved closer to the original camera measurement. VLMs use RGBDriver From the articleVision-language models (VLMs) typically operate on post-image signal processing (ISP) RGB images.RGB loses dataDriverpreprocessing discards crucial sensor evidence through clipping, suppression, or quantizationFrom the articleThis technique effectively transfers supervision signals from readily available RGB proxies to the more granular, raw measurement-domain observations, addressing a fundamental challenge in training on sensor data.problemPRISM-VL approachCoregrounds vision-language models in raw camera measurements, not just RGBFrom the article 2 mentionsA new approach, PRISM-VL, investigates whether grounding performance improves when the visual interface is moved closer to the original camera measurement.enablesRAW-derived Meas.-XYZContextdirectly incorporates raw sensor data inputs for improved groundingFrom the articleThis framework directly incorporates RAW-derived Meas.-XYZ inputs.Camera-conditioned groundingCorea key innovation for better understanding of sensor dataFrom the article 3 mentionsA key innovation is its camera-conditioned grounding mechanism and Exposure-Bracketed Supervision Aggregation.Improved performanceOutcomesignificantly improves performance on challenging visual tasksFrom the articleA new approach, PRISM-VL, investigates whether grounding performance improves when the visual interface is moved closer to the original camera measurement.Exposure-Bracketed SupervisionCoretransfers supervision from RGB proxies to raw measurement domain observationsFrom the article 2 mentionsA key innovation is its camera-conditioned grounding mechanism and Exposure-Bracketed Supervision Aggregation.Quantifiable gainsOutcomedemonstrates measurable improvements in challenging scenariosFrom the articleThis represents a substantial leap over the RGB-based Qwen3-VL-8B baseline, with gains of +0.1074 BLEU, +0.1071 ROUGE-L, and +4.46 percentage points in LLM-Judge accuracy. ## Bridging the Measurement-to-RGB Gap The researchers introduce measurement-grounded vision-language learning, instantiated as PRISM-VL. This framework directly incorporates RAW-derived Meas.-XYZ inputs. A key innovation is its camera-conditioned grounding mechanism and Exposure-Bracketed Supervision Aggregation. This technique effectively transfers supervision signals from readily available RGB proxies to the more granular, raw measurement-domain observations, addressing a fundamental challenge in training on sensor data. ## Quantifiable Gains in Challenging Scenarios PRISM-VL-8B, trained on a 150K instruction-tuning set and evaluated on a benchmark targeting low-light, HDR, visibility-sensitive, and hallucination-sensitive cases, achieved significant improvements. It reached 0.6120 BLEU and 0.4571 ROUGE-L scores, alongside an 82.66% LLM-Judge accuracy. This represents a substantial leap over the RGB-based Qwen3-VL-8B baseline, with gains of +0.1074 BLEU, +0.1071 ROUGE-L, and +4.46 percentage points in LLM-Judge accuracy. These results strongly suggest that a portion of VLM grounding errors stems directly from information lost during standard RGB rendering, underscoring the value of preserving measurement-domain evidence for enhanced multimodal reasoning. --- Original analysis from [startuphub.ai](https://www.startuphub.ai), the #1 AI startup directory.