UniDoc-RL: Finer-Grained Visual RAG

UniDoc-RL enhances LVLMs with fine-grained visual RAG via hierarchical RL, active perception, and multi-reward training, achieving state-of-the-art results.

Diagram illustrating the UniDoc-RL framework showing hierarchical action space and information refinement.
The UniDoc-RL framework proposes a hierarchical approach to visual RAG.
Contents(3)

The efficacy of Large Vision-Language Models (LVLMs) is often capped by their reliance on generic retrieval signals for external visual knowledge. This approach fails to capture the fine-grained semantics critical for sophisticated reasoning tasks. To bridge this gap, researchers have introduced UniDoc-RL, a unified reinforcement learning framework designed to tackle these limitations head-on.

Active Perception for Semantic Precision

UniDoc-RL reframes visual information acquisition as a sequential decision-making process. Its hierarchical action space allows the LVLM agent to move beyond simple document retrieval. It progressively refines visual evidence, starting with coarse-grained document retrieval and advancing to fine-grained image selection and active region cropping. This granular control enables the model to actively suppress irrelevant content and focus on information-dense areas, a crucial step for accurate reasoning.

End-to-End Training with Dense Multi-Reward Supervision

Training such a complex agent necessitates sophisticated supervision. UniDoc-RL employs a dense multi-reward scheme that provides task-aware feedback for each action taken by the agent. Coupled with Group Relative Policy Optimization (GRPO), this approach allows for effective end-to-end training of the LVLM agent, aligning its behavior with multiple objectives without the need for a separate value network. The development of a comprehensive dataset with fine-grained action annotations further supports this advanced training paradigm, pushing the boundaries of what's possible with UniDoc-RL visual RAG.

Outperforming State-of-the-Art in Visual Reasoning

Experimental results across three benchmarks demonstrate the superiority of UniDoc-RL. The framework consistently surpasses existing state-of-the-art baselines, achieving gains of up to 17.7% over prior RL-based methods. This significant performance improvement underscores the value of UniDoc-RL's approach to fine-grained visual retrieval and active perception in complex reasoning scenarios, positioning UniDoc-RL visual RAG as a leading solution.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.