papersTODAY 04:00 UTC
Agentic Visual RAG via Explicit Context Selection and Consolidation
A new arXiv paper proposes a visual retrieval-augmented generation approach that treats evidence gathering as an explicit agentic process. Instead of relying on a single retrieval step, the method selects and consolidates page images as context before reasoning over visually rich documents. The work targets settings where supporting evidence is sparse and spread across pages.