papersTODAY 04:00 UTC
UniCAR-RL targets fine-grained perception failures in multimodal math reasoning
A new arXiv paper introduces UniCAR-RL, a reinforcement learning approach aimed at improving how multimodal large language models handle math problems involving diagrams and figures. The authors argue that weak fine-grained visual perception leads models to hallucinate details early, which then causes errors to compound through the rest of the reasoning chain. The method is framed as improving perception before deeper reasoning steps are attempted.