Sample-Adaptive Strategy Routing Improves Vision Token Pruning in Multimodal LLMs
A new arXiv preprint proposes choosing among vision token pruning strategies on a per-image basis rather than applying one fixed policy to every input. By adapting the pruning approach to each sample, the method aims to lower the heavy inference costs that multimodal language models incur from processing large numbers of visual tokens while maintaining output quality.