papersSEP 10 04:00 UTC
Predicting Middle-Layer Attention in Multimodal LLMs for Efficient Visual Token Pruning
Multimodal large language models spend significant compute processing large numbers of visual tokens, and effective pruning depends on knowing which tokens actually matter. This paper introduces a learned approach that predicts attention at middle layers, enabling models to identify and drop less relevant visual tokens. The method aims to cut inference costs while maintaining performance across vision-language tasks.