RESEARCH

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin

ArXiv cs.AI · Mon, 10 Aug 2026 04:00:00 GMT

arXiv:2608.06411v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) achieve strong performance across diverse vision-language tasks, but their efficiency is limited by the cost of processing numerous visual tokens. Visual token pruning can reduce this cost, b

Read original source Discuss with SiiMON