Slide 35
Slide 35 text
参考⽂献
[紹介論文]
Y. Wang et al., "When Token Pruning Is Worse Than Random: Understanding Visual Token Information in VLLMs," Proc. IEEE/CVF Conf. Computer Vision and
Pattern Recognition (CVPR), 2026.
[Bolya+ 2023]
D. Bolya et al., "Token Merging: Your ViT but Faster," Proc. Int'l Conf. Learning Representations (ICLR), 2023.
[Li+ 2023]
J. Li et al., "BLIP-2: Bootstrapping Language-Image Pre-Training with Frozen Image Encoders and Large Language Models," Proc. 40th Int'l Conf. Machine
Learning (ICML), 2023, pp. 19730–19742.
[Chen+ 2024]
L. Chen et al., "An Image Is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models," Proc. European Conf.
Computer Vision (ECCV), 2024, pp. 19–35.
[Liu+ 2024]
H. Liu et al., "Improved Baselines with Visual Instruction Tuning," Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition (CVPR), 2024, pp.
26296–26306.
[Alvar+ 2025]
S.R. Alvar et al., "DivPrune: Diversity-Based Visual Token Pruning for Large Multimodal Models," Proc. IEEE/CVF Conf. Computer Vision and Pattern
Recognition (CVPR), 2025, pp. 9392–9401.
[Bai+ 2025]
S. Bai et al., "Qwen2.5-VL Technical Report," arXiv:2502.13923, 2025.
[Lin+ 2025]
Z. Lin et al., "Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference," Proc. AAAI Conf. Artificial Intelligence (AAAI),
2025, pp. 5334–5342.
[Wen+ 2025]
Z. Wen et al., "Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More," Proc. Conf. Empirical Methods in Natural
Language Processing (EMNLP), 2025.