Pruning Yicheng Ji 1,2 , Jun Zhang 1,2 , Heming Xia 3 , Jinpeng Chen 4 , Lidan Shou 1,2 , Gang Chen 1 , Huan Li 1,2 1 The State Key Laboratory of Blockchain and Data Security, Zhejiang University 2 Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security 3 Department of Computing, The Hong Kong Polytechnic University 4 School of Computer Science, Beijing University of Posts and Telecommunications EMNLP25 慶應義塾⼤学 杉浦孔明研究室 B4 野⼝拓海 Yic he ng J i e t a l., “S pe c V L M : Enha nc ing S pe c ula tive De c o ding o f V ide o L L M s via V e r ifie r- G uide d T o k e n P r uning,” in EM N L P , 2025 -1-
(→ appendix) • ⼀度の AR ステップで複数トークンを⽣成し⾼速化 • Draft model (軽量) : 将来のトークン列 (ドラフト) を⽣成 • Target model (元のLLM) : 尤度に基づきドラフトを検証 An Introduction to Speculative Decoding for Reducing Latency in AI Inference, NVIDIA J Lossless 理論保証: 出⼒分布は元のLLMのものと⼀致 o Video LLM では draft model が低速化 • Draft model も⼤量の video token を保持 L KV cache の読み書きがボトルネック -4-