MEGA Hub

OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs

Authors

Do you know Haoyang Huang?You can claim authorship or link another user.Do you know Wenjie Huang?You can claim authorship or link another user.Do you know Tianqi Xu?You can claim authorship or link another user.Do you know Hongyaoxing Gu?You can claim authorship or link another user.Do you know Kang Tan?You can claim authorship or link another user.Do you know Yikai Fu?You can claim authorship or link another user.Do you know Yuhao Shen?You can claim authorship or link another user.Do you know Tianyu Liu?You can claim authorship or link another user.Do you know Baolin Zhang?You can claim authorship or link another user.Do you know Jun Zhang?You can claim authorship or link another user.Do you know Xinyi Hu?You can claim authorship or link another user.Do you know Jun Dai?You can claim authorship or link another user.Do you know Shuang Ge?You can claim authorship or link another user.Do you know Lei Chen?You can claim authorship or link another user.Do you know Yue Li?You can claim authorship or link another user.Do you know Mingchen Wang?You can claim authorship or link another user.Do you know Meng Zhang?You can claim authorship or link another user.

Abstract

Emerging Omni-modal Large Language Models (OmniLLMs) enable unified understanding of text, audio, and video, but their long audio-video token sequences introduce substantial memory and inference costs. Existing compression methods mainly focus on selecting important tokens under fixed budgets, leaving the preceding budget-allocation problem underexplored. We show that direct query-to-audio/video similarity is unreliable for inter-modal budget allocation, and that uniform intra-modal budgets can miss key evidence while retaining redundant content. To address these limitations, we propose OmniDelta, a training-free, skill-driven framework that couples intent-aware inter-modal allocation with content-aware intra-modal allocation. OmniDelta first constructs audio and video skill pools to shift the fixed retained-token budget according to query demand, then reallocates modality budgets over audio segments and video frames using local complexity and temporal redundancy. The resulting local budgets can be combined with existing pruning strategies, preserving the total retained-token ratio while changing where the budget is spent. Experiments on four audio-video benchmarks with two Qwen2.5-Omni models show that OmniDelta establishes a new accuracy-efficiency Pareto frontier across pruning ratios. At 25% token retention on Qwen2.5-Omni-7B, OmniDelta reduces GPU memory by 22.0% and achieves a 1.64x end-to-end speedup over full-token inference.

Community

00

Publication notes

Author note
24 pages, 8 figures