MEGA Hub

LongCat Sparse Attention: Taming the Lightning via Streaming-aware Hierarchical Cross-Layer Indexing

Authors

Do you know Wen Zan?You can claim authorship or link another user.Do you know Jiaqi Zhang?You can claim authorship or link another user.Do you know Jianchao Tan?You can claim authorship or link another user.Do you know Hong Liu?You can claim authorship or link another user.Do you know Cunguang Wang?You can claim authorship or link another user.Do you know Xiang Li?You can claim authorship or link another user.Do you know Duyue Ma?You can claim authorship or link another user.Do you know Guanyu Wu?You can claim authorship or link another user.Do you know Yifan Lu?You can claim authorship or link another user.Do you know Fengcun Li?You can claim authorship or link another user.Do you know Yerui Sun?You can claim authorship or link another user.Do you know Peng Pei?You can claim authorship or link another user.Do you know Yuchen Xie?You can claim authorship or link another user.Do you know Xunliang Cai?You can claim authorship or link another user.

Abstract

DeepSeek Sparse Attention (DSA) enables efficient long-context modeling through its Lightning Indexer. However, practical deployment remains constrained by the indexer's expensive $O(L^2)$ scoring overhead and the hardware-inefficient, discontinuous memory-access patterns induced by its outputs. To address these system-level bottlenecks, we introduce LongCat Sparse Attention (LSA), a hardware-algorithm co-designed framework comprising three complementary and orthogonal strategies: (1) Streaming-Aware Indexing, which selectively converts scattered KV entries into hardware-aligned contiguous layouts to enable coalesced HBM access; (2) Cross-Layer Indexing, which amortizes indexing overhead by reusing the results produced by a single layer across consecutive layers, supported by cross-layer distillation; and (3) Hierarchical Indexing, which adopts a coarse-to-fine scoring scheme to progressively narrow the candidate set for each query, thereby substantially reducing indexing computation. Extensive scaling experiments, ranging from 69B-A3B to 560B-A27B models, demonstrate that LSA consistently achieves performance on par with full attention across both general-purpose and long-context benchmarks. Moreover, LSA supports native training with context lengths of up to one million tokens and underpins the development of LongCat-2.0 (1.6T-A48B). To facilitate further research, we also introduce and open-source LongCat-Flash-Lite-Sparse (69B-A3B), which integrates LSA into LongCat-Flash-Lite and incorporates an updated long-context training corpus.

Community

00