MEGA Hub

Adaptive Matrix Multiplication for Dynamic Shapes on Ascend NPUs

Authors

Do you know Yuhang Zhou?You can claim authorship or link another user.Do you know Jiang Peng?You can claim authorship or link another user.Do you know Qianyu Jiang?You can claim authorship or link another user.Do you know Zhibin Wang?You can claim authorship or link another user.Do you know Xinghui Tian?You can claim authorship or link another user.Do you know Jianwei Zhou?You can claim authorship or link another user.Do you know Songxiang Zhu?You can claim authorship or link another user.Do you know Jingyi Zhang?You can claim authorship or link another user.Do you know Junsong Wang?You can claim authorship or link another user.Do you know Chen Tian?You can claim authorship or link another user.

Abstract

Matrix Multiplication (MatMul) faces a "generalization crisis" driven by highly dynamic tensor shapes. This crisis is particularly acute on Ascend NPUs, where explicitly controlled architectures and strict physical constraints render existing GPU-centric optimizations ineffective. To resolve this, we propose AdaptCore, an adaptive framework for universally high-performance MatMul on Ascend NPUs. AdaptCore systematically decouples operator optimization into spatial tiling and instruction orchestration. It first maps dynamic shapes into a hardware-aware 2D tiling taxonomy to balance on-chip capacity limits and multi-core parallelism. Furthermore, it integrates a composable optimization library with a deterministic analytical performance model. By mathematically evaluating hardware state mutations, AdaptCore proactively selects and caches optimal implementations, enabling O(1) overhead runtime dispatching. Evaluations demonstrate that AdaptCore delivers a remarkable 1.85x mean speedup across 80,000 input shapes, and achieves up to a 1.48x acceleration in representative end-to-end models over the highly-tuned native vendor library (ACLNN).

Community

00