MEGA Hub

Beyond Prefill-Decode Disaggregation: Dissecting LLM Inference for Heterogeneous Platforms via Dynamic Operator Scheduling

Authors

Do you know Jiaqi Yang?You can claim authorship or link another user.Do you know Jiayi Li?You can claim authorship or link another user.Do you know Yihan Fu?You can claim authorship or link another user.Do you know Hongxiao Zhao?You can claim authorship or link another user.Do you know Zhan Chen?You can claim authorship or link another user.Do you know Qiuping Wu?You can claim authorship or link another user.Do you know Yuchao Yang?You can claim authorship or link another user.Do you know Bonan Yan?You can claim authorship or link another user.

Abstract

Prefill-decode disaggregation (PD) and roofline-based operator placement are common strategies for partitioning Large Language Model (LLM) inference across heterogeneous systems, but they are often insufficient in practice. End-to-end latency also depends on workload shape, runtime device contention, and persistent weight layout. We present DOPS (dynamic operator scheduling), a hardware-aware, closed-loop framework that jointly optimizes operator scheduling and blockwise weight layouts. DOPS constructs a stage-aware directed acyclic graph (DAG) and integrates two components: the Bifocal scheduler for dynamic operator-to-device placement and the Weight Layout Arbiter (WLA) for selecting hardware-efficient weight layouts under strict memory constraints. Across representative heterogeneous systems combining neural processing units (NPUs) and processing-in-memory (PIM) devices, Bifocal achieves geometric-mean speedups of 1.20$\times$ to 2.23$\times$ over the PD baseline. WLA provides an additional geometric-mean speedup of 1.28$\times$ to 1.33$\times$ over Bifocal/Linear. DOPS also supports systematic analysis of workload sensitivity and hardware scalability for LLM serving. The source code is available at https://github.com/YIAI-02/TriForm, and the visualization tool is demonstrated at https://youtu.be/Ya_oMCyYno0.

Community

00

Publication notes

Author note
To appear in MICRO 2026