MEGA Hub

FSGen: Agile Fused and Sparse Accelerator Generator with Accurate Power Model for LLM Applications

Authors

Do you know Jay Zhe-An Mok?You can claim authorship or link another user.Do you know Qijun Zhang?You can claim authorship or link another user.Do you know Zhiyao Xie?You can claim authorship or link another user.

Abstract

With the growing demand of artificial intelligence (AI) applications, large language models (LLMs) have become important workloads in many domains. The question of how to efficiently generate optimal AI chip accelerator designs remains unresolved and challenging. Currently, there is a lack of end-to-end design methodologies for efficient design space exploration (DSE). We propose FSGen, an agile framework for attention-based LLM accelerator generation with an early-stage PPA estimator. FSGen supports fused operator dataflows and sparsity with a diverse design space and finds designs with 1.4x better power efficiency or 10x speedup with similar PPA metrics compared to prior work. Pareto-optimal designs have much better performance over a wide range of LLM benchmarks and have 58x better figures of merit (FoM). Design exploration is also faster due to our PPA estimators, which have better accuracy than prior art and reduce DSE runtime drastically.

Community

00

Publication notes

Author note
Research Manuscript Published in Design Automation Conference (DAC) 63 (2026)