MEGA Hub

Are LLM-Generated GPU Kernels Production-Ready? A Trace-Driven Benchmark and Optimization Agent

Authors

Do you know Lingyun Yang?You can claim authorship or link another user.Do you know Yuxiao Wang?You can claim authorship or link another user.Do you know Shenghao Liang?You can claim authorship or link another user.Do you know Linfeng Yang?You can claim authorship or link another user.Do you know Daocheng Ying?You can claim authorship or link another user.Do you know Chunbo You?You can claim authorship or link another user.Do you know Rui Zhang?You can claim authorship or link another user.Do you know Luping Wang?You can claim authorship or link another user.Do you know Yinghao Yu?You can claim authorship or link another user.Do you know Guodong Yang?You can claim authorship or link another user.Do you know Liping Zhang?You can claim authorship or link another user.

Abstract

Existing GPU kernel generation benchmarks draw problems from synthetic or curated sources that diverge from deployed workloads. We present Atrex-Bench, a benchmark whose 30 operators and 440 shapes are sampled directly from full-cluster production inference traces of compute-limited, memory-rich GPUs. Each problem carries an importance weight derived from its share of observed GPU time, weighted by application card-hours and computed separately for the serving phases in which it runs, together with a per-problem roofline ceiling, so the aggregate score emphasizes the kernels that consume the most serving time. Evaluating six frontier coding agents on Atrex-Bench shows that even the best vanilla model reaches only ${\sim}10\%$ of the hardware roofline on production operators; and correctness alone overstates capability, since much of the apparent pass rate comes from PyTorch fallbacks rather than kernels the model wrote. To close this gap, we co-release Atrex-Kernel-Agent (AKA), a profile-driven kernel-optimization agent that combines iterative measure-revise search, optimization dropout for escaping stalled search contexts, and a layered GPU-optimization knowledge base (298 reference-kernel files and 244 optimization-knowledge documents, plus external upstream reference projects for API/ISA lookup). In a controlled case study, the agent converts zero-FlyDSL fallbacks into real kernels that match or exceed hand-tuned production baselines.

Community

00

Publication notes

Author note
Both artifacts are released as open source: Atrex-Bench (https://github.com/alibaba/atrex-bench) and Atrex-Kernel-Agent (https://github.com/alibaba/atrex-kernel-agent)