MEGA Hub

Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation

Authors

Do you know Junhao Chen?You can claim authorship or link another user.Do you know Mingjin Chen?You can claim authorship or link another user.Do you know Jingjia Mao?You can claim authorship or link another user.Do you know Lin Chen?You can claim authorship or link another user.Do you know Saining Zhang?You can claim authorship or link another user.Do you know Minglin Chen?You can claim authorship or link another user.Do you know Ruocheng Wu?You can claim authorship or link another user.Do you know Liaoyuan Fan?You can claim authorship or link another user.Do you know Wenyi Li?You can claim authorship or link another user.Do you know Mingju Gao?You can claim authorship or link another user.Do you know Henghaofan Zhang?You can claim authorship or link another user.Do you know Zhihao Li?You can claim authorship or link another user.Do you know Hao Zhao?You can claim authorship or link another user.Do you know Yufei Wang?You can claim authorship or link another user.Do you know Ruqi Huang?You can claim authorship or link another user.

Abstract

Text-to-music language models begin with a choice usually made by default: how to tokenize music. Normally entangled with backbone, data, and recipe, its effect has never been measured in isolation. We fix pretrained Qwen3.5 (0.8B-27B), data, budget, and decoding, and swap only the representation across seven tokenizations, anchoring texture metrics to each representation's model-free ceiling. The ordering is clean and surprising: representation, not model size, is the binding variable for distributional fidelity. Scaling the backbone 34x barely moves Frechet Music Distance (FMD), whereas switching representation halves it. PMT, a performance-resolution stream we release (10 ms timing, per-note velocity, multi-track texture; 609 symbols), reaches FMD 159 at 0.8B against 272-286 for beat grids (1.7-1.8x lower, up to 2.8x elsewhere; non-overlapping bootstrap CIs), so a 0.8B performance-resolution model beats a 27B beat grid. It reappears on a 26M from-scratch backbone and a second performance-resolution tokenizer: a property of the class, not one lucky vocabulary. Nor is it a finer-lattice artifact: snapping PMT's onsets to the beat grids' resolution still leaves it 67-129 FMD ahead of both (n=500). The effect is distributional; whether it is audible is a separate question, left open by our probe, with a human study pre-registered. Native caption adherence is weak but separable: a lightweight decode-time constraint doubles instrument-F1 (.28 to .60) and Correct-Key (.16 to .35) at no distributional cost. We release the harness, 25+ checkpoints, two corpora (86.6k aligned across caption/MIDI/ABC/audio; 6.25M captioned, the largest for music), and an imprinting diagnostic: published text-to-MIDI systems reproduce their training distribution near-invariant to the caption (72% vs. 71% chord-time on disjoint domains). The field's next representation claim can now be measured, not asserted.

Community

00

Publication notes

Author note
Project Page: https://yisuanwang.github.io/Agogic