Sign InOpen Brain
arXivPaperNeeds Review

Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation

Controlled text-to-MIDI tests find token representation matters more than a 34× model-size increase for distributional fidelity; performance timing also beats beat-grid tokenization.

arXiv · Aug 4, 2026
Open Source Open MarkdownOpen JSON
Source Summary

With model family, data, budget, and decoding fixed, the study swaps seven music tokenizations. A **0.8B** PMT model records **FMD 159**, versus 272–286 for beat grids, and beats a 27B beat-grid model.

Practical Implication

Builders of music generators should benchmark representations before scaling parameters. PMT retains 10 ms timing, velocity, and multi-track texture; a decode constraint also raises instrument-F1 from **0.28 to 0.60** without measured distributional cost.

Agent-Ready Context
With model family, data, budget, and decoding fixed, the study swaps seven music tokenizations. A **0.8B** PMT model records **FMD 159**, versus 272–286 for beat grids, and beats a 27B beat-grid model.

Builders of music generators should benchmark representations before scaling parameters. PMT retains 10 ms timing, velocity, and multi-track texture; a decode constraint also raises instrument-F1 from **0.28 to 0.60** without measured distributional cost.

FMD measures distributional fidelity, not whether listeners prefer the output; the human study is still pending. Native caption adherence remains weak, and the reported training-distribution imprint suggests conditioning text may influence existing systems less than expected.
Connected Context · Feed7 Judgment

Agogic shifts symbolic-music model selection upstream: representation and decoding constraints can matter more than parameter count under controlled conditions. This reinforces prior evidence that architecture-adjacent data choices can outperform added compute, while narrowing the conclusion to distributional fidelity and instrument coverage because listener preference and strong caption adherence remain unestablished.

Context Map
modelaudio#generative-media#model-selection
Uncertainty
FMD measures distributional fidelity, not whether listeners prefer the output; the human study is still pending. Native caption adherence remains weak, and the reported training-distribution imprint suggests conditioning text may influence existing systems less than expected.