# Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation

Source: [arXiv](https://arxiv.org/abs/2608.03999v1)  
Feed7 permalink: https://feed7.dev/p/2608-03999v1-1584qkh  
Published: 2026-08-04T17:56:49.000Z  
Trust: Needs Review (needs_review)

## Why Included

Controlled text-to-MIDI tests find token representation matters more than a 34× model-size increase for distributional fidelity; performance timing also beats beat-grid tokenization.

## Source Summary

With model family, data, budget, and decoding fixed, the study swaps seven music tokenizations. A **0.8B** PMT model records **FMD 159**, versus 272–286 for beat grids, and beats a 27B beat-grid model.

## Practical Implication

Builders of music generators should benchmark representations before scaling parameters. PMT retains 10 ms timing, velocity, and multi-track texture; a decode constraint also raises instrument-F1 from **0.28 to 0.60** without measured distributional cost.

## Agent-Ready Context

With model family, data, budget, and decoding fixed, the study swaps seven music tokenizations. A **0.8B** PMT model records **FMD 159**, versus 272–286 for beat grids, and beats a 27B beat-grid model.

Builders of music generators should benchmark representations before scaling parameters. PMT retains 10 ms timing, velocity, and multi-track texture; a decode constraint also raises instrument-F1 from **0.28 to 0.60** without measured distributional cost.

FMD measures distributional fidelity, not whether listeners prefer the output; the human study is still pending. Native caption adherence remains weak, and the reported training-distribution imprint suggests conditioning text may influence existing systems less than expected.

## Connected Context

Feed7 judgment across 353 accumulated Signals:

Agogic shifts symbolic-music model selection upstream: representation and decoding constraints can matter more than parameter count under controlled conditions. This reinforces prior evidence that architecture-adjacent data choices can outperform added compute, while narrowing the conclusion to distributional fidelity and instrument coverage because listener preference and strong caption adherence remain unestablished.

- [Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI](https://feed7.dev/p/data-quality-is-the-compute-multiplier-ari-morcos-datologyai-0x7k2ve) — Both show that upstream choices about how training material is represented or composed can outperform simply scaling compute; Agogic provides a controlled tokenization comparison specific to symbolic music.
- [Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language Models](https://feed7.dev/p/2607-11801v1-0m0m7wj) — IAAN changes acoustic behavior through a targeted inference-time intervention, while Agogic’s decode constraint similarly shows that measured audio capability can improve without enlarging or retraining the core model.

## Context Map

- Layer: model
- Domains: audio
- Topics: generative-media, model-selection

## Uncertainty

- FMD measures distributional fidelity, not whether listeners prefer the output; the human study is still pending. Native caption adherence remains weak, and the reported training-distribution imprint suggests conditioning text may influence existing systems less than expected.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
