{
  "schema_version": "1.1",
  "id": "s13:https://arxiv.org/abs/2608.03999v1",
  "slug": "2608-03999v1-1584qkh",
  "url": "https://feed7.dev/p/2608-03999v1-1584qkh",
  "title": "Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation",
  "why_included": "Controlled text-to-MIDI tests find token representation matters more than a 34× model-size increase for distributional fidelity; performance timing also beats beat-grid tokenization.",
  "summary": "With model family, data, budget, and decoding fixed, the study swaps seven music tokenizations. A **0.8B** PMT model records **FMD 159**, versus 272–286 for beat grids, and beats a 27B beat-grid model.",
  "practical_implication": "Builders of music generators should benchmark representations before scaling parameters. PMT retains 10 ms timing, velocity, and multi-track texture; a decode constraint also raises instrument-F1 from **0.28 to 0.60** without measured distributional cost.",
  "agent_context": "With model family, data, budget, and decoding fixed, the study swaps seven music tokenizations. A **0.8B** PMT model records **FMD 159**, versus 272–286 for beat grids, and beats a 27B beat-grid model.\n\nBuilders of music generators should benchmark representations before scaling parameters. PMT retains 10 ms timing, velocity, and multi-track texture; a decode constraint also raises instrument-F1 from **0.28 to 0.60** without measured distributional cost.\n\nFMD measures distributional fidelity, not whether listeners prefer the output; the human study is still pending. Native caption adherence remains weak, and the reported training-distribution imprint suggests conditioning text may influence existing systems less than expected.",
  "source": {
    "name": "arXiv",
    "url": "https://arxiv.org/abs/2608.03999v1",
    "published_at": "2026-08-04T17:56:49.000Z"
  },
  "source_class": "blog_post",
  "content_type": "Paper",
  "layer": "model",
  "domains": [
    "audio"
  ],
  "topics": [
    "generative-media",
    "model-selection"
  ],
  "verification": {
    "status": "needs_review",
    "label": "Needs Review",
    "method": "unverified",
    "verified_at": null
  },
  "uncertainty": [
    "FMD measures distributional fidelity, not whether listeners prefer the output; the human study is still pending. Native caption adherence remains weak, and the reported training-distribution imprint suggests conditioning text may influence existing systems less than expected."
  ],
  "connected_context": {
    "meaning": "Agogic shifts symbolic-music model selection upstream: representation and decoding constraints can matter more than parameter count under controlled conditions. This reinforces prior evidence that architecture-adjacent data choices can outperform added compute, while narrowing the conclusion to distributional fidelity and instrument coverage because listener preference and strong caption adherence remain unestablished.",
    "corpus_size": 353,
    "generated_at": "2026-08-05T10:06:33.635Z",
    "connections": [
      {
        "title": "Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI",
        "source_name": "AI Engineer",
        "source_url": "https://www.youtube.com/watch?v=_PdK6x7PQNM",
        "feed7_url": "https://feed7.dev/p/data-quality-is-the-compute-multiplier-ari-morcos-datologyai-0x7k2ve",
        "reason": "Both show that upstream choices about how training material is represented or composed can outperform simply scaling compute; Agogic provides a controlled tokenization comparison specific to symbolic music."
      },
      {
        "title": "Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language Models",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2607.11801v1",
        "feed7_url": "https://feed7.dev/p/2607-11801v1-0m0m7wj",
        "reason": "IAAN changes acoustic behavior through a targeted inference-time intervention, while Agogic’s decode constraint similarly shows that measured audio capability can improve without enlarging or retraining the core model."
      }
    ]
  },
  "lifecycle": "Current",
  "published_at": "2026-08-04T17:56:49.000Z",
  "modified_at": "2026-08-04T17:56:49.000Z",
  "supersedes": [],
  "expires_at": null,
  "formats": {
    "html": "https://feed7.dev/p/2608-03999v1-1584qkh",
    "json": "https://feed7.dev/p/2608-03999v1-1584qkh.json",
    "markdown": "https://feed7.dev/p/2608-03999v1-1584qkh.md"
  }
}