Sign InOpen Brain
arXivPaperNeeds Review

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment

X³-OPD transfers a text model’s reasoning into an audio-language model while grounding training in the student’s own acoustic interpretations, including events, prosody, and dialogue.

arXiv · Jul 23, 2026
Open Source Open MarkdownOpen JSON
Source Summary

**X³-OPD** is an on-policy distillation framework: an audio-language student generates reasoning from acoustic input, while a text teacher supplies token-level guidance from matched transcripts and verified answers.

Practical Implication

The **three-tier corpus** spans speech-rendered text problems, complex audio-event reasoning, and spoken dialogue with paralinguistic cues. Builders working on voice or audio agents should treat non-verbal sound and prosody as reasoning inputs, not merely transcription problems.

Agent-Ready Context
**X³-OPD** is an on-policy distillation framework: an audio-language student generates reasoning from acoustic input, while a text teacher supplies token-level guidance from matched transcripts and verified answers.

The **three-tier corpus** spans speech-rendered text problems, complex audio-event reasoning, and spoken dialogue with paralinguistic cues. Builders working on voice or audio agents should treat non-verbal sound and prosody as reasoning inputs, not merely transcription problems.

Results across **MMSU, MMAU, BIG Bench Audio, and MMAR** indicate better audio-grounded reasoning and chain-of-thought quality with existing abilities largely preserved under domain shift. The abstract provides no effect sizes, implementation details, or released artifacts.
Context Map
modelaudio#reasoning
Uncertainty
Results across **MMSU, MMAU, BIG Bench Audio, and MMAR** indicate better audio-grounded reasoning and chain-of-thought quality with existing abilities largely preserved under domain shift. The abstract provides no effect sizes, implementation details, or released artifacts.