Sign InOpen Brain
arXivPaperNeeds Review

Nuha-Speech: Building General-Purpose Arabic Speech-LLMs

Nuha-Speech adds an Arabic speech-QA corpus, Qwen-Omni fine-tuning, and a tailored evaluation framework—a useful blueprint for adapting speech models where language resources are scarce.

arXiv · Sep 10, 2026
Open Source Open MarkdownOpen JSON
Source Summary

Nuha-Speech combines an Arabic speech question-answering corpus with supervised fine-tuning and evaluation. The corpus contains **over 1.5 million training samples**, and training uses **Qwen-Omni variants** at multiple scales.

Practical Implication

Builders adapting speech agents to underrepresented languages should treat data construction, model tuning, and task-specific evaluation as one pipeline. Broad instruction coverage matters when existing speech resources are limited.

Agent-Ready Context
Nuha-Speech combines an Arabic speech question-answering corpus with supervised fine-tuning and evaluation. The corpus contains **over 1.5 million training samples**, and training uses **Qwen-Omni variants** at multiple scales.

Builders adapting speech agents to underrepresented languages should treat data construction, model tuning, and task-specific evaluation as one pipeline. Broad instruction coverage matters when existing speech resources are limited.

The abstract provides no benchmark results, licensing details, dialect coverage, or release terms. It therefore establishes the initiative’s scope, not how well the resulting models generalize or compare with alternatives.
Connected Context · Feed7 Judgment

This expands the speech-model landscape with an Arabic-focused data-and-tuning pipeline rather than a demonstrated model-selection result. Against candidates centered on realtime products, transcription, or training techniques, it emphasizes that underrepresented-language coverage begins with corpus construction and task-specific evaluation. Missing scores, dialect coverage, licensing, and release terms prevent judging generalization or deployability.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy AlignmentX³-OPD offers a method for transferring reasoning into audio models, whereas this Signal addresses the complementary prerequisite of building language-specific speech data and evaluations for adaptation.Gemini 3.5 Transcribe now available on AI GatewayGemini provides a deployable multilingual transcription route, but its missing accuracy evidence and this Signal’s missing benchmark results leave Arabic workload comparison unresolved and requiring task-specific testing.Hugging Face and Cerebras bring Gemma 4 to real-time voice AIThe open speech-to-speech pipeline shows how existing components can be assembled for deployment; Nuha-Speech instead concentrates on Arabic corpus creation and supervised tuning, with release terms still unknown.
Context Map
modelaudio
Uncertainty
The abstract provides no benchmark results, licensing details, dialect coverage, or release terms. It therefore establishes the initiative’s scope, not how well the resulting models generalize or compare with alternatives.