Nuha-Speech: Building General-Purpose Arabic Speech-LLMs
Nuha-Speech adds an Arabic speech-QA corpus, Qwen-Omni fine-tuning, and a tailored evaluation framework—a useful blueprint for adapting speech models where language resources are scarce.
Nuha-Speech combines an Arabic speech question-answering corpus with supervised fine-tuning and evaluation. The corpus contains **over 1.5 million training samples**, and training uses **Qwen-Omni variants** at multiple scales.
Builders adapting speech agents to underrepresented languages should treat data construction, model tuning, and task-specific evaluation as one pipeline. Broad instruction coverage matters when existing speech resources are limited.
Nuha-Speech combines an Arabic speech question-answering corpus with supervised fine-tuning and evaluation. The corpus contains **over 1.5 million training samples**, and training uses **Qwen-Omni variants** at multiple scales. Builders adapting speech agents to underrepresented languages should treat data construction, model tuning, and task-specific evaluation as one pipeline. Broad instruction coverage matters when existing speech resources are limited. The abstract provides no benchmark results, licensing details, dialect coverage, or release terms. It therefore establishes the initiative’s scope, not how well the resulting models generalize or compare with alternatives.
This expands the speech-model landscape with an Arabic-focused data-and-tuning pipeline rather than a demonstrated model-selection result. Against candidates centered on realtime products, transcription, or training techniques, it emphasizes that underrepresented-language coverage begins with corpus construction and task-specific evaluation. Missing scores, dialect coverage, licensing, and release terms prevent judging generalization or deployability.