# Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language Reasoning

Source: [arXiv](https://arxiv.org/abs/2609.10445v1)  
Feed7 permalink: https://feed7.dev/p/2609-10445v1-0tad2cs  
Published: 2026-09-09T16:56:25.000Z  
Trust: Needs Review (needs_review)

## Why Included

Tiny Aya L2-Thinker shows multilingual reasoning can transfer through data mixing, suggesting builders should evaluate whether agents reason in the user's language, not only answer in it.

## Source Summary

Tiny Aya L2-Thinker is a **3.35B-parameter** model trained through a data-centric SFT recipe. It reports an in-language reasoning rate above **93%** across **60 languages** and **6 benchmarks**, covering math, commonsense, instructions, open generation, and cultural reasoning.

## Practical Implication

For multilingual agents, test the language of intermediate reasoning as well as final answers. The reported recipe combines broad language coverage, multilingual non-reasoning data, and a sufficient English reasoning base rather than requiring reasoning supervision for every target language.

## Agent-Ready Context

Tiny Aya L2-Thinker is a **3.35B-parameter** model trained through a data-centric SFT recipe. It reports an in-language reasoning rate above **93%** across **60 languages** and **6 benchmarks**, covering math, commonsense, instructions, open generation, and cultural reasoning.

For multilingual agents, test the language of intermediate reasoning as well as final answers. The reported recipe combines broad language coverage, multilingual non-reasoning data, and a sufficient English reasoning base rather than requiring reasoning supervision for every target language.

The material gives aggregate coverage but no language-level scores, failure cases, or comparisons. Claims about transfer to held-out languages therefore need validation on the exact languages and agent tasks a product serves.

## Connected Context

Feed7 judgment across 732 accumulated Signals:

This makes training-data composition, rather than per-language reasoning supervision, the central lever for compact multilingual reasoning. It also adds a stricter product evaluation requirement: verify the language used in intermediate reasoning, not merely the final response. The aggregate result supports broad transfer, but does not yet identify which languages or agent tasks benefit reliably.

- [Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI](https://feed7.dev/p/data-quality-is-the-compute-multiplier-ari-morcos-datologyai-0x7k2ve) — It gives concrete multilingual support to the broader claim that balanced, task-shaped data mixtures can improve capability without simply increasing model size or compute.
- [OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling](https://feed7.dev/p/2608-05141v1-0kai3a6) — Both make data structure and mixture a model-capability intervention: Tiny Aya targets cross-language reasoning transfer, while OctoLong targets dependency-aware long-context coding.
- [DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data](https://feed7.dev/p/2608-13517v1-10qer54) — It broadens the compact multilingual model evidence beyond Mimir’s Danish-oriented case, while both still require language- and task-level evaluation before deployment comparisons are justified.

## Context Map

- Layer: model
- Domains: research
- Topics: reasoning, open-models

## Uncertainty

- The material gives aggregate coverage but no language-level scores, failure cases, or comparisons. Claims about transfer to held-out languages therefore need validation on the exact languages and agent tasks a product serves.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
