Sign InOpen Brain
arXivPaperNeeds Review

Do speech foundation models really learn words?

HuBERT and wav2vec 2.0 appear to encode word identity beyond local phonetics in later layers. The paper offers a cleaner probe for builders evaluating speech representations.

arXiv · Sep 9, 2026
Open Source Open MarkdownOpen JSON
Source Summary

The authors use **residualization** to remove phoneme information before probing speech representations. In **later layers**, HuBERT and wav2vec 2.0 still encode words with reasonable fidelity, suggesting their word discrimination is not solely phonetic.

Practical Implication

Builders evaluating speech encoders should test what remains after obvious low-level signals are controlled for. The same disentangling approach may also make higher-order linguistic information more useful for **word discovery**.

Agent-Ready Context
The authors use **residualization** to remove phoneme information before probing speech representations. In **later layers**, HuBERT and wav2vec 2.0 still encode words with reasonable fidelity, suggesting their word discrimination is not solely phonetic.

Builders evaluating speech encoders should test what remains after obvious low-level signals are controlled for. The same disentangling approach may also make higher-order linguistic information more useful for **word discovery**.

The material reports the general finding but provides no effect sizes, dataset details, or model comparison numbers. It therefore supports a probing method and qualitative conclusion, not a production recommendation.
Connected Context · Feed7 Judgment

This strengthens benchmark-integrity practice for representation probes: apparent word knowledge should be tested after removing an obvious phonetic shortcut. The finding suggests later speech layers retain higher-order lexical information, but mainly confirms a disentangling method; absent effect sizes and dataset details, it cannot support encoder selection or establish how useful that information is downstream.

Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization EvaluationBoth move beyond headline probe scores by intervening on internal evidence, separating genuine higher-order representation from performance driven by an easier underlying signal.LittleLearner: Language Models Under Pedagogically Controlled Knowledge ExposureResidualizing phonemes parallels LittleLearner’s controlled exposure design: each removes a plausible alternative explanation before attributing a capability to the model.FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language ModelsFriendBench’s modality analysis reinforces the same evaluation principle: aggregate success does not establish that the intended information source actually contributed to the result.Surprisal Theory is Tautological (without Rational Grounding)The critique of unconstrained surprisal supports the need for controls like residualization, because fit alone is weak evidence when a simpler signal can explain the measured pattern.
Context Map
benchmarkaudioresearch#benchmark-integrity
Uncertainty
The material reports the general finding but provides no effect sizes, dataset details, or model comparison numbers. It therefore supports a probing method and qualitative conclusion, not a production recommendation.