Sign InOpen Brain
arXivPaperNeeds Review

Augustinian BabyLM: What Ostensive Definition Can and Cannot Teach a Small Language Model

Visual embedding initialization gave a small language model durable object-property knowledge that most standard benchmarks missed, showing how broad eval suites can hide narrow causal gains.

arXiv · Sep 10, 2026
Open Source Open MarkdownOpen JSON
Source Summary

A DeBERTa model trained on **10M words** received image-derived embeddings for visually grounded tokens. The effect persisted through training and improved **COMPS in every configuration**, plus a tailored color, material, size, and shape test.

Practical Implication

When testing a new initialization or grounding method, add probes aligned with the injected knowledge and track results per affected token. Aggregate language benchmarks may miss a real but localized capability change.

Agent-Ready Context
A DeBERTa model trained on **10M words** received image-derived embeddings for visually grounded tokens. The effect persisted through training and improved **COMPS in every configuration**, plus a tailored color, material, size, and shape test.

When testing a new initialization or grounding method, add probes aligned with the injected knowledge and track results per affected token. Aggregate language benchmarks may miss a real but localized capability change.

Most BabyLM benchmarks showed no effect, and gains on the tailored test stayed confined to seeded words. Even lower held-out loss for visually seeded function and abstract words was not captured by any benchmark used, leaving the right evaluation unresolved.
Connected Context · Feed7 Judgment

This confirms that a real, localized learning effect can remain invisible to broad aggregate benchmarks. It narrows the grounding claim to seeded tokens and aligned probes, so the result supports controlled, capability-specific evaluation rather than general language improvement. The unexplained held-out-loss change also shows that even targeted suites may omit affected behaviors and should not be treated as exhaustive.

Context Map
benchmarkresearchimage#benchmark-integrity
Uncertainty
Most BabyLM benchmarks showed no effect, and gains on the tailored test stayed confined to seeded words. Even lower held-out loss for visually seeded function and abstract words was not captured by any benchmark used, leaving the right evaluation unresolved.