Augustinian BabyLM: What Ostensive Definition Can and Cannot Teach a Small Language Model
Visual embedding initialization gave a small language model durable object-property knowledge that most standard benchmarks missed, showing how broad eval suites can hide narrow causal gains.
A DeBERTa model trained on **10M words** received image-derived embeddings for visually grounded tokens. The effect persisted through training and improved **COMPS in every configuration**, plus a tailored color, material, size, and shape test.
When testing a new initialization or grounding method, add probes aligned with the injected knowledge and track results per affected token. Aggregate language benchmarks may miss a real but localized capability change.
A DeBERTa model trained on **10M words** received image-derived embeddings for visually grounded tokens. The effect persisted through training and improved **COMPS in every configuration**, plus a tailored color, material, size, and shape test. When testing a new initialization or grounding method, add probes aligned with the injected knowledge and track results per affected token. Aggregate language benchmarks may miss a real but localized capability change. Most BabyLM benchmarks showed no effect, and gains on the tailored test stayed confined to seeded words. Even lower held-out loss for visually seeded function and abstract words was not captured by any benchmark used, leaving the right evaluation unresolved.
This confirms that a real, localized learning effect can remain invisible to broad aggregate benchmarks. It narrows the grounding claim to seeded tokens and aligned probes, so the result supports controlled, capability-specific evaluation rather than general language improvement. The unexplained held-out-loss change also shows that even targeted suites may omit affected behaviors and should not be treated as exhaustive.