Sign InOpen Brain
arXivPaperNeeds Review

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure

LittleLearner offers a controlled model and corpus for studying knowledge acquisition without unknown prior exposure. Its initial results separate better use of known material from new capability.

arXiv
Open Source Open MarkdownOpen JSON
Source Summary

LITTLECURRICULUM contains **88B tokens** aligned to U.S. elementary material while excluding concepts, facts, and vocabulary taught above Grade 5. A **5B-parameter** model trained from scratch provides enough language ability for open-ended evaluation within those boundaries.

Practical Implication

This gives researchers a cleaner sandbox for testing post-training, prompting, and in-context learning. Builders evaluating adaptation methods can use the setup to distinguish retrieval or recombination of known material from acquisition beyond the training scope.

Agent-Ready Context
LITTLECURRICULUM contains **88B tokens** aligned to U.S. elementary material while excluding concepts, facts, and vocabulary taught above Grade 5. A **5B-parameter** model trained from scratch provides enough language ability for open-ended evaluation within those boundaries.

This gives researchers a cleaner sandbox for testing post-training, prompting, and in-context learning. Builders evaluating adaptation methods can use the setup to distinguish retrieval or recombination of known material from acquisition beyond the training scope.

The first experiments found better use of existing knowledge but **no increase in out-of-scope capabilities**. The material does not establish whether that result generalizes to broader corpora, larger models, or agent workflows.
Context Map
benchmarkresearch#benchmark-integrity#reasoning
Uncertainty
The first experiments found better use of existing knowledge but **no increase in out-of-scope capabilities**. The material does not establish whether that result generalizes to broader corpora, larger models, or agent workflows.