arXivPaperNeeds Review
LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure
LittleLearner offers a controlled model and corpus for studying knowledge acquisition without unknown prior exposure. Its initial results separate better use of known material from new capability.
arXiv
Source Summary
LITTLECURRICULUM contains **88B tokens** aligned to U.S. elementary material while excluding concepts, facts, and vocabulary taught above Grade 5. A **5B-parameter** model trained from scratch provides enough language ability for open-ended evaluation within those boundaries.
Practical Implication
This gives researchers a cleaner sandbox for testing post-training, prompting, and in-context learning. Builders evaluating adaptation methods can use the setup to distinguish retrieval or recombination of known material from acquisition beyond the training scope.
Agent-Ready Context
LITTLECURRICULUM contains **88B tokens** aligned to U.S. elementary material while excluding concepts, facts, and vocabulary taught above Grade 5. A **5B-parameter** model trained from scratch provides enough language ability for open-ended evaluation within those boundaries. This gives researchers a cleaner sandbox for testing post-training, prompting, and in-context learning. Builders evaluating adaptation methods can use the setup to distinguish retrieval or recombination of known material from acquisition beyond the training scope. The first experiments found better use of existing knowledge but **no increase in out-of-scope capabilities**. The material does not establish whether that result generalizes to broader corpora, larger models, or agent workflows.
Context Map
benchmarkresearch#benchmark-integrity#reasoningUncertainty
The first experiments found better use of existing knowledge but **no increase in out-of-scope capabilities**. The material does not establish whether that result generalizes to broader corpora, larger models, or agent workflows.