MEGA Hub

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure

Authors

Do you know Fanfei Li?You can claim authorship or link another user.Do you know Jana Zeller?You can claim authorship or link another user.Do you know Manuel Prada-Corral?You can claim authorship or link another user.Do you know Thaddäus Wiedemer?You can claim authorship or link another user.Do you know Prasanna Mayilvahanan?You can claim authorship or link another user.Do you know Ryan Cotterell?You can claim authorship or link another user.Do you know Wieland Brendel?You can claim authorship or link another user.

Abstract

Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to characterize. To address this challenge, we introduce LITTLECURRICULUM, a curated 88B-token pretraining corpus tailored to U.S. elementary school material, explicitly excluding concepts, facts, and vocabulary taught above Grade 5. Training a 5B-parameter LLM from scratch on LITTLECURRICULUM yields LITTLELEARNER, a model with sufficient language competence for open-ended evaluation, yet with clear knowledge and capability boundaries mapped to interpretable curriculum guidelines. We release LITTLECURRICULUM and LITTLELEARNER as a developmentally restricted sandbox to study how models acquire, represent, and use data under a well-defined training scope. We illustrate the sandbox's utility in a first suite of experiments on injecting new knowledge through post-training and in-context learning. These methods let LITTLELEARNER better utilize existing knowledge, but do not raise out-of-scope capabilities. Our findings underscore the value of this controlled environment for future investigations.

Community

00