Sign InOpen Brain
arXivPaperNeeds Review

StudentSim: Training LLM-based Student Simulators

StudentSim turns sparse user histories into individualized simulators that model both current behavior and response to guidance, a pattern for testing adaptive agents before live deployment.

arXiv
Open Source Open MarkdownOpen JSON
Source Summary

StudentSim uses pooled training followed by per-student specialization. StudentSimEval covers **60 students** across chess, English writing, and mathematics, measuring behavioral fidelity and guidance responsiveness on the same de-identified learner records.

Practical Implication

For adaptive agents, separate imitation accuracy from response-to-intervention. A simulator used for evaluation or as a reward model should reproduce a user's baseline behavior and model how that behavior changes after guidance; fluent role-play alone does not establish either property.

Agent-Ready Context
StudentSim uses pooled training followed by per-student specialization. StudentSimEval covers **60 students** across chess, English writing, and mathematics, measuring behavioral fidelity and guidance responsiveness on the same de-identified learner records.

For adaptive agents, separate imitation accuracy from response-to-intervention. A simulator used for evaluation or as a reward model should reproduce a user's baseline behavior and model how that behavior changes after guidance; fluent role-play alone does not establish either property.

StudentSim beat GPT-5.4 on both metrics across all three domains. In chess it reached **F=0.51 and R=0.91**, versus **F=0.23 and R=0.72 for GPT-5.4**; however, the proof-of-concept tutor and expert ratings do not establish generalization beyond the evaluated learner datasets and tasks.
Context Map
benchmarkresearchdata#agent-evals#agent-reliability
Uncertainty
StudentSim beat GPT-5.4 on both metrics across all three domains. In chess it reached **F=0.51 and R=0.91**, versus **F=0.23 and R=0.72 for GPT-5.4**; however, the proof-of-concept tutor and expert ratings do not establish generalization beyond the evaluated learner datasets and tasks.