papersTODAY 04:00 UTC
Teacher-Guided Curriculum Boosts Data Efficiency in RLVR Training
A new arXiv paper addresses a known failure mode in reinforcement learning with verifiable rewards (RLVR), where training problems that are too hard for a model produce uniformly failed attempts and yield no learning signal. The authors propose a teacher-guided curriculum that sequences training data so the model encounters problems it can actually solve, making the process more data-efficient. The work targets mathematical reasoning in large language models and falls within the cs.CL area.