papersSEP 10 04:00 UTC
BRACE paper proposes anchored Bellman-residual correction for stale critics in asynchronous RL
A new arXiv preprint introduces BRACE, a method aimed at value-function staleness in asynchronous reinforcement learning. As training of language models increasingly relies on asynchronous setups, delays between acting and learning bias the critic toward outdated policies, while prior asynchronous-training fixes targeted only the actor. The proposed approach applies an anchored Bellman-residual correction to keep the critic aligned with the current policy.