papersSEP 10 04:00 UTC
TRACE trains reasoning agents for causal exploration with synthesized rewards
A new arXiv paper presents TRACE, a training method that brings reinforcement learning with verifiable rewards to diagnostic reasoning, where correct answers over complex data are hard to check automatically. The approach synthesizes reward signals so agents can learn to perform causal exploration. It is listed under cs.AI with a cross-listing in cs.LG.