papersTODAY 04:00 UTC
Bellman Policy Optimization: Critic-Free RL Method for LLM Reasoning
Researchers present Bellman Policy Optimization (BPO), a reinforcement learning approach for training large language models with verifiable rewards that does not require a separate critic network. The method is derived from Policy Mirror Descent and targets autoregressive generation. It aims to improve reasoning performance in LLMs while simplifying the training setup.