papersTODAY 04:00 UTC
GRPO-QM Uses Reinforcement Learning to Guide Quantum Tomography Without Distorting Posteriors
A new arXiv preprint presents GRPO-QM, a method that applies group-relative policy optimization to quantum tomography while leaving the target posterior distribution unchanged. Instead of letting reward-based learning reshape the inference target, the approach learns only an exploration policy that selects measurements. This separates the strategy for gathering data from the statistical estimate itself.