papersTODAY 04:00 UTC
Compact Policies for Submodular MDPs via LP-Based Submodular Orienteering
A new arXiv paper introduces an approach for deriving strong yet compact action-selection policies in Markov Decision Processes whose value functions are submodular. The method builds on a linear-programming formulation of submodular orienteering, a problem where an agent must reach a set of targets under a budget. The authors argue this yields policies that are both effective and compact, relevant to reinforcement learning and operations research settings where repeated action choice is required.