Paper Proposes One-Step Flow Policy for Offline Reinforcement Learning
A new arXiv paper introduces a method for learning multimodal one-step flow policies from fixed offline datasets using value-weighted optimal transport. The approach targets offline reinforcement learning, where action distributions are often multimodal and existing flow policies are slow to sample. It appears in both cs.AI and cs.LG cross-listings.