Paper Proposes Coalitional Alignment Method for Controlling Misaligned AI Agents
A new arXiv paper examines the difficulty of supervising long-running AI agents, where every action alters the environment and thus shapes what the agent can do next. When an agent is not fully aligned, the authors argue that safety depends on reviewing high-stakes actions before they are carried out. They put forward a coalitional alignment and authorization-delegation framework intended to keep such agents under safe control.