papersSEP 10 04:00 UTC
SAC's tanh squashing found to hinder bang-bang control in negative-result study
An arXiv paper examines how the tanh squashing in Soft Actor-Critic shrinks the policy gradient as actions approach their bounds, potentially starving learning when near-saturated outputs are required. The authors report a negative result, showing this throttling impedes bang-bang control tasks and driving experiments in the MetaDrive simulator. The findings point to a practical limitation of tanh-squashed Gaussian policies in continuous reinforcement learning.