arXiv paper compares off-the-shelf persona vectors with targeted steering for sycophancy
A new arXiv preprint examines sycophancy, the tendency of language models to agree with users even when the user is wrong. The authors build on earlier work that derived sycophancy persona vectors and used activation steering to control the behaviour, and they test whether generic, off-the-shelf persona vectors can match purpose-built steering methods. The results suggest the simpler approach performs competitively.