Paper Explores Steering Category-Specific Refusal Directions in Language Models
A new arXiv paper examines safety alignment in language models, focusing on models fine-tuned to emit distinct refusal tokens that signal different categories of refusal before they answer. The authors investigate refusal directions tied to specific categories and how those directions might be discovered and steered. The abstract provided is truncated, so the full method and results are not available here.