papersTODAY 04:00 UTC
Study Measures How Unstable LLM Refusals Are Near Safety Boundaries
A research paper examines how inconsistently large language models refuse prompts, particularly when benign requests are worded similarly to risky content. The authors propose measuring this confusion within local safety boundaries to better characterize refusal reliability. The work appears as a revised arXiv preprint in the cs.CL and cs.AI categories.