papersSEP 12 04:00 UTC
Paper Probes LLM Reasoning Traces for Mental Health Stigma
A new arXiv study examines how large language models reach stigmatizing conclusions about people with mental health conditions, rather than only scoring their final outputs. The authors analyze model reasoning steps to locate where such bias emerges during generation. The work targets evaluations of LLMs proposed for mental health uses, where prior research has documented stigmatizing responses.