Study identifies repetition-induced label flips in LLM guardrail classifiers
A new arXiv paper describes "overflip," a failure mode where guardrail models used to screen malicious prompts and responses change their classification when input context is repeated. The authors focus on lightweight Transformer-based guardrails, such as DeBERTa variants, which are common in latency-sensitive deployments and are trained on short contexts. The work suggests these compact classifiers can be unreliable under repeated or padded input.