Paper Analyzes Selection Bias When Model Edits Target Localized Spans
A new arXiv paper examines what happens when human corrections are applied only to identified editable spans of a model's output. The authors decompose the localized gradient into edited and untouched portions at a fixed checkpoint, showing that selective feedback channels can amplify relative selection bias. They also study gradient geometry, target mismatch, and importance weighting as factors in this effect.