Paper Proposes Counterfactual Marginalisation to Test Model Robustness
A new arXiv paper introduces counterfactual marginalisation, a test-time procedure for measuring how much a classifier depends on nuisance variables such as demographic or acquisition-related shortcuts. The method aims to expose cases where models score well on test sets despite relying on spurious cues rather than genuine signal. The authors frame it as an evaluation tool rather than a training technique.