Accountable and uncertainty-aware evaluation of sensor-based AI under distribution shift
A new machine learning paper tackles the gap between training conditions and real-world deployment for sensor-based AI, where devices, personnel, and time periods all differ from the original training setup. The authors propose a staged evaluation methodology that quantifies uncertainty and captures performance degradation that conventional random train-test splits can hide. The work draws on data collected from multiple devices and subjects over nearly three years in an underground environment.