Checkpoint Selection and Evaluation in EEG Emotion Recognition
A study examines how choosing model checkpoints can inflate reported electroencephalography-based emotion recognition scores without any real gain in trial-level performance. The authors compare selection and scoring across separate trial pools along fixed training trajectories. The findings suggest that same-session evaluation practices can distort benchmark comparisons in this field.