papersSEP 10 04:00 UTC
Time-Series Foundation Model Benchmarks Still Reflect Pretraining Familiarity on Later Hold-Outs
A new study questions whether time-series foundation models can be fairly evaluated using test data collected after their pretraining cutoff. It finds that even a temporally later, contamination-free hold-out does not fully isolate genuine generalization, as familiarity with the underlying data distribution absorbed during pretraining persists. The result suggests the field needs evaluation practices that go beyond simply withholding recent data.