papersSEP 10 04:00 UTC
Statistical Method Proposed for Determining Sample Sizes in Machine Learning Prediction Models
Researchers have introduced a statistical framework for estimating how much data is needed to train machine learning prediction models. The approach addresses a key limitation of conventional power analysis, which normally requires the predictor-outcome relationship and effect structure to be defined in advance—something that is impractical for nonlinear models that learn complex patterns from data. The work appears in a new arXiv preprint filed under both artificial intelligence and machine learning categories.