papersSEP 10 04:00 UTC
Paper shows fixed-rollout pass@k evaluations identify only limited information
A new paper examines the common practice of extrapolating pass@k benchmark results to attempt counts larger than the number of samples actually collected per problem. Under a pooled conditional-Binomial model, the authors show that success counts from fixed-size rollouts determine only a finite number of distribution moments. The result implies such evaluations cannot fully characterize model performance well beyond the sampled regime.