papersTODAY 04:00 UTC
Paper Predicts llama.cpp Throughput From GGUF Metadata Using Roofline Models
A new arXiv preprint describes a method for estimating single-sequence inference throughput in llama.cpp directly from GGUF file metadata. The authors use roofline-shaped predictors with quantization-specific scaling factors fitted on reference models. The approach was scored on 318 phase-depth measurements drawn from 53 host-file configurations across three systems.