papersSEP 10 04:00 UTC
New protocol IBIB scores enterprise AI deployments by serving route rather than model identifier
A cross-listed arXiv paper introduces IBIB, a protocol for evaluating AI systems as they are actually deployed inside enterprises rather than as bare model checkpoints. The authors argue that real-world capability emerges from the combination of weights, serving configuration, precision, output contract, and harness, so scoring an advertised model name alone is a measurement error. After auditing 18 existing benchmarks and finding that every one grades model identifiers, the paper proposes routing-based measurement instead.