papersSEP 10 04:00 UTC
TokEval: An Evaluation Suite for Language Model Tokenizers
Researchers have released TokEval, a benchmark suite designed to compare tokenizers for language models in a systematic way. The work responds to the common practice of picking tokenizers with little scrutiny, even though tokenization decisions can influence what a model is ultimately able to do. By measuring tokenizer properties alongside their downstream effects, the suite aims to give practitioners a more rigorous basis for choosing one.