Tokenization: A Survey for Modern NLP

Abstract

This survey examines tokenization in modern language models, covering algorithms, theory, evaluation, and practical trade-offs. It explains the benefits and limitations of subword tokenization, including multilingual vocabulary allocation, segmentation ambiguities, security concerns, and the constraints of static vocabularies. The survey maps current methods and bottlenecks and identifies open questions for researchers and practitioners.

Publication
Preprint 2026
Konstantin Dobler
Konstantin Dobler
Ph.D. Student in ML & NLP

I’m an ELLIS Ph.D. student at Hasso Plattner Institute.