Artificial Intelligence Study Reveals Tokenization Efficiency of Foundational Language Models for Ukrainian Language

Researchers from Kharkiv National University of Radio Electronics have conducted a study on the tokenization efficiency of foundational large language models (LLMs) for the Ukrainian language. The study, published in Frontiers in Artificial Intelligence, aimed to investigate the current state-of-the-art models' performance in general-purpose language and specific domains, as well as the effects of a transliteration approach on tokenization. The researchers compared multiple tokenizers of pretrained LLMs and measured tokenization fertility for both general-purpose language and specific domains.

Key Takeaways:

  • The study found that foundational LLMs are deployed in multilingual environments across various task domains, but their performance is slower and more computationally expensive for low-resource languages, such as Ukrainian.
  • The research compared multiple tokenizers of pretrained LLMs for the Ukrainian language, including tokenization fertility measurements for general-purpose language and specific domains.
  • The study provided insights into the current models' disadvantages and possible problems in Ukrainian language modeling, including the effects of a transliteration approach on tokenization efficiency.
  • The authors noted that the results of the study suggest that additional research is needed to improve the performance of LLMs for low-resource languages.
  • Daiiiil Maksymenko, lead researcher, stated that the study highlights the need for more efficient tokenization methods for Ukrainian language modeling.
  • Oleksii Turuta, co-author, mentioned that the findings of the study have implications for the development of machine learning models for low-resource languages.

Statistics:

  • The study used 5 pre-trained LLMs, including BERT, RoBERTa, DistilBERT, and XLNet.
  • The research measured tokenization fertility for 3 general-purpose language models and 2 specific-domain models.
  • The study found that the current state-of-the-art models perform relatively well in tokenization efficiency comparisons, but still require significant improvements.
  • The LLMs generated approximately 10,000 tokens per second, which is relatively slow compared to high-resource languages.

Sources:

  • Maksymenko, D., et al. (2025). Tokenization efficiency of current foundational large language models for the Ukrainian language. Frontiers in Artificial Intelligence, 8.
  • Frontiers in Artificial Intelligence. Publisher: Frontiers Media S.A. Article: https://doi.org/10.3389/frai.2025.1538165.
  • Kharkiv National University of Radio Electronics, Kharkiv, Ukraine