Proceedings of International Conference on Applied Innovation in IT  ·  2026/07/22  ·  Vol. 14  ·  Issue 4  ·  pp. 1261–1268
Comparative Study of Word2Vec, FastText and Multilingual BERT for Uzbek Natural Language Processing
Abdilatif Meyliev, Farrukh Ishkobilov, Mironshoh Ortikov, Leyla Cumayeva and Javlon Ibragimov
Natural language processing research has achieved remarkable progress with the development of word embedding techniques; however, most existing studies focus on high-resource languages. Low-resource and morphologically rich languages such as Uzbek still face significant challenges related to data sparsity and vocabulary variability. This paper presents a comparative analysis of static, subword-aware, and contextual embedding models for the Uzbek language within a unified and reproducible experimental framework. Specifically, Word2Vec, FastText, and multilingual BERT (mBERT) are evaluated using intrinsic evaluation metrics, including vocabulary size, out-of-vocabulary (OOV) coverage, and average cosine similarity. The experiments are conducted on a preprocessed Uzbek text corpus that combines educational, scientific, and synthetically generated data to mitigate resource limitations. The results demonstrate that Word2Vec struggles with morphological variation and OOV words, while FastText significantly improves vocabulary coverage through subword modeling. Furthermore, mBERT produces rich contextual sentence-level representations, capturing semantic information beyond the capabilities of static embeddings. The findings highlight the importance of subword-aware and contextual embedding approaches for low-resource languages and provide practical insights into embedding selection for Uzbek natural language processing tasks. The proposed framework provides a reproducible baseline for future research on embedding models and downstream NLP applications for the Uzbek language. All experiments and results are generated using an open-source codebase, ensuring full transparency and reproducibility through a publicly available repository.
Uzbek Language Word Embeddings Word2Vec FastText Multilingual BERT Low-Resource NLP Contextual Embeddings
References
  1. T. Mikolov, K. Chen, G. Corrado, and J. Dean, “Efficient Estimation of Word Representations in Vector Space,” arXiv preprint arXiv:1301.3781, 2013, [Online]. Available: https://doi.org/10.48550/arXiv.1301.3781.
  2. J. Pennington, R. Socher, and C. D. Manning, “GloVe: Global Vectors for Word Representation,” in Proc. EMNLP, 2014, pp. 1532-1543, [Online]. Available: https://doi.org/10.3115/v1/D14-1162.
  3. J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proc. NAACL-HLT, 2019, pp. 4171-4186, [Online]. Available: https://doi.org/10.18653/v1/N19-1423.
  4. T. Pires, E. Schlinger, and D. Garrette, “How Multilingual Is Multilingual BERT?” in Proc. ACL, 2019, pp. 4996-5001, [Online]. Available: https://doi.org/10.18653/v1/P19-1493.
  5. M. E. Peters et al., “Deep Contextualized Word Representations,” in Proc. NAACL-HLT, 2018, pp. 2227-2237, [Online]. Available: https://doi.org/10.18653/v1/N18-1202.
  6. P. Bojanowski, E. Grave, A. Joulin, and T. Mikolov, “Enriching Word Vectors with Subword Information,” Transactions of the Association for Computational Linguistics, vol. 5, pp. 135-146, 2017, [Online]. Available: https://doi.org/10.1162/tacl_a_00051.
  7. A. Joulin, E. Grave, P. Bojanowski, and T. Mikolov, “Bag of Tricks for Efficient Text Classification,” in Proc. EACL, 2017, pp. 427-431, [Online]. Available: https://doi.org/10.18653/v1/E17-2068.
  8. A. Rogers, O. Kovaleva, and A. Rumshisky, “A Primer in BERTology: What We Know About How BERT Works,” Transactions of the Association for Computational Linguistics, vol. 8, pp. 842-866, 2020, [Online]. Available: https://doi.org/10.1162/tacl_a_00349.
  9. R. Cotterell et al., “On the Complexity and Typology of Inflectional Morphological Systems,” Transactions of the Association for Computational Linguistics, vol. 6, pp. 327-342, 2018, [Online]. Available: https://doi.org/10.1162/tacl_a_00026.
  10. Y. Goldberg, Neural Network Methods for Natural Language Processing. San Rafael, CA, USA: Morgan & Claypool, 2017, [Online]. Available: https://doi.org/10.2200/S00762ED1V01Y201703HLT037.
  11. D. Jurafsky and J. H. Martin, Speech and Language Processing, 3rd ed., draft, 2023, [Online]. Available: https://doi.org/10.48550/arXiv.2308.08246.
  12. M. Artetxe, S. Ruder, and D. Yogatama, “On the Cross-lingual Transferability of Monolingual Representations,” in Proc. ACL, 2020, pp. 4623-4637, [Online]. Available: https://doi.org/10.18653/v1/2020.acl-main.421.
  13. G. Lample and A. Conneau, “Cross-lingual Language Model Pretraining,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 32, 2019, [Online]. Available: https://doi.org/10.48550/arXiv.1901.07291.


Proceedings of the International Conference on Applied Innovations in IT by Anhalt University of Applied Sciences is licensed under CC BY-SA 4.0
 ·  This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License

ICAIIT 2026
International Conference on Applied Innovation in IT
Navigation
Publisher
ISSN2199-8876
Location Anhalt University of Applied Sciences
Phone +49 (0) 3496 67 5611
Address Building 01, Room 425
Bernburger Str. 55
D-06366 Köthen, Germany
Open Access License

All works are licensed under the Creative Commons Attribution-ShareAlike 4.0 International License (CC BY-SA 4.0), unless otherwise noted.

Published by ICAIIT in cooperation with Anhalt University of Applied Sciences.

© 2026 ICAIIT — International Conference on Applied Innovations in IT. Anhalt University of Applied Sciences, Köthen, Germany.
Visitors: site traffic counter