Natural language processing research has achieved remarkable progress with the development of word embedding techniques; however, most existing studies focus on high-resource languages. Low-resource and morphologically rich languages such as Uzbek still face significant challenges related to data sparsity and vocabulary variability. This paper presents a comparative analysis of static, subword-aware, and contextual embedding models for the Uzbek language within a unified and reproducible experimental framework. Specifically, Word2Vec, FastText, and multilingual BERT (mBERT) are evaluated using intrinsic evaluation metrics, including vocabulary size, out-of-vocabulary (OOV) coverage, and average cosine similarity. The experiments are conducted on a preprocessed Uzbek text corpus that combines educational, scientific, and synthetically generated data to mitigate resource limitations. The results demonstrate that Word2Vec struggles with morphological variation and OOV words, while FastText significantly improves vocabulary coverage through subword modeling. Furthermore, mBERT produces rich contextual sentence-level representations, capturing semantic information beyond the capabilities of static embeddings. The findings highlight the importance of subword-aware and contextual embedding approaches for low-resource languages and provide practical insights into embedding selection for Uzbek natural language processing tasks. The proposed framework provides a reproducible baseline for future research on embedding models and downstream NLP applications for the Uzbek language. All experiments and results are generated using an open-source codebase, ensuring full transparency and reproducibility through a publicly available repository.
Keywords
Uzbek LanguageWord EmbeddingsWord2VecFastTextMultilingual BERTLow-Resource NLPContextual Embeddings
References
T. Mikolov, K. Chen, G. Corrado, and J. Dean, “Efficient Estimation of Word Representations in Vector Space,” arXiv preprint arXiv:1301.3781, 2013, [Online]. Available: https://doi.org/10.48550/arXiv.1301.3781.
J. Pennington, R. Socher, and C. D. Manning, “GloVe: Global Vectors for Word Representation,” in Proc. EMNLP, 2014, pp. 1532-1543, [Online]. Available: https://doi.org/10.3115/v1/D14-1162.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proc. NAACL-HLT, 2019, pp. 4171-4186, [Online]. Available: https://doi.org/10.18653/v1/N19-1423.
T. Pires, E. Schlinger, and D. Garrette, “How Multilingual Is Multilingual BERT?” in Proc. ACL, 2019, pp. 4996-5001, [Online]. Available: https://doi.org/10.18653/v1/P19-1493.
M. E. Peters et al., “Deep Contextualized Word Representations,” in Proc. NAACL-HLT, 2018, pp. 2227-2237, [Online]. Available: https://doi.org/10.18653/v1/N18-1202.
P. Bojanowski, E. Grave, A. Joulin, and T. Mikolov, “Enriching Word Vectors with Subword Information,” Transactions of the Association for Computational Linguistics, vol. 5, pp. 135-146, 2017, [Online]. Available: https://doi.org/10.1162/tacl_a_00051.
A. Joulin, E. Grave, P. Bojanowski, and T. Mikolov, “Bag of Tricks for Efficient Text Classification,” in Proc. EACL, 2017, pp. 427-431, [Online]. Available: https://doi.org/10.18653/v1/E17-2068.
A. Rogers, O. Kovaleva, and A. Rumshisky, “A Primer in BERTology: What We Know About How BERT Works,” Transactions of the Association for Computational Linguistics, vol. 8, pp. 842-866, 2020, [Online]. Available: https://doi.org/10.1162/tacl_a_00349.
R. Cotterell et al., “On the Complexity and Typology of Inflectional Morphological Systems,” Transactions of the Association for Computational Linguistics, vol. 6, pp. 327-342, 2018, [Online]. Available: https://doi.org/10.1162/tacl_a_00026.
Y. Goldberg, Neural Network Methods for Natural Language Processing. San Rafael, CA, USA: Morgan & Claypool, 2017, [Online]. Available: https://doi.org/10.2200/S00762ED1V01Y201703HLT037.
D. Jurafsky and J. H. Martin, Speech and Language Processing, 3rd ed., draft, 2023, [Online]. Available: https://doi.org/10.48550/arXiv.2308.08246.
M. Artetxe, S. Ruder, and D. Yogatama, “On the Cross-lingual Transferability of Monolingual Representations,” in Proc. ACL, 2020, pp. 4623-4637, [Online]. Available: https://doi.org/10.18653/v1/2020.acl-main.421.
G. Lample and A. Conneau, “Cross-lingual Language Model Pretraining,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 32, 2019, [Online]. Available: https://doi.org/10.48550/arXiv.1901.07291.