Proceedings of International Conference on Applied Innovation in IT  ·  2026/07/22  ·  Vol. 14  ·  Issue 4  ·  pp. 1211–1216
Corpus-Based Structural and Semantic Analysis of IT Educational Terminology in a Bilingual English-Uzbek Corpus
Malika Tilavova, Oybek Akhmedov, Handoko Hum and Lobar Sultonova
This study provides a reproducible, corpus-based framework for the structural and semantic analysis of educational terminology related to information technology in a bilingual (English-Uzbek) context. A comprehensive bilingual corpus containing approximately 1.2 million tokens (620,000 in English and 580,000 in Uzbek) from information technology textbooks and educational resources was constructed. A carefully selected and manually annotated evaluation set of 5,000 lexical items from this corpus served as the basis for quantitative evaluation of the extraction and classification models. Terminology candidates were automatically extracted from the full corpus using a hybrid statistical and linguistic algorithm combining TF-IDF, C-value/NC-value, and morphosyntactic filtering. Evaluation on a carefully selected set showed high reliability: precision = 0.91, recall = 0.87, and F1 score = 0.89. A prototype bilingual matching tool using FastText subword embeddings / vectors and cosine similarity achieved an accuracy of 0.86 for matching English and Uzbek terms. In addition, a supervised XGBoost classifier was trained on the evaluation set to detect neologisms and metaphorical extensions of computer terminology, achieving an F1 score of 0.88. The proposed system provides a repeatable natural language processing pipeline for extracting, matching, and semantic classification of bilingual terminology. It serves as both a quantitative benchmark and a practical tool for integrating computer vocabulary into educational resources. This approach facilitates the systematic study of the structural and semantic dynamics of bilingual educational lexicons, thereby supporting future research and the development of practical curricula.
Educational Terminology Corpus Linguistics Natural Language Processing Terminology Extraction Bilingual Alignment Information Technology Education Neologism Detection
References
  1. K. T. Frantzi, S. Ananiadou, and H. Mima, “Automatic recognition of multi-word terms: The C-value/NC-value method,” International Journal on Digital Libraries, vol. 3, no. 2, pp. 115-130, 2000, [Online]. Available: https://doi.org/10.1007/s007999900023.
  2. P. Drouin, “Term extraction using non-technical corpora as a point of leverage,” Terminology, vol. 9, no. 1, pp. 99-115, 2003.
  3. G. Bordea, P. Buitelaar, S. Faralli, and R. Navigli, “SemEval-2015 Task 17: Taxonomy Extraction Evaluation (TExEval),” in Proc. of the 9th International Workshop on Semantic Evaluation (SemEval 2015), Denver, CO, USA, 2015, pp. 902-910, [Online]. Available: https://doi.org/10.18653/v1/S15-2151.
  4. D. Jurafsky and J. H. Martin, Speech and Language Processing, 3rd ed., draft. Stanford, CA, USA: Stanford University, 2023, [Online]. Available: https://web.stanford.edu/
  5. T. McEnery and A. Hardie, Corpus Linguistics: Method, Theory and Practice. Cambridge, U.K.: Cambridge University Press, 2012.
  6. M. M. Tilavova, “The process of word formation in modern English and the types of formation of words related to education,” Mental Enlightenment Scientific-Methodological Journal, vol. 4, no. 6, pp. 302-313, 2023, [Online]. Available: https://doi.org/10.37547/mesmj-V4-I6-42.
  7. J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. of NAACL-HLT 2019, Minneapolis, MN, USA, 2019, [Online]. Available: https://arxiv.org/abs/1810.04805.
  8. C. McAuley, J. Stewart, G. Siemens, and D. Cormier, The MOOC Model for Digital Practice, 2010, [Online]. Available: https://oerknowledgecloud.org/.
  9. I. Dagan, D. Glickman, and B. Magnini, “The PASCAL recognising textual entailment challenge,” in Proc. of the PASCAL Challenges Workshop on Recognising Textual Entailment, Southampton, U.K., 2005.
  10. P. Koehn, Statistical Machine Translation. Cambridge, U.K.: Cambridge University Press, 2010.
  11. T. Mikolov, I. Sutsukever, K. Chen, G. S. Corrado, and J. Dean, “Distributed representations of words and phrases and their compositionality,” in Advances in Neural Information Processing Systems 26 (NeurIPS 2013), 2013, pp. 3111-3119.


Proceedings of the International Conference on Applied Innovations in IT by Anhalt University of Applied Sciences is licensed under CC BY-SA 4.0
 ·  This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License

ICAIIT 2026
International Conference on Applied Innovation in IT
Navigation
Publisher
ISSN2199-8876
Location Anhalt University of Applied Sciences
Phone +49 (0) 3496 67 5611
Address Building 01, Room 425
Bernburger Str. 55
D-06366 Köthen, Germany
Open Access License

All works are licensed under the Creative Commons Attribution-ShareAlike 4.0 International License (CC BY-SA 4.0), unless otherwise noted.

Published by ICAIIT in cooperation with Anhalt University of Applied Sciences.

© 2026 ICAIIT — International Conference on Applied Innovations in IT. Anhalt University of Applied Sciences, Köthen, Germany.
Visitors: site traffic counter