Proceedings of International Conference on Applied Innovation in IT  ·  2026/06/12  ·  Vol. 14  ·  Issue 3  ·  pp. 503–520
Automated System for Objective Assessment of Educational Achievements Using Generative Artificial Intelligence
Zilola Qurbonova, Feruzakhon Ramazonova, Shokhsanam Odiljonova, Gulrux Alimova, Kanybek Isakov, Sanjarbek Jamoldinov, Vazira Jalalova, Sitoraxon Rakhmanbekova, Nigora Abdurakhimova, Charoz Orziqulova, Zilola Abidova and Dildora Baxadirova
In the context of the systemic digitalization of education, Automated Scoring Systems (ASS) based on Generative Artificial Intelligence (Generative AI) are becoming an important instrument for scalable and objective assessment of educational achievements. Despite the high efficiency and operational flexibility of such systems, the fundamental psychometric issues of their reliability and validity remain insufficiently explored. Of particular importance is the transition from Classical Test Theory (CTT) to models of modern measurement theory, specifically Item Response Theory (IRT), and particularly the Rasch model, which ensures parameter invariance and metric comparability of measurements. The purpose of this study is to conduct a comprehensive evaluation of the reliability and validity of an automated scoring system based on a generative language model using the Rasch model as an analytical framework. The study applies the one-parameter logistic Rasch model to analyze item characteristics, the distribution of examinee abilities, fit statistics, separation indices, and the detection of Differential Item Functioning (DIF). The methodology involves simulation of a dataset consisting of 1,200 students and 40 open-ended items automatically evaluated by a generative model. The analysis is performed using maximum likelihood estimation procedures and the calculation of Infit and Outfit statistics. In addition, convergent and criterion validity indicators are assessed through comparison between automated scores and expert evaluations. The obtained results demonstrate high system reliability (Person Reliability > 0.88), satisfactory item fit to the Rasch model, and the absence of systematic bias related to gender or level of academic preparation. The theoretical and practical implications of integrating generative AI into high-stakes assessment systems are discussed. This study contributes to the advancement of psychometric evaluation of AI-based assessment systems and provides a scientific foundation for their further standardization.
Generative Artificial Intelligence Automated Assessment Rasch Model Reliability Validity Item Response Theory Differential Item Functioning Educational Psychometrics.
References
  1. G. Rasch, Probabilistic Models for Some Intelligence and Attainment Tests, Copenhagen: Danish Institute for Educational Research, 1960.
  2. B. D. Wright and G. N. Masters, Rating Scale Analysis, Chicago: MESA Press, 1982.
  3. T. G. Bond and C. M. Fox, Applying the Rasch Model: Fundamental Measurement in the Human Sciences, 3rd ed., New York: Routledge, 2015.
  4. S. E. Embretson and S. P. Reise, Item Response Theory for Psychologists, Mahwah, NJ: Lawrence Erlbaum Associates, 2000.
  5. F. M. Lord, Applications of Item Response Theory to Practical Testing Problems, Hillsdale, NJ: Erlbaum, 1980.
  6. R. K. Hambleton, H. Swaminathan, and H. J. Rogers, Fundamentals of Item Response Theory, Newbury Park, CA: Sage Publications, 1991.
  7. R. J. de Ayala, The Theory and Practice of Item Response Theory, New York: Guilford Press, 2009.
  8. F. B. Baker and S.-H. Kim, The Basics of Item Response Theory Using R, New York: Springer, 2017.
  9. W. J. Boone, J. R. Staver, and M. S. Yale, Rasch Analysis in the Human Sciences, Dordrecht: Springer, 2014.
  10. J. M. Linacre, “What do Infit and Outfit, Mean-square and Standardized Mean?,” Rasch Measurement Transactions, vol. 16, no. 2, p. 878, 2002.
  11. A. Tennant and P. G. Conaghan, “The Rasch measurement model in rheumatology: What is it and why use it?,” Arthritis Care & Research, vol. 57, no. 8, pp. 1358-1362, 2007, [Online]. Available: https://doi.org/10.1002/art.23108.
  12. M. T. Kane, “Validating the Interpretations and Uses of Test Scores,” Journal of Educational Measurement, vol. 50, no. 1, pp. 1-73, 2013.
  13. W. J. van der Linden, Ed., Handbook of Item Response Theory, Boca Raton: CRC Press, 2016.
  14. T. Eckes, Introduction to Many-Facet Rasch Measurement, Frankfurt am Main: Peter Lang, 2015.
  15. American Educational Research Association (AERA), American Psychological Association (APA), and National Council on Measurement in Education (NCME), Standards for Educational and Psychological Testing, Washington, DC: AERA, 2014.
  16. Y. Attali and J. Burstein, “Automated essay scoring with e-rater V.2,” The Journal of Technology, Learning and Assessment, vol. 4, no. 3, pp. 1-30, 2006.
  17. D. M. Williamson, R. J. Mislevy, and I. I. Bejar, Automated Scoring of Complex Tasks in Computer-Based Testing, Mahwah, NJ: Lawrence Erlbaum Associates, 2006.
  18. S. P. Reise and N. G. Waller, “Item response theory and clinical measurement,” Annual Review of Clinical Psychology, vol. 5, pp. 27-48, 2009.
  19. M. D. Shermis and J. Burstein, Handbook of Automated Essay Evaluation, New York: Routledge, 2013.
  20. S. Dikli, “An overview of automated scoring of essays,” The Journal of Technology, Learning and Assessment, vol. 5, no. 1, pp. 1-35, 2006.
  21. B. D. Zumbo, A Handbook on the Theory and Methods of Differential Item Functioning, Ottawa: Directorate of Human Resources Research and Evaluation, Department of National Defense, 1999.
  22. G. Camilli and L. A. Shepard, Methods for Identifying Biased Test Items, Thousand Oaks, CA: Sage Publications, 1994.
  23. M. J. Kolen and R. L. Brennan, Test Equating, Scaling, and Linking, 3rd ed., New York: Springer, 2014.
  24. I. I. Bejar, “Automated scoring and validity,” in D. M. Williamson et al., Eds., Automated Scoring of Complex Tasks in Computer-Based Testing, Routledge, 2012.
  25. R. D. Penfield and G. Camilli, “Differential item functioning and item bias,” in C. R. Rao and S. Sinharay, Eds., Handbook of Statistics, vol. 26, pp. 125-167, 2007.
  26. D. Thissen and L. Steinberg, “A taxonomy of item response models,” Psychometrika, vol. 51, no. 4, pp. 567-577, 1986.
  27. P. W. Holland and D. T. Thayer, “Differential item performance and the Mantel-Haenszel procedure,” in H. Wainer and H. I. Braun, Eds., Test Validity, Hillsdale, NJ: Erlbaum, 1988, pp. 129-145.
  28. L. M. Rudner and T. Liang, “Automated essay scoring using Bayes’ theorem,” The Journal of Technology, Learning and Assessment, vol. 1, no. 2, pp. 1-21, 2002.
  29. L. Chen, P. Chen, and Z. Lin, “Artificial intelligence in education: A review,” IEEE Access, vol. 8, pp. 75264-75278, 2020, [Online]. Available: https://doi.org/10.1109/ACCESS.2020.2988510.
  30. B. E. Clauser and K. M. Mazor, “Using statistical procedures to identify differential item functioning test items,” Educational Measurement: Issues and Practice, vol. 17, no. 1, pp. 31-44, 1998.
  31. E. W. Wolfe and E. V. Smith, “Instrument development tools and activities for measure validation using Rasch models,” Journal of Applied Measurement, vol. 8, no. 2, pp. 204-234, 2007.


Proceedings of the International Conference on Applied Innovations in IT by Anhalt University of Applied Sciences is licensed under CC BY-SA 4.0
 ·  This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License

ICAIIT 2026
International Conference on Applied Innovation in IT
Navigation
Publisher
ISSN2199-8876
Location Anhalt University of Applied Sciences
Phone +49 (0) 3496 67 5611
Address Building 01, Room 425
Bernburger Str. 55
D-06366 Köthen, Germany
Open Access License

All works are licensed under the Creative Commons Attribution-ShareAlike 4.0 International License (CC BY-SA 4.0), unless otherwise noted.

Published by ICAIIT in cooperation with Anhalt University of Applied Sciences.

© 2026 ICAIIT — International Conference on Applied Innovations in IT. Anhalt University of Applied Sciences, Köthen, Germany.
Visitors: site traffic counter