In the context of the rapid advancement of digital educational technologies, the personalization of online learning based on learners’ emotional states has become a critical research challenge. This study proposes a multimodal emotion recognition system based on the integration of visual, auditory, and textual data using deep learning techniques. The objective of this research is to develop and evaluate the effectiveness of a multimodal architecture that combines Convolutional Neural Networks (CNN), Recurrent Neural Networks (LSTM), and transformer-based attention mechanisms to improve emotion recognition accuracy in educational environments. The experimental evaluation is conducted using widely recognized datasets, including FER2013, AffectNet, and RAVDESS. The proposed model demonstrates strong performance, achieving an accuracy of 0.89 and an F1-score of 0.86, significantly outperforming unimodal approaches. The results indicate that the integration of multiple modalities produces a synergistic effect, enhancing system robustness against noise and incomplete data. Additional experiments confirm high discriminative capability (AUC > 0.9) and model stability under conditions of missing modalities. The practical significance of this study lies in the potential application of the proposed system within adaptive educational platforms, enabling dynamic content adaptation based on learners’ emotional states and thereby improving learning effectiveness. Overall, the findings confirm that multimodal artificial intelligence represents a promising approach for the development of intelligent online learning systems and adaptive pedagogy. The proposed architecture demonstrates high accuracy, robustness, and practical applicability, highlighting the effectiveness of multimodal approaches in emotion recognition within digital learning environments.
P. Ekman, “An argument for basic emotions,” Cognition and Emotion, vol. 6, no. 3-4, pp. 169-200, 1992, [Online]. Available: https://doi.org/10.1080/02699939208411068.
R. W. Picard, Affective Computing, Cambridge, MA: MIT Press, 1997.
Z. Zhang, J. Han, J. Deng, and et al., “Multimodal emotion recognition: A review of recent advances,” Information Fusion, vol. 59, pp. 103-126, 2020.
R. A. Calvo and S. D’Mello, “Affect detection: An interdisciplinary review of models, methods, and their applications,” IEEE Transactions on Affective Computing, vol. 1, no. 1, pp. 18-37, 2010, [Online]. Available: https://doi.org/10.1109/T-AFFC.2010.1.
S. Poria, E. Cambria, R. Bajpai, and A. Hussain, “A review of affective computing: From unimodal analysis to multimodal fusion,” Information Fusion, vol. 37, pp. 98-125, 2017.
I. J. Goodfellow, D. Erhan, P. L. Carrier, and et al., “Challenges in representation learning: A report on three machine learning contests,” in Neural Information Processing, pp. 117-124, 2013.
K. Cho, B. van Merriënboer, C. Gulcehre, and et al., “Learning phrase representations using RNN encoder-decoder for statistical machine translation,” in Proceedings of EMNLP, pp. 1724-1734, 2014.
A. Mollahosseini, B. Hasani, and M. H. Mahoor, “AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild,” IEEE Transactions on Affective Computing, vol. 10, no. 1, pp. 18-31, 2019, [Online]. Available: https://doi.org/10.1109/TAFFC.2017.2740923.
A. Zadeh, M. Chen, S. Poria, E. Cambria, and L.-P. Morency, “Tensor fusion network for multimodal sentiment analysis,” in Proceedings of EMNLP, pp. 1103-1114, 2017.
C. Busso, M. Bulut, C.-C. Lee, and et al., “IEMOCAP: Interactive emotional dyadic motion capture database,” Language Resources and Evaluation, vol. 42, pp. 335-359, 2008.
S. K. D’Mello and J. Kory, “A review and meta-analysis of multimodal affect detection systems,” ACM Computing Surveys, vol. 47, no. 3, Art. no. 43, 2015.
M. Soleymani, D. Garcia, B. Jou, B. Schuller, S.-F. Chang, and M. Pantic, “A survey of multimodal sentiment analysis,” Image and Vision Computing, vol. 65, pp. 3-14, 2017.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems, vol. 25, pp. 1097-1105, 2012.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation, vol. 9, no. 8, pp. 1735-1780, 1997, [Online]. Available: https://doi.org/10.1162/neco.1997.9.8.1735.
A. Vaswani, N. Shazeer, N. Parmar, and et al., “Attention is all you need,” in Advances in Neural Information Processing Systems, vol. 30, 2017.
T. Baltrušaitis, C. Ahuja, and L.-P. Morency, “Multimodal machine learning: A survey and taxonomy,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 41, no. 2, pp. 423-443, 2019, [Online]. Available: https://doi.org/10.1109/TPAMI.2018.2798607.
R. Pekrun, “The control-value theory of achievement emotions,” Educational Psychology Review, vol. 18, no. 4, pp. 315-341, 2006, [Online]. Available: https://doi.org/10.1007/s10648-006-9029-9.
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever, “Improving language understanding by generative pre-training,” OpenAI Technical Report, 2018.
S. Li and W. Deng, “Deep facial expression recognition: A survey,” IEEE Transactions on Affective Computing, vol. 13, no. 3, pp. 1195-1215, 2022.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Proceedings of ICLR, 2015.
M. Chen, S. Wang, P. P. Liang, T. Baltrušaitis, A. Zadeh, and L.-P. Morency, “Multimodal sentiment analysis with word-level fusion and reinforcement learning,” in Proceedings of ICMI, pp. 163-171, 2017.
S. Poria, E. Cambria, R. Bajpai, and A. Hussain, “A review of affective computing: From unimodal analysis to multimodal fusion,” Information Fusion, vol. 37, pp. 98-125, 2017.
S. Koelstra, M. Mühl, M. Soleymani, and et al., “DEAP: A database for emotion analysis using physiological signals,” IEEE Transactions on Affective Computing, vol. 3, no. 1, pp. 18-31, 2012.
B. Schuller, A. Batliner, S. Steidl, and D. Seppi, “Recognising realistic emotions and affect in speech,” IEEE Signal Processing Magazine, vol. 28, no. 6, pp. 93-105, 2011.
J. Han, Z. Zhang, Z. Ren, and B. Schuller, “Implicit fusion by joint audiovisual training for emotion recognition in mono modality,” in Proceedings of ICASSP, pp. 5869-5873, 2018.
J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y. Ng, “Multimodal deep learning,” in Proceedings of ICML, pp. 689-696, 2011.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: A simple way to prevent neural networks from overfitting,” Journal of Machine Learning Research, vol. 15, pp. 1929-1958, 2014.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of CVPR, pp. 770-778, 2016.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of NAACL-HLT, pp. 4171-4186, 2019.
Y.-H. H. Tsai, S. Bai, P. P. Liang, J. Z. Kolter, L.-P. Morency, and R. Salakhutdinov, “Multimodal Transformer for unaligned multimodal language sequences,” in Proceedings of ACL, pp. 6558-6569, 2019.
H. Gunes and B. Schuller, “Categorical and dimensional affect analysis in continuous input: Current trends and future directions,” Image and Vision Computing, vol. 31, no. 2, pp. 120-136, 2013.
M. Soleymani, J. Lichtenauer, T. Pun, and M. Pantic, “A multimodal database for affect recognition and implicit tagging,” IEEE Transactions on Affective Computing, vol. 3, no. 1, pp. 42-55, 2012.