Multimodal Emotion Recognition is also vital to identifying complicated human emotions in nature where unimodal systems are mostly ineffective. The paper will create a strong deep learning model which is aimed to enhance the accuracy of recognition by optimizing preprocessing and using dual-branch architecture. Two stream architecture is used, where Convolutional Neural Networks are used to extract facial expression feature and a late fusion technique used to merge body posture and scene context. The preprocessing includes ROI localization, stream-specific normalization as well as color jittering to overcome the issue of lighting changes, and sensor noise. To solve the label imbalance in discrete classification, Focal Loss is applied, and to solve continuous Valence, Arousal, and Dominance regression with mean squared error, Mean Squared Error is used. Averaged Precision (mAP) and accuracy of the model of 0.3614 and 0.8448 respectively on EMOTIC data achieves better results compared to traditional baselines and models, including EmotiCon. These findings suggest the practicality of independent feature extraction as a means of avoiding inter-modal interference and present a system robust to noisy real-world data, in essence learning to transform linguistic and non-linguistic features into one overall impression of emotion.
R. Kosti, J. M. Alvarez, A. Recasens, and A. Lapedriza, “Context Based Emotion Recognition using EMOTIC Dataset,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 42, no. 11, pp. 2755-2766, Nov. 2020.
N. Wagner et al., “CAGE: Circumplex Affect Guided Expression Inference,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, pp. 4683-4692, Jun. 2024.
F. Zhou et al., “Fine-grained facial expression analysis using deep learning with focus on environmental variance,” Machine Vision and Applications, vol. 32, no. 6, pp. 124-138, Oct. 2021.
X. Li et al., “Single-stage Emotion Recognition with Decoupled Transformer Architecture,” arXiv preprint, Feb. 2024, [Online]. Available: https://arxiv.org/abs/2402.13850.
Z. Chen and J. Qin, “Focal Loss for Multi-label Emotion Classification,” IEEE Access, vol. 9, pp. 114681-114691, Aug. 2021.
S. Tamhane, A. Shrirao, M. Shah, and D. Patil, “Emotion Recognition Using Deep Convolutional Neural Networks,” SSRN Electronic Journal, Jan. 2022, [Online]. Available: https://doi.org/10.2139/ssrn.4091264.
F. Limami et al., “Contextual emotion detection in images: A comprehensive survey,” International Journal of Computer Vision, vol. 133, no. 2, pp. 10202-10224, Feb. 2025.
B. T. Atmaja and A. Sasou, “Valence and Arousal Recognition from Body Gestures using Multimodal Streams,” IEEE Access, vol. 12, pp. 10518-10530, Apr. 2024.
R. Shankar et al., “Reducing Interference in Early Multimodal Fusion via Independent Extraction,” Procedia Computer Science, vol. 235, pp. 335-344, Jul. 2024.
J. Ede, “Batch Normalization for Activation Normalization in Deep Emotion Networks,” Computers & Electrical Engineering, vol. 115, pp. 109410-109425, Jan. 2025.
S. Poria, D. Hazarika, N. Majumder, and R. Mihalcea, “Beneath the tip of the iceberg: Current challenges and new directions in sentiment analysis research,” Information Fusion, vol. 62, pp. 14-34, Oct. 2020.
J. Wagner, J. Kim, and E. André, “From Physiological Signals to Emotions: Implementing and Comparing Multistage Algorithms for Online Recognition,” IEEE Transactions on Affective Computing, vol. 11, no. 4, pp. 605-618, Oct. 2020.
A. Mittal, K. Bhattacharya, R. Shah, and A. Majumder, “EmotiCon: Context-Aware Multimodal Emotion Recognition Using Frege’s Principle,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 14241-14250, Jun. 2020.
V. Lohani et al., “Multimodal Emotion Recognition in the Wild: Challenges and Solutions,” IEEE Transactions on Affective Computing, vol. 13, no. 4, pp. 1968-1983, Dec. 2022.
Y. Gan et al., “Multimodal Emotion Recognition with Capturing Semantic Relationships,” Neural Computing and Applications, vol. 35, no. 12, pp. 9040-9055, Aug. 2023.
T. Zhang et al., “Context-Aware Emotion Recognition via Dual-Branch Networks and Late Fusion,” IEEE Transactions on Multimedia, vol. 25, pp. 8251-8263, Oct. 2023.
H. Liu et al., “Self-supervised Learning for Multimodal Emotion Recognition using Masked Autoencoders,” in Proceedings of the International Conference on Machine Learning (ICML), vol. 139, pp. 1-10, Jul. 2021.
Q. Xie et al., “Self-Training With Noisy Student Improves ImageNet Classification,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 10687-10698, Jun. 2020.
X. Li et al., “Generalized Focal Loss V2: Learning Reliable Localization Quality Estimation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 11632-11641, Jun. 2021.
S. Ruder, “An Overview of Multi-Task Learning in Deep Neural Networks,” Neural Computing and Applications, vol. 32, no. 4, pp. 5616-5630, Jan. 2020.
G. Zhao et al., “Facial Expression Recognition from Near-Infrared Video in Variable Lighting,” Journal of Affective Disorders, vol. 350, pp. 503-515, Feb. 2025.
S. Wang and Q. Ji, “A Survey on Emotional State Analysis from Body Posture,” in Advances in Computer Science and Engineering, Springer, Jan. 2024, pp. 185-205.
E. Ryumina, D. Ryumin, A. Axyonov, D. Ivanko, and A. Karpov, “Multi-Corpus Emotion Recognition Method Based on Cross-Modal Gated Attention Fusion,” Pattern Recognition Letters, vol. 190, pp. 192-200, 2025, [Online]. Available: https://doi.org/10.1016/j.patrec.2025.02.024.
W. de Lima Costa et al., “High-Level Context Representation for Emotion Recognition in Images,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, pp. 326-334, Jun. 2023.
M. Shafiq and Z. Gu, “Deep Residual Learning for Image Recognition: A Survey,” Applied Sciences, vol. 12, no. 18, p. 8972, Sep. 2022.
R. Taneja, J. Singh, and R. Gill, “Multimodal Emotion Recognition System Using Machine Learning and Psychological Signals: A Review,” in Soft Computing: Theories and Applications, Singapore: Springer, 2022, pp. 657-666, [Online]. Available: https://doi.org/10.1007/978-981-16-1740-9_54.
M. Liu et al., “Multi-label Learning with Deep Forest for Emotion Recognition,” Journal of Visual Communication and Image Representation, vol. 102, pp. 103305-103321, Jan. 2025.
D. Kollias and S. Zafeiriou, “ABAW: Valence-Arousal Estimation and Expression Recognition in the Wild,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 2486-2495, Jun. 2022.
E. Probierz, “Context-Aware Approach by Social Robots to Emotion Recognition in Uncontrolled Settings,” Applied Sciences, vol. 13, no. 3, pp. 1097-1112, Jan. 2023.
E. Probierz, “On emotion detection and recognition using a context-aware approach by social robots: modification of Faster R-CNN and YOLO v3 neural networks,” European Research Studies Journal, vol. 26, no. 1, pp. 572-585, 2023.