Recent advancements in computer vision have led to increased interest in complex video analysis tasks, particularly in understanding and predicting content. Recognizing human actions and activities is crucial in everyday life, as it allows for the collection of comprehensive and accurate information about human behavior through wearable or stationary devices. Numerous machine learning and deep learning techniques have been investigated in human activity recognition in order to categorize human activities. This paper presents the impact of deep learning architectures for human action and activity recognition. It outlines the structure, performance, and limitations of the deep learning techniques used. These findings are intended to resolve limitations found in previous research, paving the way for enhanced models with greater accuracy across handling varied datasets. The evaluation analysis showed that action recognition achieved an average accuracy of 89.2%, while activity recognition achieved an average accuracy of 93.9%. Furthermore, the combination of ResNets and ConvLSTM algorithms on the SKIG dataset outperformed other methods in accuracy, reaching 99.72%. Discovering and extracting semantic features and the complexity of activities are key challenges facing human activity recognition systems.
R. Gómez-Ramos, J. Duque-Domingo, E. Zalama and J. Gómez-García-Bermejo, “An unsupervised method to recognise human activity at home using non-intrusive sensors,” Electronics, vol. 12, no. 23, pp. 1-24, Nov. 2023, [Online]. Available: https://doi.org/10.3390/electronics12234772.
F. Demrozi, G. Pravadelli, A. Bihorac and P. Rashidi, “Human activity recognition using inertial, physiological and environmental sensors: A comprehensive survey,” IEEE Access, vol. 8, pp. 210816-210841, 2020.
N. Sedaghati, S. Ardebili and A. Ghaffari, “Application of human activity/action recognition: a review,” Multimed. Tools Appl., Jun. 2025, [Online]. Available: https://doi.org/10.1007/s11042-024-20576-2.
D. R. Beddiar, B. Nini, M. Sabokrou and A. Hadid, “Vision-based human activity recognition: A survey,” Multimed. Tools Appl., vol. 79, pp. 30509-30555, 2020, [Online]. Available: https://doi.org/10.1007/s11042-020-09004-3.
H. Khalaf and M. Riyadh, “Human activity recognition using inertial sensors in a smartphone: technical background (review),” Al-Nahrain J. Sci., vol. 27, no. 1, pp. 108-120, Mar. 2024, [Online]. Available: https://doi.org/10.22401/ANJS.27.1.10.
N. S. Khan and M. S. Ghani, “A survey of deep learning based models for human activity recognition,” Wirel. Pers. Commun., vol. 121, pp. 3099-3133, 2021, [Online]. Available: https://doi.org/10.1007/s11277-021-08525-w.
H. Yin, R. O. Sinnott and G. T. Jayaputera, “A survey of video-based human action recognition in team sports,” Artif. Intell. Rev., vol. 57, pp. 293-348, Sep. 2024, [Online]. Available: https://doi.org/10.1007/s10462-024-10934-9.
V. Sharma, M. Gupta, A. Kumar and D. Mishra, “Video processing using deep learning techniques: a systematic literature review,” IEEE Access, vol. 9, pp. 139489-139504, Oct. 2021, [Online]. Available: https://doi.org/10.1109/ACCESS.2021.3118541.
F. Shafizadegan, A. R. Naghsh-Nilchi and E. Shabaninia, “Multimodal vision-based human action recognition using deep learning: a review,” Artif. Intell. Rev., vol. 57, Art. no. 178, Jun. 2024, [Online]. Available: https://doi.org/10.1007/s10462-024-10730-5.
J. Wang, Y. Chen, S. Hao, X. Peng and L. Hu, “Deep learning for sensor-based activity recognition: a survey,” Pattern Recognit. Lett., vol. 119, pp. 3-11, Mar. 2019, doi: 10.1016/j.patrec.2018.02.010.
X. Chai, B. G. Lee, C. Hu, M. Pike, D. Chieng, R. Wu and W.-Y. Chung, “IoT-FAR: a multi-sensor fusion approach for IoT-based firefighting activity recognition,” Inf. Fusion, vol. 113, p. 102650, Aug. 2025, [Online]. Available: https://doi.org/10.1016/j.inffus.2024.102650.
T. F. Naik Bukht, H. Rahman, M. Shaheen, A. Algarni, N. A. Almujally and A. Jalal, “A review of video-based human activity recognition: theory, methods and applications,” Multimed. Tools Appl., Oct. 2024, [Online]. Available: https://doi.org/10.1007/s11042-024-19711-w.
X. Cheng, L. Zhang, Y. Tang, Y. Liu, H. Wu and J. He, “Real-time human activity recognition using conditionally parametrized convolutions on mobile and wearable devices,” IEEE Sens. J., vol. 21, no. 15, pp. 16661-16673, Aug. 2021.
E. Hato, Z. S. Abduljabbar and Z. J. Ahmed, “Comparative analysis for bag of features (BoF) performance,” Iraqi J. Sci., vol. 65, no. 8, pp. 4606-4622, Aug. 2024, [Online]. Available: https://doi.org/10.24996/ijs.2024.65.8.38.
S. M. Al-Selwi, M. F. Hassan, S. J. Abdulkadir, A. Muneer, E. H. Sumiea, A. Alqushaibi and M. G. Ragab, “RNN-LSTM: from applications to modeling techniques and beyond—systematic review,” J. King Saud Univ. Comput. Inf. Sci., vol. 36, p. 102068, 2024, [Online]. Available: https://doi.org/10.1016/j.jksuci.2024.102068.
Elboushaki, R. Hannane, K. Afdel and L. Koutti, “MultiD-CNN: a multi-dimensional feature learning approach based on deep convolutional networks for gesture recognition in RGB-D image sequences,” Expert Syst. Appl., vol. 139, p. 112829, Mar. 2020.
S. K. Yadav, A. Singh, A. Gupta and J. L. Raheja, “Real-time yoga recognition using deep learning,” Neural Comput. Appl., vol. 33, no. 16, pp. 13243-13256, 2021.
F. Xiao, L. Pei, L. Chu, D. Zou, W. Yu, Y. Zhu and T. Li, “A deep learning method for complex human activity recognition using virtual wearable sensors,” in Proc. Int. Conf. Hum.-Centric Comput. (HCC), 2020, pp. 455-467.
P. Agarwal and M. Alam, “A lightweight deep learning model for human activity recognition on edge devices,” Procedia Comput. Sci., vol. 167, pp. 2364-2373, 2020.
N. Jaouedi, N. Boujnah and M. S. Bouhlel, “A new hybrid deep learning model for human action recognition,” J. King Saud Univ. Comput. Inf. Sci., vol. 32, no. 4, pp. 447-453, 2020.
Z. N. Khan and J. Ahmad, “Attention-induced multi-head convolutional neural network for human activity recognition,” Appl. Soft Comput., vol. 110, p. 107671, Oct. 2021.
N. Rashid, B. U. Demirel and M. A. Al Faruque, “AHAR: adaptive CNN for energy-efficient human activity recognition in low-power edge devices,” IEEE Internet Things J., vol. 9, no. 15, pp. 13000-13012, Aug. 2022.
Y. Yang, H. Chen, Z. Liu, Y. Ly, B. Zhang, S. Wu, Z. Wang and K. Ren, “Action recognition with multi-stream motion modeling and mutual information maximization,” in Proc. Int. Joint Conf. Artif. Intell. (IJCAI), 2023, pp. 1658-1664.
C. Yang, F. Mei, T. Zang, J. Tu, N. Jiang and L. Liu, “Human action recognition using key-frame attention-based LSTM networks,” Electronics, vol. 12, no. 12, p. 2622, Jun. 2023.
S. Yosry, L. Elrefaei, R. ElKamaar and R. R. Ziedan, “Various frameworks for integrating image and video streams for spatiotemporal information learning employing 2D-3D residual networks for human action recognition,” Discov. Appl. Sci., vol. 6, no. 3, p. 141, Mar. 2024.
M. Joudaki, M. Imani and H. R. Arabnia, “A new efficient hybrid technique for human action recognition using 2D Conv-RBM and LSTM with optimized frame selection,” Technologies, vol. 13, no. 2, p. 53, Feb. 2025.
E. López-Lozada, H. Sossa, E. Rubio-Espino and J. Y. Montiel-Pérez, “Action recognition in videos through a transfer-learning-based technique,” Mathematics, vol. 12, no. 20, p. 3245, Oct. 2024.
H. Chen, Y. Pan and C. Wang, “An optimization method of human skeleton keyframes selection for action recognition,” Complex Intell. Syst., vol. 10, no. 3, pp. 4659-4673, Jun. 2024.
D. R. Rani and J. P. C., “Vision transformer-based model for human action recognition in still images,” J. Comput. Anal. Appl., vol. 33, no. 8, pp. 522-531, 2024.
D. Lamani, P. Kumar, A. Bhagyalakshmi, J. M. Shanthi, L. P. Maguluri, M. Arif, C. Dhanamjayulu, S. Kumar and B. Khan, “SVM directed machine learning classifier for human action recognition network,” Sci. Rep., vol. 15, no. 1, p. 672, Jan. 2025.