International Journal of Innovative Research in Computer and Communication Engineering

ISSN Approved Journal | Impact factor: 8.771 | ESTD: 2013 | Follows UGC CARE Journal Norms and Guidelines

| Monthly, Peer-Reviewed, Refereed, Scholarly, Multidisciplinary and Open Access Journal | High Impact Factor 8.771 (Calculated by Google Scholar and Semantic Scholar | AI-Powered Research Tool | Indexing in all Major Database & Metadata, Citation Generator | Digital Object Identifier (DOI) |


TITLE Interpretable Temporal Attention-LSTM for Skeleton-Based Human Activity Recognition
ABSTRACT Human Activity Recognition (HAR) has become an important research area due to its wide range of applications in intelligent surveillance, healthcare monitoring, assisted living, sports analytics, and human-computer interaction. Among various HAR approaches, skeleton-based activity recognition has emerged as an effective solution because skeletal representations focus on human body movements while reducing the influence of background noise, illumination changes, and appearance variations. However, accurately capturing temporal motion patterns and understanding the reasoning behind model predictions remain challenging tasks. This paper proposes an Interpretable Temporal Attention-LSTM framework for skeleton-based human activity recognition. The proposed method utilizes a Long Short-Term Memory (LSTM) network to learn temporal dependencies from sequential skeletal joint data and incorporates a temporal attention mechanism to automatically identify and emphasize the most informative frames within an activity sequence. By assigning adaptive attention weights to different time steps, the model focuses on critical motion segments that contribute most to activity classification. A spatial attention module further weighs the contribution of individual joints, and together the two mechanisms provide a transparent view of the model's decision-making process. The effectiveness of the proposed framework is evaluated on the NTU RGB+D benchmark using standard evaluation metrics, including accuracy, F1-score, and top-k accuracy. Experimental results show that the model achieves an overall accuracy of 92.1% and an F1-score of 91.8%, with ablation studies confirming the contribution of the attention mechanism and the benefit of graph-based spatial encoding. The integration of temporal attention and explainability enhances both classification performance and transparency, making the proposed framework a practical and trustworthy solution for real-world human activity recognition systems.
AUTHOR P. ROJA MALLI, CH.HARIKA Department of Computer Science and Engineering, St. Mary's Women's Engineering College, Guntur, Andhra Pradesh, India
VOLUME 186
DOI DOI: 10.15680/IJIRCCE.2026.1407032
PDF pdf/32_Interpretable Temporal Attention-LSTM for Skeleton-Based Human Activity Recognition.pdf
KEYWORDS
References [1] K. Xu, D. Cheng, X. Zhu, L. Zhao, Y. Chen, and T. Tan, “Multi-modal Attention-Based Human Activity Recognition Using Dynamic Glimpses and Skeleton Fusion,” IEEE Transactions on Image Processing, vol. 30, pp. 6467–6480, 2021.
[2] C. Cao, Y. Zhang, C. Zhang, Y. Yu, and H. Lu, “Reinforced Temporal Attention and Split-Rate Transfer for Skeleton-Based Action Recognition,” Proc. IEEE/CVF CVPR, 2019, pp. 10234–10243.
[3] Z. Zhang, P. Lan, J. Luo, and Y. Liu, “Context-Aware Attention Network for Skeleton-Based Action Recognition,” IEEE Transactions on Image Processing, vol. 31, pp. 2065–2079, 2022.
[4] Z. Li, Y. Wu, and Y. Tian, “Actional-Structural Graph Convolutional Networks for Skeleton-Based Action Recognition,” Proc. IEEE/CVF CVPR, 2019, pp. 3595–3603.
[5] D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri, “A Closer Look at Spatiotemporal Convolutions for Action Recognition,” Proc. IEEE CVPR, 2018, pp. 6450–6459.
[6] L. Shi, Y. Zhang, J. Cheng, and H. Lu, “Two-Stream Adaptive Graph Convolutional Networks for Skeleton-Based Action Recognition,” Proc. IEEE/CVF CVPR, 2019, pp. 12026–12035.
[7] W. Zhang, P. Lan, Z. Liu, and H. Cheng, “Graph-Based High-Order Relation Modeling for Skeleton-Based Action Recognition,” IEEE Transactions on Image Processing, vol. 32, pp. 1580–1594, 2023.
[8] Y. Du, W. Wang, and L. Wang, “Representation Learning of Temporal Dynamics for Skeleton-Based Action Recognition,” IEEE Transactions on Image Processing, vol. 25, no. 7, pp. 3010–3023, 2016.
[9] Y. Song, Z. Zhang, C. Shan, and L. Wang, “Stronger, Faster and More Explainable: A Graph Convolutional Baseline for Skeleton-Based Action Recognition,” Proc. ACM Multimedia, 2022, pp. 844–852.
[10] P. Wang, C. Shen, A. van den Hengel, and P. Torr, “Graph-Based 3D Skeleton Embedding for Action Recognition with Hierarchical Pooling,” IEEE Transactions on Pattern Analysis and Machine Intelligence, early access, 2024.
Copyright © IJIRCCE 2020.All right reserved