International Journal of Innovative Research in Computer and Communication Engineering

ISSN Approved Journal | Impact factor: 8.771 | ESTD: 2013 | Follows UGC CARE Journal Norms and Guidelines

| Monthly, Peer-Reviewed, Refereed, Scholarly, Multidisciplinary and Open Access Journal | High Impact Factor 8.771 (Calculated by Google Scholar and Semantic Scholar | AI-Powered Research Tool | Indexing in all Major Database & Metadata, Citation Generator | Digital Object Identifier (DOI) |


TITLE Speech Emotion Recognition Using MFCC, Autoencoders, And CNN-LSTM Networks
ABSTRACT Speech Emotion Recognition (SER) is an emerging area of artificial intelligence and human-computer interaction that seeks to automatically identify a speaker's emotional state from speech signals. Reliable SER is complicated by speaker variability, environmental noise, accent differences and overlapping emotional cues, all of which limit the accuracy of conventional recognition pipelines. This paper presents a hybrid deep learning framework for SER that combines Mel-Frequency Cepstral Coefficients (MFCC) for perceptually relevant feature extraction, an Autoencoder for unsupervised dimensionality reduction of the MFCC feature space, and a Convolutional Neural Network-Long Short-Term Memory (CNN-LSTM) network for classification. Speech signals are first pre-processed to remove noise and normalise amplitude; MFCC features are then extracted and compressed by the Autoencoder into a compact latent representation, which is passed to the CNN-LSTM classifier to recognise seven emotion classes - happy, sad, angry, fear, surprise, disgust and neutral. The framework is evaluated on a publicly available emotional speech corpus using accuracy, precision, recall and F1-score, and is benchmarked against SVM, KNN, Random Forest, CNN and LSTM baselines. Experimental results show that the proposed MFCC-Autoencoder-CNN-LSTM pipeline consistently outperforms these baselines, confirming the benefit of combining perceptual feature extraction, unsupervised feature optimisation and joint spatio-temporal modelling for robust emotion recognition.
AUTHOR CH. KIRANMAI, DR. Y. DAVID SOLOMON RAJU Department of Computer Science and Engineering, St. Mary's Women's Engineering College, Budampadu, Guntur, Andhra Pradesh, India Principal & Professor, Department of Computer Science and Engineering, St. Mary's Women's Engineering College, Budampadu, Guntur, Andhra Pradesh, India
VOLUME 187
DOI DOI: 10.15680/IJIRCCE.2026.1408024
PDF pdf/24_Speech Emotion Recognition Using MFCC, Autoencoders, And CNN-LSTM Networks.pdf
KEYWORDS
References [1] S. Latif, R. Rana, S. Younis, J. Qadir, and J. Epps, "Transfer learning for improving speech emotion classification accuracy," IEEE Access, vol. 9, pp. 18185-18195, 2021.
[2] H. M. Fayek, M. Lech, and L. Cavedon, "Evaluating deep learning architectures for speech emotion recognition," Neural Netw., vol. 92, pp. 60-68, 2021.
[3] Z. Zhang, B. Schuller, and S. Weninger, "Deep learning-based speech emotion recognition: A survey," IEEE Trans. Affective Comput., vol. 13, no. 2, pp. 617-639, 2022.
[4] Y. Kim and H. Kim, "Speech emotion recognition using CNN-LSTM hybrid networks," Expert Syst. Appl., vol. 186, pp. 115-128, 2022.
[5] J. Deng, X. Xu, and B. Schuller, "Deep neural networks for speech emotion recognition," Inf. Fusion, vol. 82, pp. 145-160, 2022.
[6] S. Hochreiter and J. Schmidhuber, "Long short-term memory," Neural Comput., vol. 9, no. 8, pp. 1735-1780, 1997.
[7] D. P. Kingma and M. Welling, "Auto-encoding variational Bayes," in Proc. Int. Conf. Learning Representations (ICLR), 2014.
[8] T. N. Sainath and B. Li, "Deep convolutional neural networks for speech recognition," IEEE Signal Process. Mag., vol. 29, no. 6, pp. 82-97, 2021.
[9] B. Logan, "Mel frequency cepstral coefficients for speech recognition," in Proc. Int. Symp. Signal Processing (ISSP), 2020, pp. 1-5.
[10] F. Eyben, M. Wollmer, and B. Schuller, "OpenSMILE: Audio feature extraction toolkit," in Proc. ACM Multimedia, 2021, pp. 1459-1462.
[11] I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning. Cambridge, MA, USA: MIT Press, 2016.
[12] R. B. Patel and S. Sharma, "Speech emotion recognition using MFCC and deep learning," IEEE Access, vol. 10, pp. 45122-45135, 2022.
[13] K. Han, D. Yu, and I. Tashev, "Speech emotion recognition using deep neural networks," in Proc. IEEE Int. Conf. Acoustics, Speech and Signal Processing (ICASSP), 2021, pp. 5800-5804.
[14] M. Neumann and N. Vu, "Improving speech emotion recognition with attention mechanisms," in Proc. INTERSPEECH, 2022, pp. 1003-1007.
[15] Y. Huang and J. Wang, "A hybrid CNN-LSTM framework for speech emotion recognition," Knowledge-Based Syst., vol. 241, pp. 108-120, 2023.
[16] P. Verma and R. Singh, "Deep learning techniques for speech emotion recognition: A comparative study," Appl. Soft Comput., vol. 126, 2023.
[17] A. Vaswani et al., "Attention is all you need," in Proc. Advances in Neural Information Processing Systems (NeurIPS), 2017.
[18] D. P. Kingma and J. Ba, "Adam: A method for stochastic optimization," in Proc. Int. Conf. Learning Representations (ICLR), 2015.
[19] TensorFlow Developers, "TensorFlow: An end-to-end open source machine learning platform." [Online]. Available: https://www.tensorflow.org
[20] Librosa Development Team, "Librosa: Python library for audio and music analysis." [Online]. Available: https://librosa.org


Copyright © IJIRCCE 2020.All right reserved