A Robust Framework for Real-Time Speech Emotion Analysis
Main Article Content
Abstract
Speech Emotion Recognition (SER) focuses on figuring out someone’s emotional state just by listening to their voice. This tech sits at the heart of friendlier human-computer conversations, smarter virtual assistants, mental health monitoring, and a whole field called affective computing. Over the years, SER research has moved from simple handcrafted features and basic machine learning to today’s powerful deep learning and multimodal approaches. In this paper, we dive into the full evolution of SER methods from 2003 to 2025. We break down the main research approaches—running from classic machine learning right up to deep neural networks and transformer-based multimodal systems. Along the way, we take a close look at common datasets and performance trends. The review also puts a spotlight on big challenges, like generalizing across different datasets, handling imbalanced data, and making models more explainable. We wrap up by introducing a hybrid Transformer–CNN design boosted with explainable AI to make SER more robust and transparent.