Executive Summary

Technology

December 10, 2025

Performance comparison between Machine Learning and Generative AI in emotion detection

Author: Carlos Eduardo Frantz Manchini — Advisor: Thiago Gentil Ramires

Summary prepared by the ResumeAI tool, an artificial intelligence solution developed by the Pecege Institute focused on synthesis and writing.

This work compared traditional machine learning models with the Hume AI Generative Artificial Intelligence solution in the task of inferring emotions from the RAVDESS audio database. The investigation sought to determine which approach offers greater accuracy and balance in classifying a diverse spectrum of emotional states, contrasting the classic methodology, based on acoustic feature engineering, with the end-to-end approach of pre-trained generative models. The analysis aims to provide subsidies for the selection of technologies in practical applications, such as mental health and human-computer interaction, where accuracy in detecting emotional nuances is fundamental.

Advances in Data Science and Artificial Intelligence (AI) allow for the processing and interpretation of unstructured data, such as text, images, and audio. According to Gartner (2020), approximately 80% of new corporate data is of this nature, making insight extraction a competitive differentiator. This large-scale processing has been enabled by Machine Learning and Deep Learning techniques, which identify complex patterns inaccessible to conventional methods (Goodfellow et al., 2016). The rise of Generative Artificial Intelligence (GenAI), with its large language models (LLMs), has accelerated this transformation.

Innovations driven by GenAI are generating impacts in various areas. In healthcare, generative models assist in diagnosing diseases such as cancer, with accuracies exceeding 99%, in contrast to the 80% of traditional methods (Sheakh et al., 2024), and accelerate the development of new treatments (Topol, 2019). In security, AI enhances facial recognition, fraud detection, and intelligent monitoring, preventing crimes and protecting assets (Li et al., 2020). In education, GenAI enables the creation of virtual assistants and adaptive learning platforms that personalize content to each student’s pace (Luckin et al., 2016).

In audio processing, AI is the foundation of voice assistants and machine translation systems (Tan et al., 2021). In this context, automatic speech emotion recognition (SER) is a complex research area. The challenge lies in the subtle nature of emotions, manifested in acoustic patterns. According to Akçay and Oğuz (2020), SER models variations in intonation (pitch), rhythm, intensity, and timbre, which carry information about the speaker’s affective state. Deep neural network-based GenAI models are being explored to process audio signals more holistically (Purwins et al., 2019).

The motivation for this study is the need to critically evaluate the performance of generative models compared to established methods. Although deep learning models, which underpin GenAI, often outperform conventional techniques in pattern recognition (Khare et al., 2024), interpretability and dependence on large data volumes are challenges (Latif et al., 2020). This study investigates the advantages of each approach, considering accuracy, robustness, computational efficiency, and generalization, to guide future implementations. The applications of SER are vast, ranging from conversational interfaces to mental health tools, where alterations in vocal prosody can be biomarkers for the detection of stress, anxiety, and depression (Cummins et al., 2015).

The research is an experimental study of an applied nature, with quantitative and comparative evaluation of two approaches for audio emotion recognition. The first, traditional, uses feature engineering and supervised Machine Learning algorithms. The second employs a Generative AI solution that processes data end-to-end. The methodology was designed for rigorous comparison, with controlled variables and standardized metrics. The experiments were implemented in Python, with specialized libraries for data manipulation, feature extraction, modeling, and evaluation.

The dataset used was the Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS), a recognized database in the SER community (Livingstone and Russo, 2018). The corpus contains 1440 audio files from 24 professional actors (12 males, 12 females) expressing eight emotions: anger, calm, disgust, fear, happiness, neutral, sadness, and surprise. RAVDESS was chosen for its balanced structure, high quality, and clear labeling. The use of the CORAA corpus (Alvim et al., 2022), in Brazilian Portuguese, was considered, but the absence of detailed emotional labels made its use unfeasible. Pre-processing was performed on the audios, including noise reduction and volume normalization, to standardize the signals and improve pattern detection, as recommended by Cowie et al. (2001).

In the traditional approach, an acoustic feature extraction pipeline was followed, as described by Akçay and Oğuz (2020). Features such as Mel-Frequency Cepstral Coefficients (MFCCs), Chroma Features, Zero-Crossing Rate (ZCR), RMS energy, and fundamental frequency (pitch) features were extracted. With this set of 81 features, three supervised learning algorithms were trained: Random Forest, Support Vector Machines (SVM), and XGBoost, following the practices of Schuller et al. (2011). The evaluation used stratified cross-validation and metrics such as accuracy, precision, F1-score, and confusion matrix.

For the generative approach, the strategy of transcribing the audio with Whisper (OpenAI, 2022) and analyzing the text with an LLM was discarded, as it would neglect prosodic nuances. Instead, the Expression Measurement – Prosody API from Hume AI (2025) was adopted, a solution that uses deep learning models to infer emotions directly from audio, eliminating manual feature extraction. The implementation consisted of sending the 1440 RAVDESS audios to the API via HTTP calls, with subsequent processing of the returned scores to determine the predominant emotion. The Librosa and pydub libraries were used for audio manipulation, Scikit-learn for metric calculation, and Matplotlib and Seaborn for visualizations.

The exploratory data analysis sought to understand the behavior of acoustic characteristics. The visualization of Log-Mel spectrograms revealed distinct patterns between high and low valence emotions. “Anger” audios exhibited energy concentrated in higher frequencies, while “sadness” audios presented a flatter pitch contour and lower intensity. These observations validated that prosodic features contained relevant information for classification, justifying feature engineering. The quantitative analysis of the 81 extracted features formed the basis for variable selection.

Feature selection was a crucial step to build a more parsimonious model less prone to overfitting. Pearson correlation analysis revealed high multicollinearity. The 12 Chroma features were replaced by three principal components (PCA) that preserved 88% of the variability. To mitigate redundancy, a feature importance-based selection method calculated by XGBoost was employed. The comparison between the full model (81 features) and a reduced one (23 features) showed the effectiveness of the approach: the full model achieved an F1-Macro of 0.648, while the reduced one obtained 0.635, a performance drop of only 2% with a 71.6% reduction in the number of variables. The compact model was adopted for the next steps.

The modeling with supervised classifiers was performed on a stratified split of the data (70% for training, 30% for testing). The Support Vector Machines (SVM) model showed the best overall performance, with an accuracy of 61.8% and an F1-Macro of 0.58 on the test set. The tree-based models, Random Forest and XGBoost, showed signs of overfitting, with a drop in performance on the test data. The superiority of SVM was attributed to its greater stability and generalization capability in a high-dimensional feature space. The F1-Macro metric was important for being robust for multiclass problems with imbalance.

The analysis of the SVM confusion matrix revealed its patterns of correct and incorrect classifications. The model was effective in recognizing emotions with distinct acoustic characteristics, such as “calm” and “surprise”. However, systematic confusions occurred between emotions with similar sound properties. The classes “neutral”, “calm”, and “sad”, characterized by low intensity and less pronounced pitch variations, showed significant overlap. Similarly, high-intensity emotions like “fear” and “disgust” were also occasionally confused, indicating that distinguishing subtle nuances remains a challenge.

In the generative pipeline, the evaluation of the Hume AI API revealed a challenge with emotional taxonomy incompatibility. While RAVDESS uses eight discrete categories, Hume AI returns scores for a different set of labels, such as “boredom” and “amusement”. To allow for direct comparison, an ontological mapping was necessary, semantically approximating Hume’s labels to RAVDESS classes (e.g., “boredom” mapped to “neutral”, “amusement” to “happiness”). This harmonization was crucial for a fair quantitative evaluation.

Even after mapping, Hume AI’s performance was significantly lower than that of supervised models, with an overall accuracy of only 29% and an F1-Macro of 0.26. Hume AI’s confusion matrix showed a specific pattern: the model identified the emotion “anger” well, with an F1-score of 0.65 for this class. However, for more subtle emotions, performance was extremely low, with F1-scores close to zero for “fear” and “disgust”. The model tended to confuse most emotions with neutral or calm categories, indicating a limitation in capturing the diversity of the emotional spectrum.

The direct comparison of the results consolidates the superiority of the traditional approach. Supervised models, led by SVM, achieved an average of 55% accuracy and 0.53 F1-Macro, with more balanced recognition among the eight classes. In contrast, the Hume API, with 29% accuracy and 0.26 F1-Macro, proved inadequate for detailed classification. While traditional models got between half and two-thirds of the predictions right, the generative solution correctly classified just over one in five samples. This reinforces that, despite greater effort in preprocessing and tuning, supervised classifiers offer a consistency that “ready-to-use” solutions have not yet matched.

Considering the practical aspects, the traditional pipeline, although more complex to implement, offers full control, interpretability, and does not incur operational costs per prediction. The Hume API offers simplicity, but its “black box” nature limits interpretability and introduces a financial cost. Processing the 1440 audios in this study cost approximately US$3.60, a factor to consider on a large scale. The trade-off is that the convenience of GenAI comes at the cost of reduced performance and ongoing operational expense, while the classic approach requires a greater initial investment in development to achieve more robust results.

This work confirmed the superiority of supervised models for audio emotion recognition. The Support Vector Machine (SVM) classifier stood out with an accuracy of 61.8% on the test set, demonstrating robust generalization and balanced classification. In contrast, Hume AI’s generative solution, although simple to operate, had inferior performance with 29% accuracy. The API showed aptitude for identifying high-intensity emotions like anger but failed to distinguish nuances between subtler categories such as fear and disgust, confusing them with neutral states. The conclusions point to a trade-off between convenience and precision. “Out-of-the-box” generative solutions may serve for prototyping or in applications focused on intense emotions. However, for scenarios requiring high accuracy and a detailed emotional spectrum, supervised models, fine-tuned for the task, remain more suitable. As future perspectives, we suggest exploring hybrid approaches, fine-tuning generative models, and applying them to more heterogeneous databases, including Brazilian Portuguese audio. It is concluded that the objective was achieved: it was demonstrated that supervised models, such as SVM, remain more accurate and robust for classifying a broad emotional spectrum in audio compared to the evaluated Generative AI solution.

References:
Akçay, M. B.; Oğuz, K. 2020. Speech emotion recognition: emotional models, databases, features, preprocessing methods, supporting modalities, and classifiers. Speech Communication 116: 56-76.
Alvim, G.; Magalhães, L. L. C.; Bigal, R. L. M.; Medeiros, H. F. G.; Souza, S. R. M.; Silva, C. F. M.; Oliveira, E. P. L.; Pardo, S. R. C. 2022. CORAA: a large corpus of spontaneous and prepared speech manually validated for speech recognition in Brazilian Portuguese. In: International Conference on Language Resources and Evaluation, 2022, Marseille, France. Proceedings… p. 5908-5915.
Cowie, R.; Douglas-Cowie, E.; Tsapatsoulis, N.; Votsis, G.; Kollias, S.; Fellenz, W.; Taylor, J. G. 2001. Emotion recognition in human-computer interaction. IEEE Signal Processing Magazine 18(1): 32-80.
Cummins, N.; Scherer, S.; Krajewski, J.; Schnieder, S.; Epps, J.; Quatieri, T. F. 2015. A review of depression and suicide risk assessment using speech analysis. Speech Communication 71: 10-49.
Gartner [GARTNER]. 2020. Market guide for text analytics. Available at: <https://www. gartner. com/en/documents/3989657>. Accessed on: Mar. 24, 2025.
Goodfellow, I.; Bengio, Y.; Courville, A. 2016. Deep Learning. The MIT Press, Cambridge, MA, USA.
Hume AI. 2025. Expression Measurement – Prosody. Available at: <https://dev. hume. ai/docs/expression-measurement>. Accessed on: Sep. 25, 2025.
Khare, S. K.; Blanes-Vidal, V.; Nadimi, E. S.; Acharya, U. R. 2024. Emotion recognition and artificial intelligence: a systematic review (2014–2023) and research recommendations. Information Fusion 102: 102019.
Latif, S.; Rana, R.; Qadir, J.; Epps, J.; Schuller, B. W. 2020. Deep representation learning in speech processing: challenges, recent advances, and future trends. Computer Speech & Language 68: 101-178.
Li, Y.; Schuckert, M.; Law, R.; Wang, J. 2020. The impact of artificial intelligence on security and privacy in smart cities. Journal of Urban Technology 27(2): 65-85.
Livingstone, S. R.; Russo, F. A. 2018. The ryerson audio-visual database of emotional speech and song (RAVDESS). PLoS ONE 13(5): e0196391.
Luckin, R.; Holmes, W. 2016. Intelligence Unleashed: An Argument for AI in Education. Pearson, London, UK.
OpenAI. 2022. Whisper: robust speech recognition via large-scale audio training. Available at: <https://github. com/openai/whisper>. Accessed on: Mar. 29, 2025.
Poria, S.; Majumder, N.; Mihalcea, R.; Hovy, E. 2019. Emotion recognition in conversation: research challenges, datasets, and recent advances. IEEE Access 7: 100943-100953.
Purwins, H.; Li, B.; Virtanen, T.; Schlüter, J.; Chang, S. Y.; Sainath, T. 2019. Deep learning for audio signal processing. IEEE Journal of Selected Topics in Signal Processing 13(2): 206-219.
Schuller, B.; Batliner, A.; Steidl, S.; Seppi, D. 2011. Recognizing realistic emotions and affect in speech: state of the art and lessons learnt from the first challenge. Speech Communication 53(9-10): 1062-1087.
Sheakh, M. A.; Azam, S.; Tahosin, M. S.; Karim, A.; Montaha, S.; Fahim, K. U.; De Boer, F. 2024. ECgMLP: a novel gated MLP model for enhanced endometrial cancer diagnosis. Computer Methods and Programs in Biomedicine Update 5: 100181.
Tan, X.; Qin, T.; Soong, F.; Liu, T.-Y. 2021. A survey on neural speech synthesis. Available at: <https://arxiv. org/abs/2106.15561>. Accessed on: Sep. 27, 2025.
Topol, E. 2019. Deep Medicine: How Artificial Intelligence Can Make Healthcare Human Again. Basic Books, New York, NY, USA.


Executive summary from the Final Project of the Specialization in Data Science and Analytics from the MBA USP/Esalq

Learn more about the course; click here:

Who edited this article

Most recent

You may also like