Multimodal Emotion Detection System: Enhancing Human-Computer Interaction through Facial, Text, and Voice Analysis
DOI:
https://doi.org/10.65138/ijramt.2026.v7i6.3266Abstract
This paper proposes a system for detecting emotions across multiple modalities, merging facial, textual, and vocal data to improve human-computer interaction. The system operates three distinct processing channels to handle each data type: facial expressions undergo examination via a Haar Cascade classifier and a convolutional neural network, textual input is assessed with natural language processing methods, and audio signals are initially converted to text before undergoing sentiment analysis. These processing systems produce distinct emotion classifications, which are subsequently merged into an ultimate decision via a multimodal blending approach. The proposed method addresses the limitations of unimodal approaches by capturing complementary emotional cues from diverse sources, thereby improving robustness and accuracy. Additionally, the system is designed as a real-time web interface with Python, Flask, and OpenCV, which supports smooth user interaction. Our research advances affective computing by showing the concrete advantages of merging multiple modalities, especially in cases where single-modality approaches might be inadequate because of noise or uncertainty. The experimental findings underscore the system’s capacity to adjust to diverse input conditions, which renders it appropriate for applications including virtual assistants, mental health monitoring, and interactive learning environments. The modular architecture also supports potential future additions, for example adding other modalities or improving the fusion approach. This research advances the development of more intuitive and responsive human-computer interfaces.
Downloads
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Padmaja Guthula, Rajasekharam Bonthu

This work is licensed under a Creative Commons Attribution 4.0 International License.
