AiTA Lab, Faculty of Information Technology
FPT University, Ho Chi Minh Campus
Ho Chi Minh City 71320, Vietnam
Nguyen Minh Nhut
B.Sc. in Artificial Intelligence, FPT University, Ho Chi Minh Campus
I build multimodal and human-centered AI systems that help machines understand emotion, speech, and intent. My work sits at the intersection of cross-modal representation learning, graph neural networks, and robust speech processing — with a recurring interest in what happens when labels are scarce.
Research Focus
Multimodal Emotion Recognition
Cross-modal fusion across audio, text, and vision — graph attention, gated fusion, and adaptive attention for decoding expressive cues.
Speech & Audio Intelligence
Robust representation learning for emotion, speaker traits, and conversational analytics on top of transformer speech encoders.
Human-Centered AI
Datasets, interfaces, and evaluation loops that keep people at the centre of intelligent systems and make their behaviour legible.
Currently Exploring
- Emotion in conversational context Modelling emotional dynamics across a human–AI dialogue, so a system can perceive, track, and adapt to affective cues as the conversation unfolds.
- Scalable graph-based multimodal architectures Efficient graph and hypergraph networks for real-time inference over heterogeneous modalities — speech, text, and vision — in emotionally rich interactions.
- Semi-supervised and contrastive learning Using unlabelled multimodal data, including federated settings, to improve robustness and generalisation when annotation is limited.
News
Selected Papers
- Enhancing multimodal emotion recognition with dynamic fuzzy membership and attention fusionEngineering Applications of Artificial Intelligence, Feb 2026
- Multimodal fusion in speech emotion recognition: A comprehensive review of methods and technologiesEngineering Applications of Artificial Intelligence, Jan 2026
- In Proceedings of the 2025 Asia-Pacific Network Operations and Management Symposium (APNOMS), Sep 2025