Multimodal attention and robust interlocutor tracking in natural HRI
DOI:
https://doi.org/10.17979/ja-cea.2026.47.13811Keywords:
Intelligent robotics, Multi-modal interaction, Perception and sensing, Mobile robots, Human and vehicle interactionAbstract
This paper presents a multimodal architecture designed to enhance naturalness in Human-Robot Interaction. To overcome spatial ambiguities, a microphone array is integrated to estimate the direction of arrival of sound omnidirectionally. This allows the robot to orient itself naturally through the coordination of its head and mobile base. Furthermore, an efficient visual recognition method is proposed that shifts the computational burden from continuous facial analysis to full-body tracking. This approach utilizes an identity cache that successfully maintains the interlocutor's identity during temporary occlusions or face turns. The system is completed with a conversational agent managed by Large Language Models and an expressivity module that synchronizes the mouth with speech and adapts the facial expression to the dialogue context. The architecture has been successfully validated on a social robot from the MAPIR-UMA research group.
References
Alansari, M., Alnuaimi, K., Ganapathi, I., Alansari, S., Javed, S., Shoufan, A., Zweiri, Y., Werghi, N., 2025. Efficientfacev2s: A lightweight model and a benchmarking approach for drone-captured face recognition. Expert Systems with Applications 273, 126786. DOI: https://doi.org/10.1016/j.eswa.2025.126786
Cañete, A., Quemada-Torres, E., Ruiz-Sarmiento, J.-R., Moreno, F. Á., Gonzalez-Jimenez, J., 2024. Sistema multimodal para la orientación de robots móviles hacia su interlocutor. Jornadas de Automática 45. DOI: https://doi.org/10.17979/ja-cea.2024.45.10939
Desai, D., Mehendale, N., 2022. A review on sound source localization systems. Archives of Computational Methods in Engineering 29, 4631–4642. DOI: https://doi.org/10.1007/s11831-022-09747-2
Ge, Z., Liu, S., Wang, F., Li, Z., Sun, J., 2021. Yolox: Exceeding yolo series in 2021. arXiv preprint arXiv:2107.08430. DOI: https://doi.org/10.48550/arXiv.2107.08430
Khalifa, A., Abdelrahman, A. A., Strazdas, D., Hintz, J., Hempel, T., Al-Hamadi, A., 2022. Face recognition and tracking framework for human-robot interaction. Applied Sciences 12 (11), 5568. DOI: https://doi.org/10.3390/app12115568
Khan, S. S., Sengupta, D., Ghosh, A., Chaudhuri, A., 2024. Mtcnn++: A cnn-based face detection algorithm inspired by mtcnn. The Visual Computer 40, 899–917. DOI: https://doi.org/10.1007/s00371-023-02822-0
Kumar, A., Kaur, A., Kumar, M., 2019. Face detection techniques: a review. Artificial Intelligence Review 52 (2), 927–948. DOI: https://doi.org/10.1007/s10462-018-9650-2
Matheus, K., Ramnauth, R., Scassellati, B., Salomons, N., 2025. Long-term interactions with social robots: Trends, insights, and recommendations. J. Hum.-Robot Interact. 14 (3). DOI: https://doi.org/10.1145/3729539
Picovoice, 2026. Porcupine: On-device wake word detection powered by deep learning. https://github.com/Picovoice/porcupine, repositorio de GitHub. Accedido el 21 de abril de 2026.
Quemada-Torres, E., Jun. 2025. Sistema multimodal de interacción humano robot basado en técnicas de ia. Trabajo fin de grado, Universidad de Málaga. URL: https://hdl.handle.net/10630/40843
Scripka, D., 2026. openwakeword: An open-source audio wake word detection framework. https://github.com/dscripka/openWakeWord, repositorio de GitHub. Accedido el 21 de abril de 2026.
Wang, Y.-Q., 2014. An Analysis of the Viola-Jones Face Detection Algorithm. Image Processing On Line 4, 128–148. DOI: https://doi.org/10.5201/ipol.2014.104
Zangemeister, W. H., Stark, L., 1982. Gaze latency: Variable interactions of head and eye latency. Experimental Neurology 75 (2), 389–406. DOI: https://doi.org/10.1016/0014-4886(82)90169-8
Zhang, K., Zhang, Z., Li, Z., Qiao, Y., 2016. Joint face detection and alignment using multitask cascaded convolutional networks. IEEE Signal Processing Letters 23 (10), 1499–1503. DOI: https://doi.org/10.1109/LSP.2016.2603342
Zhang, Y., Sun, P., et al., 2022. Bytetrack: Multi-object tracking by associating every detection box. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer, pp. 1–21. DOI: https://doi.org/10.1007/978-3-031-20047-2_1
Downloads
Published
Issue
Section
License
Copyright (c) 2026 A. Valencia-Rojas, P. Hormigo-Jiménez, E. Quemada-Torres, F.-A. Moreno, J. González-Jiménez

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.