Memoria Episódica Jerárquica para Razonamiento Espacial en Robots Asistenciales
DOI:
https://doi.org/10.17979/ja-cea.2026.47.13708Palabras clave:
Ingeniería de sistemas cognitivos, Robótica inteligente, Computación centrada en el ser humano, Percepción y sensorización, Robots móvilesResumen
Conseguir una interpretación semántica robusta de entornos domésticos representa uno de los principales desafíos en robótica de servicio, especialmente con alta densidad de objetos. Esta tarea se ve dificultada por la disposición estructural del entorno, donde las barreras arquitectónicas fragmentan la percepción y complican la organización lógica de la memoria. En este trabajo, presentamos un sistema de memoria episódica jerárquica diseñado para vincular el razonamiento de los modelos de lenguaje (LLM) a la geometría espacial. A diferencia de los métodos basados en proximidad, nuestra propuesta segmenta el espacio en zonas lógicas definidas por el campo de visión efectivo, obligando al sistema a respetar la arquitectura física. El sistema implementa una búsqueda mediante anclas semánticas y un algoritmo de desambiguación espacial para el conteo preciso de entidades. Los resultados muestran que esta arquitectura permite realizar tareas de búsqueda, conteo y explicación con alta fiabilidad, mitigando las alucinaciones del LLM mediante una validación visual proactiva que garantiza un razonamiento situado y veraz.
Referencias
H. Ali, P. Allgeuer, C. Mazzola, G. Belgiovine, B. C. Kaplan, L. Gajdošech, S. Wermter, "Robots can multitask too: Integrating a memory architecture and LLMs for enhanced cross-task robot action generation," in Proc. IEEE-RAS Int. Conf. Humanoid Robots (Humanoids), 2024, pp. 811–818.
E. W. Dijkstra, "A note on two problems in connexion with graphs," Numerische Mathematik, vol. 1, no. 1, pp. 269–271, 1959, doi: 10.1007/BF01386390.
V. S. Dorbala and D. Manocha, "Memctrl: Using MLLMs as active memory controllers on embodied agents," arXiv preprint arXiv:2601.20831, 2026.
N. Funk, J. Tarrio, S. Papatheodorou, M. Popović, P. F. Alcantarilla, and S. Leutenegger, "Multi-resolution 3D mapping with explicit free space representation for fast and accurate mobile robot motion planning," IEEE Robotics and Automation Letters, vol. 6, no. 2, pp. 3553–3560, 2021.
Q. Gu, A. Kuwajerwala, S. Morin, K. M. Jatavallabhula, B. Sen, A. Agarwal, C. Rivera, W. Paul, K. Ellis, R. Chellappa, et al., "Conceptgraphs: Open-vocabulary 3D scene graphs for perception and planning," in Proc. IEEE Int. Conf. Robotics and Automation (ICRA), 2024, pp. 5021–5028.
C. Huang, O. Mees, A. Zeng, and W. Burgard, "Visual language maps for robot navigation," in Proc. IEEE Int. Conf. Robotics and Automation (ICRA), 2023, pp. 10608–10615.
L. Keselman, J. Iselin Woodfill, A. Grunnet-Jepsen, and A. Bhowmik, "Intel RealSense stereoscopic depth cameras," in Proc. IEEE Conf. Computer Vision and Pattern Recognition Workshops, 2017, pp. 1–10.
Z. Liao, Zhang, Y., J. Luo, and W. Yuan, "TSM: Topological scene map for representation in indoor environment understanding," IEEE Access, vol. 8, pp. 185870–185884, 2020.
Y. Mao, H. Ye, W. Dong, C. Zhang, and H. Zhang, "Meta-memory: Retrieving and integrating semantic-spatial memories for robot spatial reasoning," arXiv preprint arXiv:2509.20754, 2025.
R. Martins, D. Bersan, M. F. Campos, and E. R. Nascimento, "Extending maps with semantic and contextual object information for robot navigation: a learning-based framework using visual and depth cues," Journal of Intelligent & Robotic Systems, vol. 99, no. 3, pp. 555–569, 2020.
A. Mora, A. Prados, A. Mendez, G. Espinoza, P. Gonzalez, B. Lopez, V. Muñoz, L. Moreno, S. Garrido, and R. Barber, "ADAM: a robotic companion for enhanced quality of life in aging populations," Frontiers in Neurorobotics, vol. 18, p. 1337608, 2024.
OpenAI, "Presentamos GPT-5," 2025. [En línea]. Disponible en: https://openai.com/index/introducing-gpt-5/. [Accedido: 05-may-2026].
J. Plewnia and T. Asfour, "Combining episodic memory and LLMs for the verbalization of robot experiences," in Proc. IEEE-RAS Int. Conf. Humanoid Robots (Humanoids), 2025, pp. 531–538.
N. Reimers et al., "all-mpnet-base-v2: Sentence transformer model," Hugging Face, 2021. [En línea]. Disponible en: https://huggingface.co/sentence-transformers/all-mpnet-base-v2. [Accedido: 05-may-2026].
S. Shaji, F. Huppertz, A. Mitrevski, and S. Houben, "From language to action: Can LLM-based agents be used for embodied robot cognition?" arXiv preprint arXiv:2603.03148, 2026.
T. Vincenty, "Direct and inverse solutions of geodesics on the ellipsoid with application of geocentric Hexagonal co-ordinates," Survey Review, vol. 23, no. 176, pp. 88–93, 1975, doi: 10.1179/sre.1975.23.176.88.
Z. Wang and G. Tian, "Hybrid offline and online task planning for service robot using object-level semantic map and probabilistic inference," Information Sciences, vol. 593, pp. 78–98, 2022.
J. H. Ward, Jr., "Hierarchical grouping to optimize an objective function," Journal of the American Statistical Association, vol. 58, no. 301, pp. 236–244, 1963, doi: 10.1080/01621459.1963.10500845.
Q. Xie, S. Y. Min, P. Ji, Y. Yang, T. Zhang, K. Xu, A. Bajaj, R. Salakhutdinov, M. Johnson-Roberson, and Y. Bisk, "Embodied-RAG: General non-parametric embodied memory for retrieval and generation," arXiv preprint arXiv:2409.18313, 2024.
J. Yang, S. Yang, A. W. Gupta, R. Han, L. Fei-Fei, and S. Xie, "Thinking in space: How multimodal large language models see, remember, and recall spaces," in Proc. Comput. Vis. Pattern Recognit. Conf. (CVPR), 2025, pp. 10632–10643.
J. Zha, Y. Fan, X. Yang, C. Gao, and X. Chen, "How to enable LLM with 3D capacity? a survey of spatial reasoning in LLM," arXiv preprint arXiv:2504.05786, 2025.
J. Zhang, S. Wu, X. Ma, and S. Schwertfeger, "Generation of indoor open street maps for robot navigation from CAD files," arXiv preprint arXiv:2507.00552, 2025.
Descargas
Publicado
Número
Sección
Licencia
Derechos de autor 2026 Alicia Mora, Lucía Lishan, Gonzalo Espinoza, Luis Moreno, Ramón Barber

Esta obra está bajo una licencia internacional Creative Commons Atribución-NoComercial-CompartirIgual 4.0.