Hierarchical Episodic Memory for Spatial Reasoning in Assistive Robots

Authors

  • Alicia Mora Universidad Carlos III de Madrid
  • Lucía Lishan Universidad Carlos III de Madrid
  • Gonzalo Espinoza Universidad Carlos III de Madrid
  • Luis Moreno Universidad Carlos III de Madrid
  • Ramón Barber Universidad Carlos III de Madrid

DOI:

https://doi.org/10.17979/ja-cea.2026.47.13708

Keywords:

Cognitive systems engineering, Intelligent robotics, Human-centered computing, Perception and sensing, Mobile robots

Abstract

Achieving a robust semantic interpretation of domestic environments represents one of the main challenges in service robotics, particularly in settings with a high density of objects. This task is made more difficult by the structural layout of the environment, where architectural barriers fragment perception and complicate the logical organisation of memory. In this work, we present a hierarchical episodic memory system designed to link the reasoning of language models (LLMs) to spatial geometry. Unlike proximity-based methods, our proposal segments the space into logical zones defined by the effective field of view, forcing the system to respect the physical architecture. The system implements a search using semantic anchors and a spatial disambiguation algorithm for the accurate counting of entities. The results show that this architecture enables search, counting and explanation tasks to be performed with high reliability, mitigating LLM hallucinations through proactive visual validation that ensures situated and accurate reasoning.

References

H. Ali, P. Allgeuer, C. Mazzola, G. Belgiovine, B. C. Kaplan, L. Gajdošech, S. Wermter, "Robots can multitask too: Integrating a memory architecture and LLMs for enhanced cross-task robot action generation," in Proc. IEEE-RAS Int. Conf. Humanoid Robots (Humanoids), 2024, pp. 811–818.

E. W. Dijkstra, "A note on two problems in connexion with graphs," Numerische Mathematik, vol. 1, no. 1, pp. 269–271, 1959, doi: 10.1007/BF01386390.

V. S. Dorbala and D. Manocha, "Memctrl: Using MLLMs as active memory controllers on embodied agents," arXiv preprint arXiv:2601.20831, 2026.

N. Funk, J. Tarrio, S. Papatheodorou, M. Popović, P. F. Alcantarilla, and S. Leutenegger, "Multi-resolution 3D mapping with explicit free space representation for fast and accurate mobile robot motion planning," IEEE Robotics and Automation Letters, vol. 6, no. 2, pp. 3553–3560, 2021.

Q. Gu, A. Kuwajerwala, S. Morin, K. M. Jatavallabhula, B. Sen, A. Agarwal, C. Rivera, W. Paul, K. Ellis, R. Chellappa, et al., "Conceptgraphs: Open-vocabulary 3D scene graphs for perception and planning," in Proc. IEEE Int. Conf. Robotics and Automation (ICRA), 2024, pp. 5021–5028.

C. Huang, O. Mees, A. Zeng, and W. Burgard, "Visual language maps for robot navigation," in Proc. IEEE Int. Conf. Robotics and Automation (ICRA), 2023, pp. 10608–10615.

L. Keselman, J. Iselin Woodfill, A. Grunnet-Jepsen, and A. Bhowmik, "Intel RealSense stereoscopic depth cameras," in Proc. IEEE Conf. Computer Vision and Pattern Recognition Workshops, 2017, pp. 1–10.

Z. Liao, Zhang, Y., J. Luo, and W. Yuan, "TSM: Topological scene map for representation in indoor environment understanding," IEEE Access, vol. 8, pp. 185870–185884, 2020.

Y. Mao, H. Ye, W. Dong, C. Zhang, and H. Zhang, "Meta-memory: Retrieving and integrating semantic-spatial memories for robot spatial reasoning," arXiv preprint arXiv:2509.20754, 2025.

R. Martins, D. Bersan, M. F. Campos, and E. R. Nascimento, "Extending maps with semantic and contextual object information for robot navigation: a learning-based framework using visual and depth cues," Journal of Intelligent & Robotic Systems, vol. 99, no. 3, pp. 555–569, 2020.

A. Mora, A. Prados, A. Mendez, G. Espinoza, P. Gonzalez, B. Lopez, V. Muñoz, L. Moreno, S. Garrido, and R. Barber, "ADAM: a robotic companion for enhanced quality of life in aging populations," Frontiers in Neurorobotics, vol. 18, p. 1337608, 2024.

OpenAI, "Presentamos GPT-5," 2025. [En línea]. Disponible en: https://openai.com/index/introducing-gpt-5/. [Accedido: 05-may-2026].

J. Plewnia and T. Asfour, "Combining episodic memory and LLMs for the verbalization of robot experiences," in Proc. IEEE-RAS Int. Conf. Humanoid Robots (Humanoids), 2025, pp. 531–538.

N. Reimers et al., "all-mpnet-base-v2: Sentence transformer model," Hugging Face, 2021. [En línea]. Disponible en: https://huggingface.co/sentence-transformers/all-mpnet-base-v2. [Accedido: 05-may-2026].

S. Shaji, F. Huppertz, A. Mitrevski, and S. Houben, "From language to action: Can LLM-based agents be used for embodied robot cognition?" arXiv preprint arXiv:2603.03148, 2026.

T. Vincenty, "Direct and inverse solutions of geodesics on the ellipsoid with application of geocentric Hexagonal co-ordinates," Survey Review, vol. 23, no. 176, pp. 88–93, 1975, doi: 10.1179/sre.1975.23.176.88.

Z. Wang and G. Tian, "Hybrid offline and online task planning for service robot using object-level semantic map and probabilistic inference," Information Sciences, vol. 593, pp. 78–98, 2022.

J. H. Ward, Jr., "Hierarchical grouping to optimize an objective function," Journal of the American Statistical Association, vol. 58, no. 301, pp. 236–244, 1963, doi: 10.1080/01621459.1963.10500845.

Q. Xie, S. Y. Min, P. Ji, Y. Yang, T. Zhang, K. Xu, A. Bajaj, R. Salakhutdinov, M. Johnson-Roberson, and Y. Bisk, "Embodied-RAG: General non-parametric embodied memory for retrieval and generation," arXiv preprint arXiv:2409.18313, 2024.

J. Yang, S. Yang, A. W. Gupta, R. Han, L. Fei-Fei, and S. Xie, "Thinking in space: How multimodal large language models see, remember, and recall spaces," in Proc. Comput. Vis. Pattern Recognit. Conf. (CVPR), 2025, pp. 10632–10643.

J. Zha, Y. Fan, X. Yang, C. Gao, and X. Chen, "How to enable LLM with 3D capacity? a survey of spatial reasoning in LLM," arXiv preprint arXiv:2504.05786, 2025.

J. Zhang, S. Wu, X. Ma, and S. Schwertfeger, "Generation of indoor open street maps for robot navigation from CAD files," arXiv preprint arXiv:2507.00552, 2025.

Downloads

Published

2026-09-01

Issue

Section

Robótica