Open-Vocabulary Semantic Mapping via 3D Gaussian Splatting
DOI:
https://doi.org/10.17979/ja-cea.2026.47.13803Keywords:
Autonomous Mobile Robots, Perception and sensing, Map building, Intelligent robotics, Machine LearningAbstract
Open-vocabulary semantic maps are key to the autonomy of mobile robots in complex, unstructured environments. Compared to traditional techniques, open-vocabulary semantic maps represent a significant improvement, as they are not limited by a predefined set of categories. For their construction, geometric representations based on 3D Gaussian Splatting enriched with low-level semantic descriptors have gained popularity. However, robotic operations demand high-level labels, such as textual descriptions and categories. This work presents an evolution of OpenSplat3D that integrates Large Vision-Language Models (LVLMs) for the description and categorization of object instances, whose consistency is validated within the latent space of Vision-Language Models (VLMs). Additionally, the geometric reconstruction has been optimized through dense initialization and the use of a depth-based loss function. Both proposals are validated using the Replica and ScanNet datasets, yielding accurate, semantically enriched maps.
References
Ambrosio-Cestero, G., Matez, J., Ruiz-Sarmiento, J., Gonzalez-Jimenez, J., 2024. Container based architecture for mobile robotics. En: Actas de las XLV Jornadas de Automática.
Dai, A., Chang, A. X., Savva, M., Halber, M., Funkhouser, T., Nießner, M., 2017. ScanNet: Richly-annotated 3d reconstructions of indoor scenes. En: Proc. CVPR, pp. 5828–5839.
Kerbl, B., Kopanas, G., Leimkühler, T., Drettakis, G., 2023. 3D Gaussian Splatting for real-time radiance field rendering. ACM Trans. Graph. 42 (4), 1–14.
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., et al., 2023. Segment anything. arXiv preprint arXiv:2304.02643.
Li, H., Wu, Y., Meng, J., Gao, Q., Zhang, Z., Wang, R., Zhang, J., 2025. InstanceGaussian: Appearance-semantic joint gaussian representation for 3d instance-level perception. En: Proc. CVPR.
Matez-Bandera, J.-L., Ojeda, P., Monroy, J., Gonzalez-Jimenez, J., Ruiz-Sarmiento, J.-R., 2024. Voxeland: Probabilistic instance-aware semantic mapping with evidence-based uncertainty quantification. arXiv preprint arXiv:2411.08727.
Piekenbrinck, J., Schmidt, C., Hermans, A., Vaskevicius, N., Linder, T., Leibe, B., 2025. OpenSplat3D: Open-Vocabulary 3D Instance Segmentation using Gaussian Splatting. En: Proc. CVPR, pp. 5246–5255.
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., et al., 2021. Learning transferable visual models from natural language supervision. En: Proc. ICML. PMLR, pp. 8748–8763.
Ravi, N., Gabeur, V., Hu, Y.-T., Hu, R., Ryali, C., Ma, T., et al., 2024. SAM 2: Segment anything in images and videos. arXiv preprint arXiv:2408.00714.
Rubio, F., Valero, F., Llopis-Albert, C., 2019. A review of mobile robots: Concepts, methods, theoretical framework, and applications. Int. J. Adv. Robot. Syst. 16 (2), 1–22.
Ruiz-Sarmiento, J.-R., Galindo, C., Gonzalez-Jimenez, J., 2017. Building multiversal semantic maps for mobile robot operation. Knowl. Based Syst. 119, 257–272. DOI: https://doi.org/10.1016/j.knosys.2016.12.016
Straub, J., Whelan, T., Ma, L., Chen, Y., Wijmans, E., Green, S., et al., 2019. The Replica dataset: A digital replica of indoor spaces. arXiv preprint arXiv:1906.05797.
Takmaz, A., Fedele, E., Sumner, R. W., Pollefeys, M., Tombari, F., Engelmann, F., 2023. OpenMask3D: Open-Vocabulary 3D Instance Segmentation. En: Proc. NeurIPS.
Wu, Y., Meng, J., Li, H., Wu, C., Shi, Y., Cheng, X., et al., 2024. OpenGaussian: Towards point-level 3d gaussian-based open vocabulary understanding. En: Proc. NeurIPS, pp. 19114–19138.
Xu, X., Xiong, T., Ding, Z., Tu, Z., 2023. MasQCLIP for open-vocabulary universal image segmentation. En: Proc. ICCV, pp. 887–898.
Zhu, R., Yu, M., Xu, L., Jiang, L., Li, Y., Zhang, T., et al., 2025. ObjectGS: Object-aware scene reconstruction and scene understanding via gaussian splatting. En: Proc. ICCV.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 J. Serrano-Mesa, J. Moncada-Ramirez, P. Ojeda, G. Ambrosio-Cestero, J.R. Ruiz-Sarmiento, J. Gonzalez-Jimenez

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.