Visual localization using cross-view techniques with satellite imagery

Authors

  • Míriam Máximo Gutiérrez Instituto de Investigación en Ingeniería de Elche (I3E), Universidad Miguel Hernández de Elche, Avda. de la Universidad s\/n, 03202 Elche (Alicante), España. https://orcid.org/0009-0005-5297-9059
  • Judith Vilella Cantos Instituto de Investigación en Ingeniería de Elche (I3E), Universidad Miguel Hernández de Elche, Avda. de la Universidad s\/n, 03202 Elche (Alicante), España. https://orcid.org/0009-0008-7220-7963
  • María Flores Tenza Instituto de Investigación en Ingeniería de Elche (I3E), Universidad Miguel Hernández de Elche, Avda. de la Universidad s\/n, 03202 Elche (Alicante), España. https://orcid.org/0000-0003-1117-0868
  • Luis Payá Castelló Instituto de Investigación en Ingeniería de Elche (I3E), Universidad Miguel Hernández de Elche, Avda. de la Universidad s\/n, 03202 Elche (Alicante), España. https://orcid.org/0000-0002-3045-4316
  • David Valiente García Instituto de Investigación en Ingeniería de Elche (I3E), Universidad Miguel Hernández de Elche, Avda. de la Universidad s\/n, 03202 Elche (Alicante), España. https://orcid.org/0000-0002-2245-0542
  • Mónica Ballesta Galdeano Instituto de Investigación en Ingeniería de Elche (I3E), Universidad Miguel Hernández de Elche, Avda. de la Universidad s\/n, 03202 Elche (Alicante), España. https://orcid.org/0000-0002-8029-5085

DOI:

https://doi.org/10.17979/ja-cea.2026.47.13547

Keywords:

Perception and sensing, Image processing, Programming and Vision, Sensor integration and perception, Position estimation, Deep Learning, Neural Networks, Mobile Robot

Abstract

Using open-access resources, such as satellite imagery employed as maps for mobile robot localization, represents a significant advantage in terms of temporal and economic efficiency, offering extensive global coverage without the need for prior mapping. Furthermore, autonomous navigation of mobile robots strictly requires high precision in both position and orientation. To this end, this article adapts a deep learning architecture originally designed to estimate the two-dimensional pose through the correlation of ground-level and satellite images. Specifically, the ground-level input has been modified to spherical projections generated from 3D LiDAR point clouds. The deployment of this sensor substantially enhances the system's robustness against illumination changes, thereby overcoming the limitations of conventional cameras. Finally, an experimental evaluation conducted using the KITTI dataset successfully demonstrates the viability of the proposed approach.

References

Alfaro, M., Cabrera, J. J., Reinoso, ´O., Gil, A., Pay´a, L., 2025. Localizaci´on visual mediante im´agenes omnidireccionales y técnicas de fusi´on temprana. Jornadas de Autom´atica (46). DOI: 10.17979/ja-cea.2025.46.12239

Chen, X., L¨abe, T., Milioto, A., R¨ohling, T., Vysotska, O., Haag, A., Behley, J., Stachniss, C., 2020. OverlapNet: Loop Closing for LiDAR-based SLAM. In: Proceedings of Robotics: Science and Systems (RSS). DOI: 10.15607/RSS.2020.XVI.009

Fervers, F., Bullinger, S., Bodensteiner, C., Arens, M., Stiefelhagen, R., 2022. Continuous Self-Localization on Aerial Images Using Visual and Lidar Sensors. In: 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). pp. 7028–7035. DOI: 10.1109/IROS47612.2022.9982195

Fu, M., Zhu, M., Yang, Y., Song, W., Wang, M., 2020. Lidar-based vehicle localization on the satellite image via a neural network. Robotics and Autonomous Systems 129, 103519. DOI: https://doi.org/10.1016/j.robot.2020.103519

Geiger, A., Lenz, P., Stiller, C., Urtasun, R., 2013. Vision meets robotics: The KITTI dataset. The international journal of robotics research 32 (11), 1231–1237. DOI: 10.1177/0278364913491297

Guo, J., Amayri, M., Bouguila, N., Liu, X., Fan, W., 2026. Enhancing rotation-invariant 3d learning with global pose awareness and attention mechanisms. In: Proceedings of the AAAI Conference on Artificial Intelligence. Vol. 40. pp. 4403–4411.

Hu, S., Feng, M., Nguyen, R. M., Lee, G. H., 2018. CVM-Net: Cross-View Matching Network for Image-Based Ground-to-Aerial Geo-Localization. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 7258–7267. DOI: 10.1109/CVPR.2018.00758

Kang, S., Liao, M. Y., Xia, Y., Wysocki, O., Jutzi, B., Cremers, D., 2025. OPAL: Visibility-aware LiDAR-to-OpenStreetMap Place Recognition via Adaptive Radial Fusion. arXiv preprint arXiv:2504.19258. DOI: 10.48550/arXiv.2504.19258

Li, L., Kong, X., Zhao, X., Huang, T., Li, W., Wen, F., Zhang, H., Liu, Y., 2022. RINet: Efficient 3D Lidar-Based Place Recognition Using Rotation Invariant Neural Network. IEEE Robotics and Automation Letters 7 (2), 4321–4328. DOI: 10.1109/LRA.2022.3150499

Macario Barros, A., Michel, M., Moline, Y., Corre, G., Carrel, F., 2022. A Comprehensive Survey of Visual SLAM Algorithms. Robotics 11 (1), 24. DOI: 10.3390/robotics11010024

Shi, Y., Li, H., 2022. Beyond cross-view image retrieval: Highly accurate vehicle localization using satellite image. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 17010–17020. DOI: 10.1109/CVPR52688.2022.01650

Shi, Y., Liu, L., Yu, X., Li, H., 2019. Spatial-aware feature aggregation for image based cross-view geo-localization. Advances in Neural Information Processing Systems 32.

Sinnott, R. W., 1984. Virtues of the haversine. Sky and Telescope 68 (2), 159.

Wang, S., Zhang, Y., Perincherry, A., Vora, A., Li, H., 2023. View Consistent Purification for Accurate Cross-View Localization. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 8197–8206. DOI: 10.1109/ICCV51070.2023.00753

Xia, Z., Booij, O., Kooij, J. F., 2023. Convolutional Cross-View Pose Estimation. IEEE Transactions on Pattern Analysis and Machine Intelligence 46 (5), 3813–3831. DOI: 10.1109/TPAMI.2023.3346924

Zhang, X., Wang, L., Su, Y., 2021. Visual place recognition: A survey from deep learning perspective. Pattern Recognition 113, 107760. DOI: https://doi.org/10.1016/j.patcog.2020.107760

Downloads

Published

2026-09-01

Issue

Section

Visión por Computador