A platform for the navigation of Unmanned Aerial Vehiclesbased on Reinforcement Learning

Authors

DOI:

https://doi.org/10.17979/ja-cea.2026.47.13789

Keywords:

Reinforcement learning control, Autonomous Vehicles, Platform for learning and validation

Abstract

This article introduces a platform for developing use cases for the automated control of unmanned aerial vehicles (UAVs), utilising the AirSim simulator. This platform allows for the generation of realistic flight scenarios involving multiple UAVs.
The proposed platform facilitates the construction of use cases for the development, validation, and verification of reinforcement learning (RL) and neural networks in critical real-time systems. These algorithms can learn to navigate dynamic environments without human intervention. The platform was validated using two UAVs with neural networks that were trained using the Soft Actor-Critical (SAC) algorithm.

References

DLR, 2026. Stable baselines3, institute of robotics and mechatronics. URL: https://www.dlr.de/en/rm/research/publications-and-downloads/software/stable-baselines3.

ECSS, 2017a. European Cooperation for Space Standardization. ECSS-Q-ST-40C Space product assurance — Safety.

ECSS, 2017b. European Cooperation for Space Standardization. ECSS-Q-ST-80C Rev.1 – Space Product Assurance — Software Product Assurance.

ECSS, 2021. European Cooperation for Space Standardization. ECSS-E-ST-40C Rev.1, Dir.1, Space engineering — Software.

Haarnoja, T., Zhou, A., Abbeel, P., Levine, S., 2018. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. CoRR abs/1801.01290. URL: http://arxiv.org/abs/1801.01290, doi:10.48550/arXiv.1801.01290, arXiv:1801.01290.

Pérez-Muñoz, A.G., López-García, G., García-Quijano, H., Alonso, A., 2025a. Development and validation of a safe reinforcement learning drone controller. Jornadas de Autom´ atica 10.17979/ja-cea.2025.46.12154.

Pérez-Muñoz, A.G., López-García,., García-Villoria, I., Alonso, A., Porras-Hermoso, A. , Pérez, M.S., 2025b. Feasibility of deep reinforcement learning for the real-time attitude control of a satellite system. Journal of Systems Architecture 167, 103513. URL: https://www.sciencedirect.com/science/article/pii/S1383762125001857, doi:https://doi.org/10.1016/j.sysarc.2025.103513.

Rierson, L., 2013. Developing safety-critical software. CRC Press, Boca Raton, FL.

Sutton, R.S., Barto, A.G., 2018. Reinforcement Learning: An Introduction, 2nd edition. A Bradford Book, Cambridge, MA, USA.

Downloads

Published

2026-09-01

Issue

Section

Computadores y Control