Una aproximación preliminar al despliegue asíncrono de software de aprendizaje por refuerzo en sistemas físicos

Autores/as

DOI:

https://doi.org/10.17979/ja-cea.2026.47.13751

Palabras clave:

Manipuladores robóticos, Tiempo de paso variable, Despliegue asíncrono, Arquitectura de software distribuida, Aprendizaje por Refuerzo

Resumen

El aprendizaje por refuerzo (RL) suele asumir tiempos de paso fijos en el bucle de control, ignorando las discrepancias temporales entre simulación y realidad, así como el efecto de las latencias existentes en dicho software. Esto puede degradar el rendimiento al transferir la simulación a la realidad (sim-to-real gap). Este trabajo presenta una solución desacoplada que permite una ejecución de RL asíncrona, agnóstica al algoritmo y con tratamiento explícito del tiempo. Al aislar el algoritmo del agente en bucles independientes, se superan las limitaciones de los marcos de desarrollo síncronos tradicionales. Esto facilita la integración con robots físicos y simuladores, permitiendo implantar estrategias de RL con tiempo de paso variable. Nuestra herramienta es compatible con bibliotecas como Stable-Baselines3, simuladores como MuJoCo o CoppeliaSim y robots como el Panda de Franka Emika. Una validación experimental preliminar ha demostrado su efectividad con dicho robot manipulador, y la versión inicial del software ha sido puesta a disposición de la comunidad en GitHub: https://github.com/uncore-team/rl_spin_decoupler.

Referencias

Al Mahmud, S., Kamarulariffin, A., Ibrahim, A. M., Mohideen, A. J. H., 2024. Advancements and challenges in mobile robot navigation: A comprehensive review of algorithms and potential for self-learning approaches. Journal of Intelligent & Robotic Systems 110 (3), 120. URL: https://doi.org/10.1007/s10846-024-02149-5 DOI: 10.1007/s10846-024-02149-5

Dulac-Arnold, G., Mankowitz, D. J., Hester, T., 2019. Challenges of real-world reinforcement learning. ArXiv abs/1904.12901.

Elguea-Aguinaco, Í., Serrano-Muñoz, A., Chrysostomou, D., Inziarte-Hidalgo, I., Bøgh, S., Arana-Arexolaleiba, N., 2023. A review on reinforcement learning for contact-rich robotic manipulation tasks. Robotics and Computer-Integrated Manufacturing 81, 102517.

Elsner, J., 2023. Taming the panda with python: A powerful duo for seamless robotics programming and integration. SoftwareX 24, 101532. URL: https://www.sciencedirect.com/science/article/pii/S2352711023002285 DOI: 10.1016/j.softx.2023.101532

Fernández-Madrigal, J.-A., Navarro, A., Asenjo, R., Cruz-Martín, A., 2020. Characterization, statistical analysis and method selection in the two-clocks synchronization problem for pairwise interconnected sensors. Sensors 20 (17). URL: https://www.mdpi.com/1424-8220/20/17/4808 DOI: 10.3390/s20174808

Gamma, E., Helm, R., Johnson, R., Vlissides, J., 1994. Design Patterns: Elements of Reusable Object-Oriented Software. Addison-Wesley, Reading, MA.

Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., Abbeel, P., Levine, S., 2019. Soft actor-critic algorithms and applications. URL: https://arxiv.org/abs/1812.05905

Haddadin, S., 2024. The franka emika robot: A standard platform in robotics research. IEEE Robotics & Automation Magazine.

Ibarz, J., Tan, J., Finn, C., Kalakrishnan, M., Pastor, P., Levine, S., 2021. How to train your robot with deep reinforcement learning: lessons we have learned. The International Journal of Robotics Research 40, 698 – 721.

IMECH.UMA, 2025. Instituto Universitario de Investigación en Ingeniería Mecatrónica y Sistemas Ciberfísicos, Universidad de Málaga, España. https://www.imech-uma.es/, accedido el 14 de julio de 2025.

Le, H., Saeedvand, S., Hsu, C.-C., 2024. A comprehensive review of mobile robot navigation using deep reinforcement learning algorithms in crowded environments. Journal of Intelligent & Robotic Systems 110 (4), 158. URL: https://doi.org/10.1007/s10846-024-02198-w DOI: 10.1007/s10846-024-02198-w

Mahmood, A. R., Korenkevych, D., Komer, B., Bergstra, J., 2018. Setting up a reinforcement learning task with a real-world robot. 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 4635–4640. URL: https://api.semanticscholar.org/CorpusID:3971262

Postel, J., 1981. Transmission control protocol. RFC 793, accessed: 2025-05-22. URL: https://www.rfc-editor.org/rfc/rfc793

Prasuna, R. G., Potturu, S. R., 2024. Deep reinforcement learning in mobile robotics – a concise review. Multimedia Tools and Applications 83 (28), 70815–70836.

Puterman, M. L., 2014. Markov Decision Processes: Discrete Stochastic Dynamic Programming, 1st Edition. John Wiley & Sons.

Raffin, A., Hill, A., Gleave, A., Kanervisto, A., Ernestus, M., Dormann, N., 2021. Stable-baselines3: Reliable reinforcement learning implementations. Journal of Machine Learning Research 22 (268), 1–8. URL: http://jmlr.org/papers/v22/20-1364.html

Robotics, C., 2024. Coppeliasim. https://www.coppeliarobotics.com/, accessed: 2025-05-08.

Salvato, E., Fenu, G., Medvet, E., Pellegrino, F. A., 2021. Crossing the reality gap: a survey on sim-to-real transferability of robot controllers in reinforcement learning. IEEE Access PP, 1–1. URL: https://api.semanticscholar.org/CorpusID:243882665

Todorov, E., Erez, T., Tassa, Y., 2012. Mujoco: A physics engine for model-based control. In: IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). pp. 5026–5033. DOI: 10.1109/IROS.2012.6386109

Towers, M., Kwiatkowski, A., Terry, J., Balis, J. U., De Cola, G., Deleu, T., Goulão, M., Kallinteris, A., Krimmel, M., KG, A., et al., 2024. Gymnasium: A standard interface for reinforcement learning environments. arXiv preprint arXiv:2407.17032.

Van Rossum, G., Drake, F. L., 2009. Python 3 Reference Manual. CreateSpace, Scotts Valley, CA.

Yuan, Y., Mahmood, A. R., 2022. Asynchronous reinforcement learning for real-time control of physical robots. URL: https://arxiv.org/abs/2203.12759

Zhao, W., Queralta, J. P., Westerlund, T., 2020. Sim-to-real transfer in deep reinforcement learning for robotics: a survey. 2020 IEEE Symposium Series on Computational Intelligence (SSCI), 737–744. URL: https://api.semanticscholar.org/CorpusID:221971078

Zhu, K., Zhang, T., 2021. Deep reinforcement learning based mobile robot navigation: A review. Tsinghua Science and Technology 26 (5), 674–691. DOI: 10.26599/TST.2021.9010012

Descargas

Publicado

01-09-2026

Número

Sección

Control Inteligente