Análisis del tiempo de paso en el aprendizaje por refuerzo para robots móviles
DOI:
https://doi.org/10.17979/ja-cea.2026.47.13780Palabras clave:
Aprendizaje por Refuerzo Profundo, Duración del Tiempo de Paso, Robots Móviles, Brecha de Simulación a Realidad, Arquitecturas Asíncronas, Eficiencia de Muestreo, Dinámica Temporal, Robótica InteligenteResumen
El uso del Aprendizaje por Refuerzo Profundo (DRL) en robótica ha experimentado un gran crecimiento en la última decada, pero el impacto de sus aspectos temporales al aplicarse a sistemas físicos -principalmente la duración del tiempo de paso (timestep)- sigue poco explorado. Aunque algunos enfoques lo incluyen en el espacio de acción, ofrecen un análisis limitado, omitiendo elementos clave como la asincronía del software o las latencias del sistema. Estos factores causan desviaciones entre los tiempos de paso nominales y reales, perjudicando el aprendizaje. Este artículo propone tareas de navegación para robots móviles diseñadas para analizar los efectos de la duración del tiempo de paso, respaldadas por más de 2000 horas de simulación. Hasta donde sabemos, esta es la primera evaluación multimétrica de la influencia del tiempo de paso en robótica basada en DRL. Los resultados confirman que existen tiempos de paso óptimos que maximizan la eficiencia del aprendizaje y el éxito en la tarea.
Referencias
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viégas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., Zheng, X., 2015. TensorFlow: Large-scale machine learning on heterogeneous systems. Software available from tensorflow.org. URL: https://www.tensorflow.org/ .
Chen, X., Wang, J., 2022. Inhomogeneous deep q-network for time sensitive applications. Artificial Intelligence 312, 103757.
de Jesus, J. C., Kich, V. A., Kolling, A. H., Grando, R. B., de Souza Leite Cuadros, M. A., Gamarra, D. F. T., 2021. Soft actor-critic for navigation of mobile robots. Journal of Intelligent & Robotic Systems 102.
Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., Abbeel, P., Levine, S., 2019. Soft actor-critic algorithms and applications.
Ibarz, J., Tan, J., Finn, C., Kalakrishnan, M., Pastor, P., Levine, S., 2021. How to train your robot with deep reinforcement learning: lessons we have learned. The International Journal of Robotics Research 40, 698 – 721.
IMECH.UMA, 2025. University Research Institute in Mechatronics Engineering and Cyber-Physical Systems, University of Málaga,Spain. https:
//www.imech-uma.es/, accessed on July 14, 2025.
Kuo, P.-H., Huang, C.-T., Chang, C.-W., Feng, P.-H., Lin, Y.-S., 2025. Design and implementation of a soft actor–critic controller for a robotic arm. Engineering Applications of Artificial Intelligence 151, 110589.
Le, H., Saeedvand, S., Hsu, C.-C., 2024. A comprehensive review of mobile robot navigation using deep reinforcement learning algorithms in crowded environments. Journal of Intelligent & Robotic Systems 110 (4), 158.
Lee, J., 2013. Turtlebot 2: The new standard hardware reference platform. In: ROSCon 2013. Open Robotics. URL: http://dx.doi.org/10.36288/roscon2013-899053
Mahmood, A. R., Korenkevych, D., Komer, B., Bergstra, J., 2018. Setting up a reinforcement learning task with a real-world robot. 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 4635–4640.
Puterman, M. L., 2014. Markov Decision Processes: Discrete Stochastic Dynamic Programming, 1st Edition. John Wiley & Sons.
Raffin, A., Hill, A., Gleave, A., Kanervisto, A., Ernestus, M., Dormann, N., 2021. Stable-baselines3: Reliable reinforcement learning implementations. Journal of Machine Learning Research 22 (268), 1–8.
Robotics, C., 2024. Coppeliasim. https://www.coppeliarobotics.com/ , accessed: 2025-05-08.
Salvato, E., Fenu, G., Medvet, E., Pellegrino, F. A., 2021. Crossing the reality gap: a survey on sim-to-real transferability of robot controllers in reinforcement learning. IEEE Access PP, 1–1.
Towers, M., Kwiatkowski, A., Terry, J., Balis, J. U., De Cola, G., Deleu, T., Goulão, M., Kallinteris, A., Krimmel, M., KG, A., et al., 2024. Gymnasium: A standard interface for reinforcement learning environments. arXiv preprint arXiv:2407.17032.
UNCORE-Team, 2025. rl spin decoupler: A library for decoupling agent and environment control loops in RL. https://github.com/uncore-team/rl_spin_decoupler/tree/main , accessed: 2025-05-13.
Wang, D., Beltrame, G., 2024. Variable time step reinforcement learning for robotic applications.
Yuan, Y., Mahmood, A. R., 2022. Asynchronous reinforcement learning for real-time control of physical robots.
Zhou, H., Huang, A., Azizzadenesheli, K., Childers, D., Lipton, Z., 2024. Timing as an action: Learning when to observe and act. In: Dasgupta, S., Mandt, S., Li, Y. (Eds.), Proceedings of The 27th International Conference on Artificial Intelligence and Statistics. Vol. 238 of Proceedings of Machine Learning Research. PMLR, pp. 3979–3987.
Descargas
Publicado
Número
Sección
Licencia
Derechos de autor 2026 Adrián Bañuls Arias, Juan-Antonio Fernández-Madrigal, Ana Cruz-Martín, Vicente Arévalo-Espejo, Cipriano Galindo, Juan-Manuel Gandarias

Esta obra está bajo una licencia internacional Creative Commons Atribución-NoComercial-CompartirIgual 4.0.