Ascendiendo a la montaña: aprendizaje por refuerzo colaborativo entre agentes con objetivos opuestos

Autores/as

  • Violeta Tejera-Munguía Universidad Autónoma de Madrid
  • Juan Jesús Roldán-Gómez Universidad Autónoma de Madrid
  • José Luis Jorro-Aragoneses Universidad Autónoma de Madrid

DOI:

https://doi.org/10.17979/ja-cea.2026.47.13783

Palabras clave:

Inteligencia Artificial, Aprendizaje por Refuerzo, Optimización

Resumen

Este trabajo explora el concepto de Estupidez Artificial como complemento de la Inteligencia Artificial en el contexto del aprendizaje por refuerzo. Para ello, se realizan una serie de experimentos con el algoritmo Q-Learning en el entorno Mountain Car, comparando un agente inteligente con un agente estúpido que desarrollan comportamientos opuestos. Se proponen dos estrategias de entrenamiento: una reformativa, donde el agente inteligente se entrena a partir del estúpido, y otra colaborativa, donde ambos se entrenan simultáneamente compartiendo información. Los resultados muestran que la incorporación de información procedente del agente estúpido mejora el rendimiento del agente inteligente, acelerando su convergencia y mejorando su rendimiento.

Referencias

Bădică, A., Bădică, C., Ivanović, M., Logofătu, D., 2022. Experiments with solving mountain car problem using state discretization and q-learning. In: Intelligent Information and Database Systems: 14th Asian Conference,

ACIIDS 2022, Ho Chi Minh City, Vietnam, November 28–30, 2022, Proceedings, Part I. Springer-Verlag, Berlin, Heidelberg, p. 142–155. URL: https://doi.org/10.1007/978-3-031-21743-2_12 DOI: 10.1007/978-3-031-21743-212

Climent, L., Longhi, A., Arbelaez, A., Mancini, M., 2024. A framework for designing reinforcement learning agents with dynamic difficulty adjustment in single-player action video games. Entertainment Computing 50, 100686. URL: https://www.sciencedirect.com/science/article/pii/S1875952124000545 DOI: https://doi.org/10.1016/j.entcom.2024.100686

Economist, 1992. Artificial stupidity. The Economist 324 (7770), 14. Grollman, D. H., Billard, A., 2011. Donut as i do: Learning from failed demonstrations. In: 2011 IEEE international conference on robotics and automation. IEEE, pp. 3804–3809.

Lidén, L., 2003. Artificial stupidity: the art of intentional mistakes. Vol. 2. Charles River Media.

Moore, A.W., 1990. Efficient memory-based learning for robot control. Tech. rep., University of Cambridge.

Roldán-Gómez, J. J., 2022. Artificial Stupidity in Robotics: Something Unwanted or Somehow Useful? Lecture notes in networks and systems, 26–37. URL: https://doi.org/10.1007/978-3-031-21062-4_3DOI: 10.1007/978-3-031-21062-4{3

Sutton, R. S., Barto, A. G., 11 2018. Reinforcement Learning, second edition. MIT Press.

Teja Chavali, S., Tej Kandavalli, C., Sugash, T. M., Amudha, J., 2022. Modelling a reinforcement learning agent for mountain car problem using q–learning with tabular discretization. In: 2022 IEEE 2nd Mysore Sub Section International Conference (MysuruCon). pp. 1–5. DOI: 10.1109/MysuruCon55714.2022.9972352

Watkins, C. J. C. H., Dayan, P., 5 1992. Q-learning. Vol. 8. URL: https://doi.org/10.1007/bf00992698 OI: 10.1007/bf00992698

Descargas

Publicado

01-09-2026

Número

Sección

Robótica