Aprendizaje Temporal de Recompensas a partir de Expertos para Aprendizaje por Refuerzo Inverso

Autores/as

DOI:

https://doi.org/10.17979/ja-cea.2026.47.13710

Palabras clave:

Aprendizaje por refuerzo inverso, Aprendizaje por imitación, Aprendizaje de recompensa, Manipulación robótica, Aprendizaje por refuerzo profundo

Resumen

El aprendizaje de funciones de recompensa a partir de demostraciones expertas elimina la necesidad de diseñar manualmente la recompensa en el aprendizaje por refuerzo aplicado a la manipulación robótica. Sin embargo, los métodos existentes requieren anotaciones de calidad sobre las trayectorias, etiquetas de éxito de episodio, o producen recompensas implícitas difíciles de inspeccionar. Este artículo presenta TEXB-IRL (Temporal EXpert-Based reward learning for Inverse Reinforcement Learning), que aprende una función de recompensa neuronal densa directamente a partir de demostraciones expertas sin etiquetar. TEXB-IRL explota el orden temporal intratrayectoria mediante dos objetivos complementarios: una pérdida de consistencia temporal que impone una recompensa monótonamente creciente a lo largo de las trayectorias expertas, y una pérdida de separación experto-agente que ancla la escala de la recompensa. La política se optimiza con PPO sobre la recompensa aprendida. Experimentos preliminares en simulación con el robot PAL TIAGo++ muestran resultados competitivos frente algortimos del estado del arte.

Referencias

Abbeel, P., & Ng, A. Y. (2004). Apprenticeship learning via inverse reinforcement learning. In Proceedings of the 21st International Conference on Machine Learning (ICML). ACM.

Brown, D. S., Coleman, R., Srinivasan, R., & Niekum, S. (2020). Safe imitation learning via fast Bayesian reward inference from preferences. In Proceedings of the 37th International Conference on Machine Learning (ICML) (pp. 1165–1177).

Brown, D. S., Goo, W., Narayanan, P., & Niekum, S. (2019). Extrapolating beyond suboptimal demonstrations in inverse reinforcement learning from observations. In Proceedings of the 36th International Conference on Machine Learning (ICML) (pp. 783–792).

Finn, C., Levine, S., & Abbeel, P. (2016). Guided cost learning: Deep inverse optimal control via policy optimization. In Proceedings of the 33rd International Conference on Machine Learning (ICML) (pp. 49–58).

Fu, J., Luo, K., & Levine, S. (2018). Learning robust rewards with adversarial inverse reinforcement learning. In International Conference on Learning Representations (ICLR).

Ho, J., & Ermon, S. (2016). Generative adversarial imitation learning. In Advances in Neural Information Processing Systems (NeurIPS) (Vol. 29).

Li, Y., Gao, Y., Yang, N., & Xia, S. (2026). TW-CRL: Time-weighted contrastive reward learning for efficient inverse reinforcement learning. Proceedings of the AAAI Conference on Artificial Intelligence, 40(28), 23301–23309. DOI: 10.1609/aaai.v40i28.39499

Naranjo-Campos, F. J., et al. (2024). Expert-trajectory-based features for apprenticeship learning via inverse reinforcement learning. Applied Sciences, 14(24), 11131. DOI: 10.3390/app142411131

Ng, A. Y., & Russell, S. J. (2000). Algorithms for inverse reinforcement learning. In Proceedings of the 17th International Conference on Machine Learning (ICML) (pp. 663–670). Morgan Kaufmann.

Ravichandar, H., Polydoros, A. S., Chernova, S., & Billard, A. (2020). Recent advances in robot learning from demonstration. Annual Review of Control, Robotics, and Autonomous Systems, 3, 297–330. DOI: 10.1146/annurev-control-100819-063206

Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347.

Ziebart, B. D., Maas, A., Bagnell, J. A., & Dey, A. K. (2008). Maximum entropy inverse reinforcement learning. In Proceedings of the 23rd AAAI Conference on Artificial Intelligence (pp. 1433–1438).

Descargas

Publicado

01-09-2026

Número

Sección

Robótica