Temporal Expert-Based Reward Learning for Inverse Reinforcement Learning

Authors

DOI:

https://doi.org/10.17979/ja-cea.2026.47.13710

Keywords:

Imitation learning, Inverse reinforcement learning, Reward learning, Robotic manipulation, Deep reinforcement learning

Abstract

Learning reward functions from expert demonstrations removes the need for manual reward engineering in reinforcement learning applied to robotic manipulation. Existing methods, however, require trajectory quality annotations, episode success labels, or produce implicit rewards that are difficult to inspect. This paper presents TEXB-IRL (Temporal EXpert-Based reward learning for Inverse Reinforcement Learning), which learns a dense neural reward function directly from unlabeled expert demonstrations. TEXB-IRL exploits intra-trajectory temporal ordering through two complementary objectives: a temporal consistency loss that enforces monotonically increasing reward along expert trajectories, and an expert-agent separation loss that anchors the reward scale. The policy is optimized with PPO over the learned reward. Preliminary experiments in simulation with the PAL TIAGo++ robot show competitive results against state of the art algorithms.

References

Abbeel, P., & Ng, A. Y. (2004). Apprenticeship learning via inverse reinforcement learning. In Proceedings of the 21st International Conference on Machine Learning (ICML). ACM.

Brown, D. S., Coleman, R., Srinivasan, R., & Niekum, S. (2020). Safe imitation learning via fast Bayesian reward inference from preferences. In Proceedings of the 37th International Conference on Machine Learning (ICML) (pp. 1165–1177).

Brown, D. S., Goo, W., Narayanan, P., & Niekum, S. (2019). Extrapolating beyond suboptimal demonstrations in inverse reinforcement learning from observations. In Proceedings of the 36th International Conference on Machine Learning (ICML) (pp. 783–792).

Finn, C., Levine, S., & Abbeel, P. (2016). Guided cost learning: Deep inverse optimal control via policy optimization. In Proceedings of the 33rd International Conference on Machine Learning (ICML) (pp. 49–58).

Fu, J., Luo, K., & Levine, S. (2018). Learning robust rewards with adversarial inverse reinforcement learning. In International Conference on Learning Representations (ICLR).

Ho, J., & Ermon, S. (2016). Generative adversarial imitation learning. In Advances in Neural Information Processing Systems (NeurIPS) (Vol. 29).

Li, Y., Gao, Y., Yang, N., & Xia, S. (2026). TW-CRL: Time-weighted contrastive reward learning for efficient inverse reinforcement learning. Proceedings of the AAAI Conference on Artificial Intelligence, 40(28), 23301–23309. DOI: 10.1609/aaai.v40i28.39499

Naranjo-Campos, F. J., et al. (2024). Expert-trajectory-based features for apprenticeship learning via inverse reinforcement learning. Applied Sciences, 14(24), 11131. DOI: 10.3390/app142411131

Ng, A. Y., & Russell, S. J. (2000). Algorithms for inverse reinforcement learning. In Proceedings of the 17th International Conference on Machine Learning (ICML) (pp. 663–670). Morgan Kaufmann.

Ravichandar, H., Polydoros, A. S., Chernova, S., & Billard, A. (2020). Recent advances in robot learning from demonstration. Annual Review of Control, Robotics, and Autonomous Systems, 3, 297–330. DOI: 10.1146/annurev-control-100819-063206

Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347.

Ziebart, B. D., Maas, A., Bagnell, J. A., & Dey, A. K. (2008). Maximum entropy inverse reinforcement learning. In Proceedings of the 23rd AAAI Conference on Artificial Intelligence (pp. 1433–1438).

Downloads

Published

2026-09-01

Issue

Section

Robótica