Predicción de trayectorias en cola larga con visión-lenguaje-acción y segmentación

Autores/as

DOI:

https://doi.org/10.17979/ja-cea.2026.47.13800

Palabras clave:

Vehículos autónomos, Integración de sensores y percepción, Aprendizaje y adaptación en vehículos autónomos, Aprendizaje automático, Sistemas inteligentes de transporte, Predicción de trayectorias y planificación de rutas

Resumen

La conducción autónoma sigue fallando en situaciones raras y críticas, la llamada cola larga, donde los modelos genéricos tienen poca cobertura de datos. Este trabajo aborda la predicción de trayectorias condicionada por instrucciones mediante dos familias de métodos, evaluadas con la puntuación multi-maniobra (Multi-Maneuver Score, MMS), que premia el seguimiento de la instrucción a través de varios futuros válidos, no la proximidad a una referencia. La primera induce a un modelo de lenguaje visual (VLM) general a emitir waypoints con un prompt de razonamiento estructurado, restringiéndolo a la calzada mediante la segmentación semántica (SAM3) y una reparación determinista. La segunda adopta un modelo visión--lenguaje--acción (VLA) específico de conducción, con una selección multi-hipótesis guiada por un sustituto de la métrica y un desempate conservador por viabilidad de calzada. El enfoque VLA eleva la MMS de 4,34 a 5,15 y, con la selección, hasta 5,52, la más alta de la clasificación pública a fecha de redacción.

Referencias

Caesar, H., Bankiti, V., Lang, A. H., Vora, S., Liong, V. E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., Beijbom, O., 2020. nuscenes: A multimodal dataset for autonomous driving. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 11621–11631. DOI: 10.1109/CVPR42600.2020.01164

Carion, N., Gustafson, L., Hu, Y.-T., Debnath, S., Hu, R., Suris, D., Ryali, C., Alwala, K. V., Khedr, H., Huang, A., et al., 2025. Sam 3: Segment anything with concepts. arXiv preprint arXiv:2511.16719. DOI: 10.48550/arXiv.2511.16719

Choudhary, T., Dewangan, V., Chandhok, S., Priyadarshan, S., Jain, A., Singh, A. K., Srivastava, S., Jatavallabhula, K. M., Krishna, K. M., 2024. Talk2bev: Language-enhanced bird’s-eye view maps for autonomous driving. In: 2024 IEEE International Conference on Robotics and Automation (ICRA). IEEE, pp. 16345–16352. DOI: 10.1109/ICRA57147.2024.10611485

Ettinger, S., Cheng, S., Caine, B., Liu, C., Zhao, H., Pradhan, S., Chai, Y., Sapp, B., Qi, C. R., Zhou, Y., et al., 2021. Large scale interactive motion forecasting for autonomous driving: The waymo open motion dataset. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 9710–9719. DOI: 10.1109/ICCV48922.2021.00957

Google DeepMind, 2026. Gemma 4 model card. https://ai.google.dev/gemma/docs/core/model_card_4, accessed: 2026-05-27.

Hu, Y., Yang, J., Chen, L., Li, K., Sima, C., Zhu, X., Chai, S., Du, S., Lin, T., Wang, W., et al., 2023. Planning-oriented autonomous driving. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 17853–17862. DOI: 10.1109/CVPR52729.2023.01712

Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., et al., 2023. Segment anything. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 4015–4026. DOI: 10.1109/ICCV51070.2023.00371

Mao, J., Qian, Y., Ye, J., Zhao, H., Wang, Y., 2023. Gpt-driver: Learning to drive with gpt. arXiv preprint arXiv:2310.01415. DOI: 10.48550/arXiv.2310.01415

Shao, H., Hu, Y., Wang, L., Song, G., Waslander, S. L., Liu, Y., Li, H., 2024. Lmdrive: Closed-loop end-to-end driving with large language models. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 15120–15130. DOI: 10.1109/CVPR52733.2024.01432

Sima, C., Renz, K., Chitta, K., Chen, L., Zhang, H., Xie, C., Beißwenger, J., Luo, P., Geiger, A., Li, H., 2024. Drivelm: Driving with graph visual question answering. In: European conference on computer vision. Springer, pp. 256–274. DOI: 10.1007/978-3-031-72943-0 15

Tian, X., Gu, J., Li, B., Liu, Y., Wang, Y., Zhao, Z., Zhan, K., Jia, P., Lang, X., Zhao, H., 2024. Drivevlm: The convergence of autonomous driving and large vision-language models. arXiv preprint arXiv:2402.12289. DOI: 10.48550/arXiv.2402.12289

Wagner, R., Tas, O. S., Villa, J., Hauser, F., Shen, Y., Steiner, M., Strutz, D., Fernandez, C., Kinzig, C., Guitierrez-Cabello, G. S., Königshof, H., Immel, F., Schwarzkopf, R., Rack, N. A., Rösch, K., Wang, K., Pauls, J.-H., Lauer, M., Gilitschenski, I., Caesar, H., Stiller, C., 2026. Longtail driving scenarios with reasoning traces: The kitscenes longtail dataset. URL: https://arxiv.org/abs/2603.23607 DOI: 10.48550/arXiv.2603.23607

Wang, Y., Luo, W., Bai, J., Cao, Y., Che, T., Chen, K., Chen, Y., Diamond, J., Ding, Y., Ding, W., et al., 2025. Alpamayo-r1: Bridging reasoning and action prediction for generalizable autonomous driving in the long tail. arXiv preprint arXiv:2511.00088. DOI: 10.48550/arXiv.2511.00088

Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al., 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35, 24824–24837. DOI: 10.52202/068431-1800

Wilson, B., Qi, W., Agarwal, T., Lambert, J., Singh, J., Khandelwal, S., Pan, B., Kumar, R., Hartnett, A., Pontes, J. K., et al., 2023. Argoverse 2: Next generation datasets for self-driving perception and forecasting. arXiv preprint arXiv:2301.00493. DOI: 10.48550/arXiv.2301.00493

Xu, Z., Zhang, Y., Xie, E., Zhao, Z., Guo, Y., Wong, K.-Y. K., Li, Z., Zhao, H., 2024. Drivegpt4: Interpretable end-to-end autonomous driving via large language model. IEEE Robotics and Automation Letters 9 (10), 8186–8193. DOI: 10.1109/LRA.2024.3440097

Descargas

Publicado

01-09-2026

Número

Sección

Visión por Computador