Predicción de trayectorias en cola larga con visión-lenguaje-acción y segmentación
DOI:
https://doi.org/10.17979/ja-cea.2026.47.13800Palabras clave:
Vehículos autónomos, Integración de sensores y percepción, Aprendizaje y adaptación en vehículos autónomos, Aprendizaje automático, Sistemas inteligentes de transporte, Predicción de trayectorias y planificación de rutasResumen
La conducción autónoma sigue fallando en situaciones raras y críticas, la llamada cola larga, donde los modelos genéricos tienen poca cobertura de datos. Este trabajo aborda la predicción de trayectorias condicionada por instrucciones mediante dos familias de métodos, evaluadas con la puntuación multi-maniobra (Multi-Maneuver Score, MMS), que premia el seguimiento de la instrucción a través de varios futuros válidos, no la proximidad a una referencia. La primera induce a un modelo de lenguaje visual (VLM) general a emitir waypoints con un prompt de razonamiento estructurado, restringiéndolo a la calzada mediante la segmentación semántica (SAM3) y una reparación determinista. La segunda adopta un modelo visión--lenguaje--acción (VLA) específico de conducción, con una selección multi-hipótesis guiada por un sustituto de la métrica y un desempate conservador por viabilidad de calzada. El enfoque VLA eleva la MMS de 4,34 a 5,15 y, con la selección, hasta 5,52, la más alta de la clasificación pública a fecha de redacción.
Referencias
Caesar, H., Bankiti, V., Lang, A. H., Vora, S., Liong, V. E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., Beijbom, O., 2020. nuscenes: A multimodal dataset for autonomous driving. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 11621–11631. DOI: 10.1109/CVPR42600.2020.01164
Carion, N., Gustafson, L., Hu, Y.-T., Debnath, S., Hu, R., Suris, D., Ryali, C., Alwala, K. V., Khedr, H., Huang, A., et al., 2025. Sam 3: Segment anything with concepts. arXiv preprint arXiv:2511.16719. DOI: 10.48550/arXiv.2511.16719
Choudhary, T., Dewangan, V., Chandhok, S., Priyadarshan, S., Jain, A., Singh, A. K., Srivastava, S., Jatavallabhula, K. M., Krishna, K. M., 2024. Talk2bev: Language-enhanced bird’s-eye view maps for autonomous driving. In: 2024 IEEE International Conference on Robotics and Automation (ICRA). IEEE, pp. 16345–16352. DOI: 10.1109/ICRA57147.2024.10611485
Ettinger, S., Cheng, S., Caine, B., Liu, C., Zhao, H., Pradhan, S., Chai, Y., Sapp, B., Qi, C. R., Zhou, Y., et al., 2021. Large scale interactive motion forecasting for autonomous driving: The waymo open motion dataset. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 9710–9719. DOI: 10.1109/ICCV48922.2021.00957
Google DeepMind, 2026. Gemma 4 model card. https://ai.google.dev/gemma/docs/core/model_card_4, accessed: 2026-05-27.
Hu, Y., Yang, J., Chen, L., Li, K., Sima, C., Zhu, X., Chai, S., Du, S., Lin, T., Wang, W., et al., 2023. Planning-oriented autonomous driving. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 17853–17862. DOI: 10.1109/CVPR52729.2023.01712
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., et al., 2023. Segment anything. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 4015–4026. DOI: 10.1109/ICCV51070.2023.00371
Mao, J., Qian, Y., Ye, J., Zhao, H., Wang, Y., 2023. Gpt-driver: Learning to drive with gpt. arXiv preprint arXiv:2310.01415. DOI: 10.48550/arXiv.2310.01415
Shao, H., Hu, Y., Wang, L., Song, G., Waslander, S. L., Liu, Y., Li, H., 2024. Lmdrive: Closed-loop end-to-end driving with large language models. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 15120–15130. DOI: 10.1109/CVPR52733.2024.01432
Sima, C., Renz, K., Chitta, K., Chen, L., Zhang, H., Xie, C., Beißwenger, J., Luo, P., Geiger, A., Li, H., 2024. Drivelm: Driving with graph visual question answering. In: European conference on computer vision. Springer, pp. 256–274. DOI: 10.1007/978-3-031-72943-0 15
Tian, X., Gu, J., Li, B., Liu, Y., Wang, Y., Zhao, Z., Zhan, K., Jia, P., Lang, X., Zhao, H., 2024. Drivevlm: The convergence of autonomous driving and large vision-language models. arXiv preprint arXiv:2402.12289. DOI: 10.48550/arXiv.2402.12289
Wagner, R., Tas, O. S., Villa, J., Hauser, F., Shen, Y., Steiner, M., Strutz, D., Fernandez, C., Kinzig, C., Guitierrez-Cabello, G. S., Königshof, H., Immel, F., Schwarzkopf, R., Rack, N. A., Rösch, K., Wang, K., Pauls, J.-H., Lauer, M., Gilitschenski, I., Caesar, H., Stiller, C., 2026. Longtail driving scenarios with reasoning traces: The kitscenes longtail dataset. URL: https://arxiv.org/abs/2603.23607 DOI: 10.48550/arXiv.2603.23607
Wang, Y., Luo, W., Bai, J., Cao, Y., Che, T., Chen, K., Chen, Y., Diamond, J., Ding, Y., Ding, W., et al., 2025. Alpamayo-r1: Bridging reasoning and action prediction for generalizable autonomous driving in the long tail. arXiv preprint arXiv:2511.00088. DOI: 10.48550/arXiv.2511.00088
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al., 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35, 24824–24837. DOI: 10.52202/068431-1800
Wilson, B., Qi, W., Agarwal, T., Lambert, J., Singh, J., Khandelwal, S., Pan, B., Kumar, R., Hartnett, A., Pontes, J. K., et al., 2023. Argoverse 2: Next generation datasets for self-driving perception and forecasting. arXiv preprint arXiv:2301.00493. DOI: 10.48550/arXiv.2301.00493
Xu, Z., Zhang, Y., Xie, E., Zhao, Z., Guo, Y., Wong, K.-Y. K., Li, Z., Zhao, H., 2024. Drivegpt4: Interpretable end-to-end autonomous driving via large language model. IEEE Robotics and Automation Letters 9 (10), 8186–8193. DOI: 10.1109/LRA.2024.3440097
Descargas
Publicado
Número
Sección
Licencia
Derechos de autor 2026 Javier Borau Bernad, Diego Caballero García-Alcaide, Raúl Fernández Matellán, Martín Santiago Soto, José María Armingol Moreno, Araceli Sanchis de Miguel

Esta obra está bajo una licencia internacional Creative Commons Atribución-NoComercial-CompartirIgual 4.0.