Long-Tail Trajectory Prediction with Vision-Language-Action Models and Semantic Segmentation
DOI:
https://doi.org/10.17979/ja-cea.2026.47.13800Keywords:
Autonomous vehicles, Trajectory and Path Planning, Sensor integration and perception, Learning and adaptation in autonomous vehicles, Machine learning, Intelligent transportation systemsAbstract
Autonomous driving still fails in rare, safety-critical situations, the so-called long tail, where generic models have little data coverage. This work addresses instruction-conditioned trajectory prediction through two families of methods, evaluated with the Multi-Maneuver Score (MMS), which rewards following the instruction across several valid futures rather than proximity to a single reference. The first prompts a general-purpose visual language model (VLM) to emit waypoints via structured reasoning, constraining it to the road with semantic segmentation (SAM3) and deterministic repair. The second adopts a driving-specific vision--language--action (VLA) model with multi-hypothesis selection guided by a surrogate of the official metric and a conservative road-feasibility tie-break. The VLA approach raises the MMS from 4.34 to 5.15, nearly halving the L2 error, and selection increases it to 5.52, the highest score on the challenge's public leaderboard at the time of writing.
References
Caesar, H., Bankiti, V., Lang, A. H., Vora, S., Liong, V. E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., Beijbom, O., 2020. nuscenes: A multimodal dataset for autonomous driving. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 11621–11631. DOI: 10.1109/CVPR42600.2020.01164
Carion, N., Gustafson, L., Hu, Y.-T., Debnath, S., Hu, R., Suris, D., Ryali, C., Alwala, K. V., Khedr, H., Huang, A., et al., 2025. Sam 3: Segment anything with concepts. arXiv preprint arXiv:2511.16719. DOI: 10.48550/arXiv.2511.16719
Choudhary, T., Dewangan, V., Chandhok, S., Priyadarshan, S., Jain, A., Singh, A. K., Srivastava, S., Jatavallabhula, K. M., Krishna, K. M., 2024. Talk2bev: Language-enhanced bird’s-eye view maps for autonomous driving. In: 2024 IEEE International Conference on Robotics and Automation (ICRA). IEEE, pp. 16345–16352. DOI: 10.1109/ICRA57147.2024.10611485
Ettinger, S., Cheng, S., Caine, B., Liu, C., Zhao, H., Pradhan, S., Chai, Y., Sapp, B., Qi, C. R., Zhou, Y., et al., 2021. Large scale interactive motion forecasting for autonomous driving: The waymo open motion dataset. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 9710–9719. DOI: 10.1109/ICCV48922.2021.00957
Google DeepMind, 2026. Gemma 4 model card. https://ai.google.dev/gemma/docs/core/model_card_4, accessed: 2026-05-27.
Hu, Y., Yang, J., Chen, L., Li, K., Sima, C., Zhu, X., Chai, S., Du, S., Lin, T., Wang, W., et al., 2023. Planning-oriented autonomous driving. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 17853–17862. DOI: 10.1109/CVPR52729.2023.01712
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., et al., 2023. Segment anything. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 4015–4026. DOI: 10.1109/ICCV51070.2023.00371
Mao, J., Qian, Y., Ye, J., Zhao, H., Wang, Y., 2023. Gpt-driver: Learning to drive with gpt. arXiv preprint arXiv:2310.01415. DOI: 10.48550/arXiv.2310.01415
Shao, H., Hu, Y., Wang, L., Song, G., Waslander, S. L., Liu, Y., Li, H., 2024. Lmdrive: Closed-loop end-to-end driving with large language models. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 15120–15130. DOI: 10.1109/CVPR52733.2024.01432
Sima, C., Renz, K., Chitta, K., Chen, L., Zhang, H., Xie, C., Beißwenger, J., Luo, P., Geiger, A., Li, H., 2024. Drivelm: Driving with graph visual question answering. In: European conference on computer vision. Springer, pp. 256–274. DOI: 10.1007/978-3-031-72943-0 15
Tian, X., Gu, J., Li, B., Liu, Y., Wang, Y., Zhao, Z., Zhan, K., Jia, P., Lang, X., Zhao, H., 2024. Drivevlm: The convergence of autonomous driving and large vision-language models. arXiv preprint arXiv:2402.12289. DOI: 10.48550/arXiv.2402.12289
Wagner, R., Tas, O. S., Villa, J., Hauser, F., Shen, Y., Steiner, M., Strutz, D., Fernandez, C., Kinzig, C., Guitierrez-Cabello, G. S., Königshof, H., Immel, F., Schwarzkopf, R., Rack, N. A., Rösch, K., Wang, K., Pauls, J.-H., Lauer, M., Gilitschenski, I., Caesar, H., Stiller, C., 2026. Longtail driving scenarios with reasoning traces: The kitscenes longtail dataset. URL: https://arxiv.org/abs/2603.23607 DOI: 10.48550/arXiv.2603.23607
Wang, Y., Luo, W., Bai, J., Cao, Y., Che, T., Chen, K., Chen, Y., Diamond, J., Ding, Y., Ding, W., et al., 2025. Alpamayo-r1: Bridging reasoning and action prediction for generalizable autonomous driving in the long tail. arXiv preprint arXiv:2511.00088. DOI: 10.48550/arXiv.2511.00088
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al., 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35, 24824–24837. DOI: 10.52202/068431-1800
Wilson, B., Qi, W., Agarwal, T., Lambert, J., Singh, J., Khandelwal, S., Pan, B., Kumar, R., Hartnett, A., Pontes, J. K., et al., 2023. Argoverse 2: Next generation datasets for self-driving perception and forecasting. arXiv preprint arXiv:2301.00493. DOI: 10.48550/arXiv.2301.00493
Xu, Z., Zhang, Y., Xie, E., Zhao, Z., Guo, Y., Wong, K.-Y. K., Li, Z., Zhao, H., 2024. Drivegpt4: Interpretable end-to-end autonomous driving via large language model. IEEE Robotics and Automation Letters 9 (10), 8186–8193. DOI: 10.1109/LRA.2024.3440097
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Javier Borau Bernad, Diego Caballero García-Alcaide, Raúl Fernández Matellán, Martín Santiago Soto, José María Armingol Moreno, Araceli Sanchis de Miguel

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.