Long-Tail Trajectory Prediction with Vision-Language-Action Models and Semantic Segmentation

Authors

DOI:

https://doi.org/10.17979/ja-cea.2026.47.13800

Keywords:

Autonomous vehicles, Trajectory and Path Planning, Sensor integration and perception, Learning and adaptation in autonomous vehicles, Machine learning, Intelligent transportation systems

Abstract

Autonomous driving still fails in rare, safety-critical situations, the so-called long tail, where generic models have little data coverage. This work addresses instruction-conditioned trajectory prediction through two families of methods, evaluated with the Multi-Maneuver Score (MMS), which rewards following the instruction across several valid futures rather than proximity to a single reference. The first prompts a general-purpose visual language model (VLM) to emit waypoints via structured reasoning, constraining it to the road with semantic segmentation (SAM3) and deterministic repair. The second adopts a driving-specific vision--language--action (VLA) model with multi-hypothesis selection guided by a surrogate of the official metric and a conservative road-feasibility tie-break. The VLA approach raises the MMS from 4.34 to 5.15, nearly halving the L2 error, and selection increases it to 5.52, the highest score on the challenge's public leaderboard at the time of writing.

References

Caesar, H., Bankiti, V., Lang, A. H., Vora, S., Liong, V. E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., Beijbom, O., 2020. nuscenes: A multimodal dataset for autonomous driving. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 11621–11631. DOI: 10.1109/CVPR42600.2020.01164

Carion, N., Gustafson, L., Hu, Y.-T., Debnath, S., Hu, R., Suris, D., Ryali, C., Alwala, K. V., Khedr, H., Huang, A., et al., 2025. Sam 3: Segment anything with concepts. arXiv preprint arXiv:2511.16719. DOI: 10.48550/arXiv.2511.16719

Choudhary, T., Dewangan, V., Chandhok, S., Priyadarshan, S., Jain, A., Singh, A. K., Srivastava, S., Jatavallabhula, K. M., Krishna, K. M., 2024. Talk2bev: Language-enhanced bird’s-eye view maps for autonomous driving. In: 2024 IEEE International Conference on Robotics and Automation (ICRA). IEEE, pp. 16345–16352. DOI: 10.1109/ICRA57147.2024.10611485

Ettinger, S., Cheng, S., Caine, B., Liu, C., Zhao, H., Pradhan, S., Chai, Y., Sapp, B., Qi, C. R., Zhou, Y., et al., 2021. Large scale interactive motion forecasting for autonomous driving: The waymo open motion dataset. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 9710–9719. DOI: 10.1109/ICCV48922.2021.00957

Google DeepMind, 2026. Gemma 4 model card. https://ai.google.dev/gemma/docs/core/model_card_4, accessed: 2026-05-27.

Hu, Y., Yang, J., Chen, L., Li, K., Sima, C., Zhu, X., Chai, S., Du, S., Lin, T., Wang, W., et al., 2023. Planning-oriented autonomous driving. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 17853–17862. DOI: 10.1109/CVPR52729.2023.01712

Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., et al., 2023. Segment anything. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 4015–4026. DOI: 10.1109/ICCV51070.2023.00371

Mao, J., Qian, Y., Ye, J., Zhao, H., Wang, Y., 2023. Gpt-driver: Learning to drive with gpt. arXiv preprint arXiv:2310.01415. DOI: 10.48550/arXiv.2310.01415

Shao, H., Hu, Y., Wang, L., Song, G., Waslander, S. L., Liu, Y., Li, H., 2024. Lmdrive: Closed-loop end-to-end driving with large language models. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 15120–15130. DOI: 10.1109/CVPR52733.2024.01432

Sima, C., Renz, K., Chitta, K., Chen, L., Zhang, H., Xie, C., Beißwenger, J., Luo, P., Geiger, A., Li, H., 2024. Drivelm: Driving with graph visual question answering. In: European conference on computer vision. Springer, pp. 256–274. DOI: 10.1007/978-3-031-72943-0 15

Tian, X., Gu, J., Li, B., Liu, Y., Wang, Y., Zhao, Z., Zhan, K., Jia, P., Lang, X., Zhao, H., 2024. Drivevlm: The convergence of autonomous driving and large vision-language models. arXiv preprint arXiv:2402.12289. DOI: 10.48550/arXiv.2402.12289

Wagner, R., Tas, O. S., Villa, J., Hauser, F., Shen, Y., Steiner, M., Strutz, D., Fernandez, C., Kinzig, C., Guitierrez-Cabello, G. S., Königshof, H., Immel, F., Schwarzkopf, R., Rack, N. A., Rösch, K., Wang, K., Pauls, J.-H., Lauer, M., Gilitschenski, I., Caesar, H., Stiller, C., 2026. Longtail driving scenarios with reasoning traces: The kitscenes longtail dataset. URL: https://arxiv.org/abs/2603.23607 DOI: 10.48550/arXiv.2603.23607

Wang, Y., Luo, W., Bai, J., Cao, Y., Che, T., Chen, K., Chen, Y., Diamond, J., Ding, Y., Ding, W., et al., 2025. Alpamayo-r1: Bridging reasoning and action prediction for generalizable autonomous driving in the long tail. arXiv preprint arXiv:2511.00088. DOI: 10.48550/arXiv.2511.00088

Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al., 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35, 24824–24837. DOI: 10.52202/068431-1800

Wilson, B., Qi, W., Agarwal, T., Lambert, J., Singh, J., Khandelwal, S., Pan, B., Kumar, R., Hartnett, A., Pontes, J. K., et al., 2023. Argoverse 2: Next generation datasets for self-driving perception and forecasting. arXiv preprint arXiv:2301.00493. DOI: 10.48550/arXiv.2301.00493

Xu, Z., Zhang, Y., Xie, E., Zhao, Z., Guo, Y., Wong, K.-Y. K., Li, Z., Zhao, H., 2024. Drivegpt4: Interpretable end-to-end autonomous driving via large language model. IEEE Robotics and Automation Letters 9 (10), 8186–8193. DOI: 10.1109/LRA.2024.3440097

Downloads

Published

2026-09-01

Issue

Section

Visión por Computador