Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models

Haohan Chi*,1 · Huan-ang Gao*,1 · Ziming Liu†,2 · Jianing Liu1 · Chenyu Liu1 · Jinwei Li1 · Kaisen Yang1 · Yangcheng Yu1 · Zeda Wang1 · Wenyi Li1 · Leichen Wang2 · Xingtao Hu2 · Hao Sun2 · Hang Zhao3 · Hao Zhao†,1

1AIR, Tsinghua University · 2Bosch Research · 3IIIS, Tsinghua University

*Equal contribution · Corresponding author

NeurIPS 2025

Visual abstract of the Impromptu VLA dataset and its benchmark results
Visual abstract. The Impromptu VLA Dataset contains over 80K clips curated from 8 open-source driving datasets, organised around four challenging unstructured categories.

Abstract

Vision-Language-Action (VLA) models for autonomous driving show promise but falter in unstructured corner case scenarios, largely due to a scarcity of targeted benchmarks. To address this, we introduce Impromptu VLA. Our core contribution is the Impromptu VLA Dataset: over 80,000 meticulously curated video clips, distilled from over 2M source clips sourced from 8 open-source large-scale datasets. This dataset is built upon our novel taxonomy of four challenging unstructured categories and features rich, planning-oriented question-answering annotations and action trajectories.

Crucially, experiments demonstrate that VLAs trained with our dataset achieve substantial performance gains on established benchmarks — improving closed-loop NeuroNCAP scores and collision rates, and reaching near state-of-the-art L2 accuracy in open-loop nuScenes trajectory prediction. Furthermore, our Q&A suite serves as an effective diagnostic, revealing clear VLM improvements in perception, prediction, and planning.

Data Processing and Annotation Pipeline

Six-stage data processing and annotation pipeline for the Impromptu VLA dataset
From dataset collection and taxonomy-driven scenario mining, through frequency alignment and clip selection, to multi-task annotation generation with Qwen2.5-VL and human verification.

Open-loop Results on nuScenes

L2 trajectory prediction error (metres, lower is better). Training on a subset of Impromptu VLA before fine-tuning on nuScenes improves every horizon.

Open-loop trajectory prediction L2 errors on nuScenes
Model1 s2 s3 sAvg.
3B Base + nuScenes0.140.300.580.34
3B Base + Impromptu + nuScenes0.130.270.520.30
7B Base + nuScenes0.130.280.550.32
7B Base + Impromptu + nuScenes0.130.270.530.30

BibTeX

BibTeX
@inproceedings{chi2025impromptu,
  title     = {Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models},
  author    = {Chi, Haohan and Gao, Huan-ang and Liu, Ziming and Liu, Jianing and Liu, Chenyu and Li, Jinwei and Yang, Kaisen and Yu, Yangcheng and Wang, Zeda and Li, Wenyi and Wang, Leichen and Hu, Xingtao and Sun, Hao and Zhao, Hang and Zhao, Hao},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
  year      = {2025}
}

Acknowledgements

Outcome of an industry–academia collaboration between Bosch Research and Tsinghua University, sponsored by Bosch Research. Figures on this page are reproduced from the paper, which is released under CC BY-NC-ND.