Good Papers

The Road Ahead in Autonomous Driving: The KITScenes Multimodal Dataset

KITScenes Multimodal provides European autonomous driving data with high-fidelity synchronized sensors, complete topologically connected 3D HD maps, and four embodied AI benchmarks.

Richard Schwarzkopf, Fabian Immel, Alexander Blumberg, Jonas Merkert, Nils A Rack, Kaiwen Wang, Fabian Konstantinidis, Julian Truetsch, Carlos Fernandez, Annika Bätz, Kevin Rösch, Marlon Steiner, Willi Poh, Yinzhe Shen, Felix Hauser, Dominik Strutz, Jaime Villa, Gleb Stepanov, Royden Wagner, Omer Sahin Tas, Frank Bieder, Holger Caesar, Jan-Hendrik Pauls, Christoph Stiller

Published 2026Sydney Poster Session 3 · Wed, Dec 9, 10:00 AM–1:00 PM local time · Hall 1-4▲ 19 on Hugging FaceCode ★ 26arXiv ↗OpenReview ↗

76%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel10/20reviewers recommend it
lenient 5/5
medium 3/10
strict 2/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Existing autonomous driving datasets have enabled major progress, but fall short in sensor fidelity, map completeness, or geographic diversity. We present KITScenes Multimodal, a European dataset built around high-fidelity sensors and maps. Our fully synchronized sensor suite combines high-resolution global-shutter cameras, long-range lidar beyond 400m, 4D imaging radar, and redundant GNSS/INS localization. Our HD maps are, to our knowledge, the most complete of any sensor dataset, validated through autonomous driving trials on open-source software. For the first time in a public dataset, all driving-relevant traffic elements, such as traffic lights, are mapped in 3D to a reprojection-accurate level with full topological connectivity. Recorded in cities with irregular street layouts and mixed traffic modes, our dataset complements existing datasets by broadening the available geographic diversity. We also introduce four benchmarks, each advancing spatial learning for embodied AI: online HD map construction, long-range depth estimation, novel view synthesis, and end-to-end driving. Project page: https://kitscenes.com/