⚡MapLightning: Online Vectorized HD Map Construction with 1D Map Tokens

Carnegie Mellon University

TL;DR: MapLightning replaces dense BEV with compact 1D map tokens for online vectorized HD map construction: 16.7× fewer tokens, 53% less memory, 40+ FPS inference, and 10%+ higher mAP.

Aggregating multi-view image tokens into a scene representation

Aggregating multi-view image tokens into a scene representation

(a) BEV-based methods project image features onto a dense grid of fixed BEV cells, few of which contain map elements. (b) Vanilla cross-attention compresses image tokens into compact map tokens, with no image-to-image or map-to-map interaction. (c) Our self-attention attends to and refines all tokens jointly.

Overview of MapLightning

Overview of MapLightning

A shared image backbone encodes each of the C camera views over T timesteps into image tokens, which are flattened and augmented with learned positional, camera, and time embeddings. We concatenate them with K learnable map tokens and jointly refine both with a self-attention mapper. The image tokens are then discarded, and the compact map tokens serve as the scene representation, which task-specific decoders attend to with full cross-attention.

Video Results

Speed:

Example reconstruction results

Comparisons with baselines