From video to a global map

VidMap

Exploiting Temporal Structure for Video-Based Structure-from-Motion

Zador Pataki1 · Paul-Edouard Sarlin2 · Marc Pollefeys1,3

1ETH Zurich   2Google   3Microsoft Spatial AI Lab

Top-down reconstruction captured from the interactive viewer with VidMap, ground truth, and baseline trajectories
VidMap Ground truth DA3-Long VIPE LOGER LingBotMap

VidMap reconstructs challenging long videos by integrating video-aware matching, metric depth priors, and robust loop-closure handling into global optimization

Scroll to explore

Large-scale reconstruction

VidMap’s camera poses, combined with dense triangulation, recover dense geometry along a kilometer-long route through Zurich’s old town.

Reconstructing videos from the web

From fast motion to unfamiliar environments, VidMap recovers accurate camera trajectories and 3D structure across diverse videos.

Racing drone

Cycling tour

Forest walk

Video game

How it works

Dense image matching propagates sparse tracks across keyframes, building long, accurate tracks. Multi-flow propagation further reduces drift by selecting reliable matches from multiple preceding keyframes.

Tracker recall on ETH3D SLAM and LaMAR

Visually similar places can create false loop closures. Provenance-aware losses distinguish sequential tracks from loop-closure matches, preserving reliable temporal constraints while suppressing visual aliasing.

Reconstruction with and without provenance-aware losses
Bundle Adjustment + Filtering
VidMapGround truthOutlierInlier
Matched control with loop-closure-specific losses.

Monocular depth priors regularize global positioning, resolving geometric ambiguities and limiting scale drift when camera motion provides weak multi-view constraints.

Depth ablation reconstruction
Bundle Adjustment + Filtering
VidMapGround truth
Full VidMap reconstruction.

Accuracy across benchmarks

State-of-the-art reconstruction of long, complex sequences across diverse visual domains.

Compact paired calibrated and uncalibrated benchmark plots

Higher is better throughout. The top row varies the allowed position error; the bottom row varies the length of trajectory being evaluated. Red denotes VidMap in these plots.

Explore VidMap

Explore the method, reproduce the experiments, or use the code on your own videos.