MR.ScaleMaster: Scale-Consistent Collaborative Mapping from Crowd-Sourced Monocular Videos
arXiv:2604.11372v4 Announce Type: replace Abstract: Crowd-sourced cooperative mapping combines monocular sessions from different front-ends, each an independently reconstructed keyframe sequence. Each session has its own local coordinate frame and may follow an incompatible scale convention. We present MR.ScaleMaster, a backend that accepts image, Sim(3) pose, and point-map packets without requiring a common reconstruction model, camera intrinsics, or front-end-provided metric scale. Our Cross-
Overview
arXiv:2604.11372v4 Announce Type: replace Abstract: Crowd-sourced cooperative mapping combines monocular sessions from different front-ends, each an independently reconstructed keyframe sequence. Each session has its own local coordinate frame and may follow an incompatible scale convention. We present MR.ScaleMaster, a backend that accepts image, Sim(3) pose, and point-map packets without requiring a common reconstruction model, camera intrinsics, or front-end-provided metric scale. Our Cross-Front-End Loop Factor (CFL) uses a shared matcher for image correspondences but retrieves matched 3D points from the input point maps, so its scale estimates the inter-session ratio. Scale Preconditioning (SPC) initializes session scales for a Sim(3) anchor-node graph, which then corrects residual scale and drift. For front-ends exposing an incremental scale trajectory, an agent-side Scale Collapse Alarm (SCA) rejects or rolls back false intra-session loops that would otherwise collapse the session scale. We evaluate seven front-ends on KITTI and five on CODa. Based on ground-truth path lengths, session scales in our heterogeneous KITTI setting differ by 57-103x. CFL reduces mean ATE by 37% and inter-session scale error by 60% relative to loop factors from independent pairwise reconstructions. Adding SPC raises the reductions to 74% and 77%, respectively. On CODa, five-session fusion improves mean ATE over single-session runs for all five front-ends, while joint fusion registers 15 sessions from three different front-ends in a single map. Code will be released.
Source
Originally published at arxiv.org.
Related Articles
Source: https://arxiv.org/abs/2604.11372