Vision Transformers trained without any geometric supervision are increasingly used as backbones for geometric tasks, yet it is unclear how much 3D information their features actually retain. This project probes frozen DINOv2 features directly, asking what geometry self-supervised representations implicitly encode and whether it can be extracted without large learned models.
Experiments use the TUM RGB-D SLAM dataset, which provides synchronised RGB streams with ground-truth camera trajectories from a motion capture system, giving calibrated intrinsics and ground-truth relative poses for supervising and evaluating fundamental matrix estimation. Places365 is used for tasks that need no multi-view geometry. DINOv2 (ViT-B/14) is kept entirely frozen throughout.
Gabriel López-Asiaín
Toyota Technological Institute at Chicago