Pre-ECCV mini-workshop

Workshop on Geometry, Learning and 3D Understanding 2026

An afternoon of talks and discussion on learning-based 3D vision, spanning image matching, reconstruction, geometric reasoning, and structured scene representations. Attendance is open to everyone.

Schedule

Preliminary schedule (subject to change)

  1. Welcome coffee 15 minutes
  2. Na Zhao Talk and questions
  3. David Nordström Talk and questions
  4. Break 15 minutes
  5. Fadi Khatib Talk and questions
  6. Gustav Hanning Talk and questions

Speakers

Na Zhao

Na Zhao

Singapore University of Technology and Design

Generalizable, Semantic, and Editable 3D Scene Reconstruction

Recent advances have made high-quality 3D reconstruction increasingly practical. However, photorealistic rendering alone is not enough for many downstream applications: reconstructed scenes should also generalize to unseen environments, capture meaningful semantic structure, and support intuitive editing and manipulation. In this talk, I will present our recent work toward this broader goal of generalizable, semantic, and editable 3D scene reconstruction. I will first discuss generalizable reconstruction from sparse views, focusing on how explicit geometric constraints and learned multi-view correspondence can be combined for robust and efficient reconstruction. I will then show how geometry can serve not only as an output, but also as a foundation for cross-view semantic reasoning, progressing from geometry–semantics synergy to open-vocabulary 3D semantic fields from unposed images. Finally, I will introduce compositional scene representations that combine interpretable geometric primitives with high-fidelity 3D Gaussians, enabling part-aware and physically meaningful editing. Together, these works explore how 3D reconstruction can evolve toward building structured, understandable, and actionable representations of the 3D world.

David Nordström

David Nordström

Chalmers University of Technology

Learning to Understand 3D

While the NLP community has converged on a single objective for language understanding: next token prediction, no such convegence has been achieved in vision, even less so in 3D vision. To understand 3D, there are many competing approaches at the moment: (i) build a classic optimization pipeline on-top of possibly learned priors (e.g. colmap), (ii) directly infer 3D by training on loads of labeled data (e.g. VGGT) and (iii) train on single images in SSL and gain 3D understanding as a by-product. In this talk I will talk about some ongoing/past work in each of the categories above. In particular, I will focus on the third category: SSL for 3D, as I believe SSL is our only promise for a scalable and general approach. I will start by talking about SfM pipelines where our matchers are used, then talk about feedforward reconstruction, which I show also understands image matching, and finally our work on SSL for 3D.

Fadi Khatib

Fadi Khatib

Weizmann Institute of Science

Learning and Geometry for Structure-from-Motion

In this talk, I will present some of our recent work on learning-based methods for 3D reconstruction, with a focus on Structure-from-Motion and camera pose estimation. I will discuss how we use deep learning to address different estimation problems that arise throughout the reconstruction pipeline, from robust relative geometry to global camera recovery. A central question in this work is how to benefit from learned models without discarding the geometric structure that makes classical methods accurate and reliable. I will present several examples of this approach, including methods for robust estimation and large-scale pose averaging, and discuss what we have learned about the strengths and limitations of combining learning with multi-view geometry.

Gustav Hanning

Gustav Hanning

Lund University

PolyLayout: Multi-room Manhattan Layout Estimation

Gustav will present PolyLayout, his ECCV 2026 paper on multi-room Manhattan layout estimation. The talk will introduce the problem, outline the proposed approach, and discuss how it reconstructs consistent indoor geometry across multiple rooms.