Vision-Based 3D Semantic Occupancy Prediction Achieves State-of-the-Art Results

A team of researchers at the University of Bonn has made a breakthrough in robotic perception, developing a novel method for semantic occupancy prediction using only vision data. This approach, called Sfmocc, does not require manual or LiDAR-based labels, achieving state-of-the-art results for 3D occupancy prediction among methods trained without labels. The research has been peer-reviewed and published in IEEE Robotics and Automation Letters.

Key Takeaways:

  • The researchers proposed a novel method for semantic occupancy prediction using only vision data, eliminating the need for manual or LiDAR-based labels.
  • The approach, called Sfmocc, leverages all available training images of a sequence and uses bundle adjustment to align images and estimate camera poses.
  • The method computes semantic maps from a pre-trained open-vocabulary image model and generates occupancy pseudo labels to optimize for 3D semantic occupancy prediction.
  • Sfmocc achieves state-of-the-art results for 3D occupancy prediction among methods trained without labels.
  • The research was conducted by a team of researchers at the University of Bonn, led by Rodrigo Marcuzzi.
  • The team consisted of additional authors Lucas Nunes, Elias Marks, Louis Wiesmann, Thomas Laebe, Jens Behley, and Cyrill Stachniss.
  • The research has been funded by the German Research Foundation (DFG) and the Federal Ministry of Education & Research (BMBF).

Statistics:

  • 3D occupancy voxel grids were predicted without any manual or LiDAR-based labels.
  • The research achieved state-of-the-art results for 3D occupancy prediction among methods trained without labels.
  • The approach was evaluated on a dataset of urban environments.
  • The research was published in IEEE Robotics and Automation Letters, Vol. 10, No. 5, pp. 5074-5081, and was peer-reviewed.

Sources:

  • [1] Sfmocc: Vision-based 3D Semantic Occupancy Prediction In Urban Environments. IEEE Robotics and Automation Letters, 2025;10(5):5074-5081.
  • [2] Investigators at University of Bonn Describe Findings in Robotics and Automation (Sfmocc: Vision-based 3D Semantic Occupancy Prediction In Urban Environments). Robotics & Machine Learning. May 19, 2025; p 216.
  • [3] German Research Foundation (DFG)
  • [4] Federal Ministry of Education & Research (BMBF)