Researchers have found a way to make AI vision systems work with full 360-degree camera images, not just narrow photos.
The new framework, called OmniAct3D, takes AI detection models that were trained on ordinary flat photos and adapts them to equirectangular projection images, the stretched-out format used to represent a full spherical scene in one flat picture. That format warps geometry in ways that confuse models built for normal photos. OmniAct3D fixes this with three components: one that models the curved viewing rays of a spherical camera, one that grounds each detection guess in the wider panoramic scene before converting it into a spatial action, and one that re-examines object regions at higher resolution to nail down which way objects are facing. In tests on two benchmarks, Spheriverse and PanoMMOcc, it beat the previous best 3D detector by 2.96 points on one metric and beat an unadapted baseline model by nearly 25 points on another.
This matters because most 3D object detection research still assumes a robot or car is looking at one direction at a time, stitching together separate camera feeds if it needs a fuller view. Panoramic cameras collapse that into a single image, which is cheaper and simpler, but only if the AI underneath can actually parse the distorted geometry. If this approach holds up outside the two benchmarks tested here, it suggests labs won't need to retrain detection models from scratch for every new camera rig, which is a real cost saver for warehouse robots and delivery drones.
The code is promised on GitHub, but nobody outside the paper's benchmarks has kicked the tires yet.