A new research model called DepthEvidence teaches an AI to measure distances and reason about them in plain language, in a single system.
Researchers built DepthEvidence, a 4-billion-parameter multimodal model that estimates metric depth - the actual distance to objects, not just relative closeness - directly from images. A camera-aware decoder generates full-resolution depth maps, then converts those measurements into tokens the language model can read alongside object labels. The team also built a new benchmark, Depth-VQA, to test whether models can answer depth questions, compare distances, and combine spatial and numerical constraints in one decision. Tested across nine datasets, DepthEvidence posted the best average accuracy on dense depth estimation among the models compared, while holding its own against specialized depth estimators built for that task alone.
Most vision-language models can describe a scene but stumble when asked exactly how far something is or whether one object sits farther away than another - the kind of grounded reasoning robotics and AR applications need. By folding depth estimation and language reasoning into one model instead of bolting a separate depth sensor's output onto a chatbot, DepthEvidence suggests the two tasks can share a single reasoning pipeline without one degrading the other. General visual question-answering performance held steady too, which is the detail that usually breaks in these combined systems.
It is a research paper, not a product, and 'competitive with specialized estimators' is doing some diplomatic work - the dedicated depth models it was measured against still exist for a reason.