Turns out the sharper the photo, the less arguing humans do about it.
A new arXiv study tracked how 187 annotators and reviewers labeled building damage across three imagery types from nine disasters: 20,041 buildings in drone photos, 20,695 in crewed-aviation shots, and 33,392 in satellite images, for a combined 74,128 labeled buildings. Every label went through a single-reviewer pass and then a consensus committee. The committee revised annotations far more often for lower-resolution sources: 25.27% of crewed-aviation labels and 36.95% of satellite labels got changed, versus a much lower rate for drone imagery. Even after one round of individual review, the gap persisted: 6.85% of drone labels, 14.05% of crewed-aviation labels, and 20.86% of satellite labels still needed committee correction.
Most post-disaster damage datasets are built from a single imagery source, so nobody has had to reckon with how differently humans perform across resolutions in the same pipeline. If satellite images - the cheapest and fastest to obtain after a disaster - carry roughly three times the error rate of drone photos even after review, then agencies leaning on satellite-only crowdsourcing for early damage estimates may be shipping shakier numbers than they realize. That is a real problem when those estimates guide where responders send resources first.
The paper's own fix isn't more scrutiny of satellite imagery for its own sake - it's routing more reviewer time to the blurriest pictures, a fairly unglamorous recommendation that says a lot about where crowd-sourced AI datasets actually break.