[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-ai-models-attempt-to-score-olympic-diving-like-judges":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},6843,"ai-models-attempt-to-score-olympic-diving-like-judges","AI Models Attempt to Score Olympic Diving Like Judges","Standalone vision-language models barely track real judges, but a four-model ensemble more than doubles that correlation, hitting 0.67.","Researchers just asked AI models to judge Olympic diving, and mostly, they flopped.\n\nResearchers tested several open-source vision-language models in a zero-shot setup against the AQA-7 diving benchmark, asking them to reason about form and generate phase-level sub-scores. On their own, the models' scores correlated weakly with real judges' scores, with Spearman correlations below 0.32. The team then built a framework combining the models' text explanations, via TF-IDF vectorization and dimensionality reduction, with an ensemble regression layer. Stacking four models this way pushed the correlation up to 0.67, and notably, the models' written reasoning proved more useful than their raw numeric sub-scores.\n\nThis isn't AI replacing judges - it's proof that off-the-shelf vision-language models are mediocre judges alone but decent assistants when their reasoning is treated as a feature rather than a verdict. It also suggests that letting models explain themselves in text, then quantifying that text statistically, can beat trusting their final numeric scores directly - a lesson that likely applies well beyond diving boards, to any domain where AI renders a judgment call.\n\nA 0.67 correlation beats guessing, but it's still far short of the consistency a panel of human judges delivers, so don't expect algorithmic scorecards at the next Olympics anytime soon.","[\"ai\",\"computer vision\",\"sports tech\",\"research\"]","2026-09-18T04:00:00.000Z","2026-09-18T19:09:34.879Z","2026-09-18T19:09:46.792Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Fix the dek's claim that combining models 'nearly doubles' the correlation — the body correctly states the ensemble more than doubles it (0.67 vs below 0.32), so the dek and body contradict each other on the central figure.","resolved","ai",[30,32,33,34],"computer vision","sports tech","research",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.19354",0,{"sections":41},[42,45,49,54,59,63,67,72,76,81,86,91,96,101],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",4031,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",654,{"name":50,"slug":51,"count":52,"latest_published_at":53},"Policy","policy",338,"2026-09-11T04:00:00.000Z",{"name":55,"slug":56,"count":57,"latest_published_at":58},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":18},"Hardware","hardware",155,{"name":64,"slug":65,"count":66,"latest_published_at":18},"Science","science",121,{"name":68,"slug":69,"count":70,"latest_published_at":71},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":73,"slug":74,"count":75,"latest_published_at":18},"Dev Tools","dev-tools",78,{"name":77,"slug":78,"count":79,"latest_published_at":80},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":102,"slug":103,"count":104,"latest_published_at":105},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]