[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-researchers-build-a-chess-benchmark-to-grade-ai-prompt-tweaks":10,"sections":49},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":39,"tags":40,"sources":44,"feedback":48,"feedback_at":22,"cost_usd":48,"total_tokens":48},9463,"researchers-build-a-chess-benchmark-to-grade-ai-prompt-tweaks","Researchers Build a Chess Benchmark to Grade AI Prompt Tweaks","A new chess benchmark shows prompt optimization can boost frozen LLMs, but even the best performer solves only about 55 percent of its puzzles.","A new benchmark uses chess puzzles to test whether tweaking a prompt, not retraining a model, can make AI systems noticeably smarter.\n\nResearchers built the benchmark from 1,118 Lichess puzzles to study automatic prompt optimization, the practice of rewriting instructions fed to a frozen model rather than retraining it. They ran six optimization algorithms against eight target models, grading results with exact-match scoring and chess-engine evaluation of alternative moves. The paper names the strongest model tested as \"Gemini 3.5 Flash,\" used as the meta-model that drafts and critiques candidate prompts - that name doesn't match any model Google has publicly released, so we're flagging it as stated in the unreviewed paper rather than a confirmed, verifiable fact. Even that top performer solved only about 55 percent of the puzzles, and the researchers say the entire study cost around $800 to run.\n\nMost LLM benchmarks go stale fast: models memorize test sets or scores saturate within months. This one is designed to regenerate itself with fresh puzzles and adjustable difficulty, which matters more than any single score, since it gives researchers a cheap, hard-to-game way to compare prompt-optimization methods as models keep improving.\n\nChess has been a convenient stand-in for machine reasoning since Deep Blue; here it's demoted to an $800 scorecard for prompt engineering, not a measure of raw intelligence.","[\"ai\",\"llm-benchmarks\",\"prompt-engineering\",\"chess\"]","2026-10-02T04:00:00.000Z","2026-10-02T20:31:32.559Z","2026-10-02T20:31:37.250Z","published",null,[24,30,35],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The dek claims 'most models still can't crack half the puzzles,' but the source and body only substantiate a sub-50% implication for context around the top-scoring model (55%) — per-model breakdowns for the other seven models are never given, so rewrite the dek\u002Fheadline to only claim what's actually substantiated (the strongest model tested barely topping half).","resolved",{"id":31,"reviewer":32,"round":33,"reason":34,"status":29},"publisher-r2","publisher",2,"The article cites 'Gemini 3.5 Flash' as the meta-model, but no such model exists in Google's Gemini lineup (1.0\u002F1.5\u002F2.0\u002F2.5), making this a likely factual error.",{"id":36,"reviewer":26,"round":37,"reason":38,"status":29},"editor-r3",3,"The draft still states 'Gemini 3.5 Flash' as a real, strongest model without flagging that no such model exists in Google's public Gemini lineup — rewrite to note this name is as given in the (unreviewed) paper and could not be independently verified, rather than presenting it as a confirmed fact.","ai",[39,41,42,43],"llm-benchmarks","prompt-engineering","chess",[45],{"name":46,"url":47},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.00416",0,{"sections":50},[51,54,58,63,68,73,78,83,88,93,98,103,108,113],{"name":52,"slug":39,"count":53,"latest_published_at":18},"AI",5765,{"name":55,"slug":56,"count":57,"latest_published_at":18},"Security","security",831,{"name":59,"slug":60,"count":61,"latest_published_at":62},"Policy","policy",437,"2026-10-01T18:10:00.000Z",{"name":64,"slug":65,"count":66,"latest_published_at":67},"Deals","deals",317,"2026-10-01T22:00:00.000Z",{"name":69,"slug":70,"count":71,"latest_published_at":72},"Hardware","hardware",198,"2026-10-01T17:38:48.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Science","science",168,"2026-10-01T18:35:55.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":82},"Consumer Tech","consumer-tech",155,"2026-10-01T19:54:10.000Z",{"name":84,"slug":85,"count":86,"latest_published_at":87},"Dev Tools","dev-tools",96,"2026-10-01T16:57:03.000Z",{"name":89,"slug":90,"count":91,"latest_published_at":92},"Software","software",93,"2026-09-30T21:41:11.000Z",{"name":94,"slug":95,"count":96,"latest_published_at":97},"Startups","startups",90,"2026-10-01T21:55:22.000Z",{"name":99,"slug":100,"count":101,"latest_published_at":102},"Gaming","gaming",53,"2026-10-02T02:50:39.000Z",{"name":104,"slug":105,"count":106,"latest_published_at":107},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":109,"slug":110,"count":111,"latest_published_at":112},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":114,"slug":115,"count":116,"latest_published_at":117},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]