[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-benchmark-shows-llms-struggle-to-build-ramsey-graphs":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},5341,"new-benchmark-shows-llms-struggle-to-build-ramsey-graphs","New Benchmark Shows LLMs Struggle to Build Ramsey Graphs","A new dataset built to test genuine graph-construction reasoning rather than memorized answers finds LLMs solve fewer than 4 in 10 hard problems.","A new benchmark says today's large language models still can't reliably build even small mathematical graphs.\n\nResearchers introduced RamseyGadgets, a dataset of 70 graph construction problems drawn from Ramsey theory, built so solutions are small enough (at most 10 vertices) to be checked automatically with SAT solvers. The problems are deliberately obscure spins on Ramsey-good graphs - constructions that avoid specific monochromatic subgraphs - chosen precisely because they are underexplored, so models can't just recall a textbook answer. The team tested five open-source LLMs and found none topped 37.70% accuracy on the hardest tier of problems. The dataset is designed to expand easily, since swapping which subgraphs are banned generates a fresh set of problems.\n\nMath benchmarks for LLMs are increasingly compromised by models that have memorized known constructions from training data rather than reasoning them out. By digging into underexplored corners of Ramsey theory, RamseyGadgets tries to isolate genuine problem-solving from recall - a distinction that matters as labs tout reasoning gains on benchmarks like GSM8K or MATH that are now thoroughly picked over. A sub-40% success rate on a purpose-built benchmark suggests graph construction reasoning remains a real weak spot, even for models that do well on competition math.\n\nThe researchers also tested giving models hints, which is a quiet admission that without help, these systems are still closer to guessing than proving.","[\"ai-research\",\"llm-benchmarks\",\"graph-theory\",\"open-source-ai\"]","2026-08-18T04:00:00.000Z","2026-08-18T15:57:43.382Z","2026-08-18T15:57:55.250Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Verify the model identifier 'Gemma-4-31B' against real-world Gemma releases before publishing—no Gemma 4 generation or 31B size variant exists, so this proper noun cannot be confirmed and must be corrected or removed.","resolved","ai",[32,33,34,35],"ai-research","llm-benchmarks","graph-theory","open-source-ai",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.14999",0,{"sections":42},[43,47,51,56,61,66,71,76,81,85,90,95,100,105],{"name":44,"slug":30,"count":45,"latest_published_at":46},"AI",3293,"2026-08-20T04:00:00.000Z",{"name":48,"slug":49,"count":50,"latest_published_at":46},"Security","security",435,{"name":52,"slug":53,"count":54,"latest_published_at":55},"Policy","policy",210,"2026-08-19T09:32:27.000Z",{"name":57,"slug":58,"count":59,"latest_published_at":60},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":62,"slug":63,"count":64,"latest_published_at":65},"Hardware","hardware",140,"2026-08-19T18:25:42.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":70},"Consumer Tech","consumer-tech",95,"2026-08-18T16:05:00.000Z",{"name":72,"slug":73,"count":74,"latest_published_at":75},"Science","science",90,"2026-08-19T18:41:02.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Software","software",73,"2026-08-18T07:51:50.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":18},"Dev Tools","dev-tools",69,{"name":86,"slug":87,"count":88,"latest_published_at":89},"Startups","startups",47,"2026-08-19T19:13:46.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Gaming","gaming",41,"2026-07-09T04:00:00.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"General","general",33,"2026-08-18T22:18:13.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":106,"slug":107,"count":108,"latest_published_at":109},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]