[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-ai-agents-still-fumble-real-business-decisions-study-finds":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},8526,"ai-agents-still-fumble-real-business-decisions-study-finds","AI Agents Still Fumble Real Business Decisions, Study Finds","A new arXiv preprint benchmark finds today's AI agents fail to reliably handle multi-step enterprise tasks like consulting cases and supply-chain planning.","A new benchmark says AI agents still can't run a business.\n\nResearchers behind a new arXiv preprint, posted September 30 and not yet peer-reviewed, introduce EnterpriseBench, a testing suite for large language model agents on enterprise strategic reasoning. It folds existing enterprise and financial QA datasets into one foundational suite, then adds three interactive scenarios: a management-consulting simulation where the agent diagnoses a client's problem through multi-turn questioning, a version of the classic Beer Game supply-chain simulation that tests inventory decisions under delayed feedback, and an \"Enterprise Digital Twin\" that simulates workforce, risk, and project planning. The authors ran nine agent methods across four backbone models, and none handled the full range of tasks reliably.\n\nMost enterprise AI benchmarks measure whether a model can pull a number out of a filing or answer a finance trivia question - useful, but not the same as making a call on incomplete information and living with the consequences over several turns. EnterpriseBench targets that gap directly, and the fact that no method-model combination held up consistently is a useful reality check for anyone pitching autonomous \"AI employees\" this year.\n\nIt's a preprint, so treat the numbers as a first read rather than a verdict - but the test design itself is a fair admission that most agent benchmarks have been measuring the wrong thing.","[\"ai-agents\",\"llm-benchmarks\",\"enterprise-ai\",\"arxiv\"]","2026-09-30T04:00:00.000Z","2026-09-30T09:00:43.654Z","2026-09-30T09:00:50.001Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Attribute the findings to their actual source — name it as a new arXiv preprint (not yet peer-reviewed) rather than just 'a new benchmark says,' so readers can verify and weigh the claims appropriately.","resolved","ai",[32,33,34,35],"ai-agents","llm-benchmarks","enterprise-ai","arxiv",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.37658",0,{"sections":42},[43,46,50,54,59,64,69,74,79,84,89,94,99,104],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",5029,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",780,{"name":51,"slug":52,"count":53,"latest_published_at":18},"Policy","policy",417,{"name":55,"slug":56,"count":57,"latest_published_at":58},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":68},"Science","science",154,"2026-09-28T13:19:18.000Z",{"name":70,"slug":71,"count":72,"latest_published_at":73},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":78},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":80,"slug":81,"count":82,"latest_published_at":83},"Dev Tools","dev-tools",89,"2026-09-29T17:15:00.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":105,"slug":106,"count":107,"latest_published_at":108},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]