[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-fx-work-35b-beats-bigger-ai-models-after-training-on-20k-tasks":10,"sections":49},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":38,"tags":39,"sources":44,"feedback":48,"feedback_at":22,"cost_usd":48,"total_tokens":48},8934,"fx-work-35b-beats-bigger-ai-models-after-training-on-20k-tasks","Fx-Work-35B Beats Bigger AI Models After Training on 20K Tasks","A 35-billion-parameter model trained on just 20,000 synthetic work tasks outperformed far larger AI systems on benchmarks measuring real job skills.","A 35-billion-parameter model just beat a 1.6-trillion-parameter rival on tests of real office work.\n\nResearchers built WorkGenesis, a framework that generates realistic job tasks by pulling real-world documents tied to O*NET occupational categories, then constructing a work request and grading rubric around each one. A verification step renders a sample deliverable for every task and checks whether unmet rubric items trace back to the agent, the task design, or the rubric itself, feeding failures back until the task passes muster. The team used this pipeline to synthesize 20,000 units of training work, then fine-tuned a 35-billion-parameter model, Fx-Work-35B, on that dataset using standard supervised fine-tuning. Across three benchmarks for workplace agents - GDPvalAA-v2, APEX-Agents-AA, and JobBench - Fx-Work-35B averaged a score of 31.00, ahead of comparable-sized models' 24.79 average, and ahead of DeepSeek-V4-Pro-Preview, a model roughly 45 times its size.\n\nThe result is a data-quality story, not a bigger-is-better one: a much smaller model trained on carefully verified, real-world-grounded tasks outpaced a far larger general-purpose system. That matters because expert-written training scenarios for office work are slow and expensive to produce by hand, and models trained on loosely generated synthetic tasks often learn to satisfy poorly specified rubrics rather than do the actual work.\n\nThe paper doesn't say whether Fx-Work-35B's weights or the WorkGenesis pipeline are available to outside researchers, so for now this is a result to watch, not a tool you can download and test yourself.","[\"ai agents\",\"synthetic training data\",\"llm benchmarks\",\"model efficiency\"]","2026-10-01T04:00:00.000Z","2026-10-01T11:31:44.422Z","2026-10-01T11:31:51.382Z","published",null,[24,30,34],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The headline calls the training data 'Synthetic Work Tasks,' but the body (and source) explicitly contrasts WorkGenesis with unconstrained synthesis, grounding tasks in real-world documents instead — fix the headline so it doesn't contradict the article's own description of the method.","resolved",{"id":31,"reviewer":26,"round":32,"reason":33,"status":29},"editor-r2",2,"Change the headline's plural 'Giant Rivals' to singular (e.g. 'Giant Rival') since the body names only one giant-scale model beaten (DeepSeek-V4-Pro-Preview) — the other comparison set is same-size baselines, not giants.",{"id":35,"reviewer":26,"round":36,"reason":37,"status":29},"editor-r3",3,"The headline and tag assert the model is 'open'\u002Fopen-source, but nothing in the body or source material confirms Fx-Work-35B's weights or code are publicly released — either add evidence for the open claim or drop 'Open' from the headline and the 'open-source ai' tag.","ai",[40,41,42,43],"ai agents","synthetic training data","llm benchmarks","model efficiency",[45],{"name":46,"url":47},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.39325",0,{"sections":50},[51,54,59,64,69,74,79,84,89,93,98,103,108,113],{"name":52,"slug":38,"count":53,"latest_published_at":18},"AI",5351,{"name":55,"slug":56,"count":57,"latest_published_at":58},"Security","security",801,"2026-09-30T22:18:23.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Policy","policy",429,"2026-10-01T02:26:17.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":68},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":70,"slug":71,"count":72,"latest_published_at":73},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":78},"Science","science",157,"2026-09-30T15:00:56.000Z",{"name":80,"slug":81,"count":82,"latest_published_at":83},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":90,"slug":91,"count":87,"latest_published_at":92},"Software","software","2026-09-30T21:41:11.000Z",{"name":94,"slug":95,"count":96,"latest_published_at":97},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":99,"slug":100,"count":101,"latest_published_at":102},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":104,"slug":105,"count":106,"latest_published_at":107},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":109,"slug":110,"count":111,"latest_published_at":112},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":114,"slug":115,"count":116,"latest_published_at":117},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]