AI/ ai agents · synthetic training data · llm benchmarks · model efficiency

Fx-Work-35B Beats Bigger AI Models After Training on 20K Tasks

A 35-billion-parameter model trained on just 20,000 synthetic work tasks outperformed far larger AI systems on benchmarks measuring real job skills.

A 35-billion-parameter model just beat a 1.6-trillion-parameter rival on tests of real office work.

Researchers built WorkGenesis, a framework that generates realistic job tasks by pulling real-world documents tied to O*NET occupational categories, then constructing a work request and grading rubric around each one. A verification step renders a sample deliverable for every task and checks whether unmet rubric items trace back to the agent, the task design, or the rubric itself, feeding failures back until the task passes muster. The team used this pipeline to synthesize 20,000 units of training work, then fine-tuned a 35-billion-parameter model, Fx-Work-35B, on that dataset using standard supervised fine-tuning. Across three benchmarks for workplace agents - GDPvalAA-v2, APEX-Agents-AA, and JobBench - Fx-Work-35B averaged a score of 31.00, ahead of comparable-sized models' 24.79 average, and ahead of DeepSeek-V4-Pro-Preview, a model roughly 45 times its size.

The result is a data-quality story, not a bigger-is-better one: a much smaller model trained on carefully verified, real-world-grounded tasks outpaced a far larger general-purpose system. That matters because expert-written training scenarios for office work are slow and expensive to produce by hand, and models trained on loosely generated synthetic tasks often learn to satisfy poorly specified rubrics rather than do the actual work.

The paper doesn't say whether Fx-Work-35B's weights or the WorkGenesis pipeline are available to outside researchers, so for now this is a result to watch, not a tool you can download and test yourself.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →