AI/ ai · ai-agents · benchmarks · training-data

A New Dataset Teaches AI Agents to Do Real Office Work

WorkForge generates thousands of realistic office environments to train AI agents, producing sharp benchmark gains for models fine-tuned on it.

Researchers built nearly 17,000 fake office environments to teach AI models how to actually finish work, not just talk about it.

The new framework, called WorkForge, pulls real files, spreadsheets, and documents and assembles them into workspaces spanning 40 professional domains and 60 file types. For each workspace, it extracts verifiable facts about the content, then uses those facts to write task instructions, solution plans, and automated checkers that confirm whether an agent's work is actually correct. The team built 16,700 of these environments and used them to post-train Qwen3.5-35B-A3B-Base, a large mixture-of-experts language model, pushing its GDPVal score from 45.5 to 73.6 and its APEX Score from 5.0 to 21.3. The same training data also helped a separate, smaller model, Qwen3.5-27B, become highly competitive with - and in some cases beat - stronger rivals.

Work agents have been stuck in a familiar bind: hand-built training environments are too expensive to scale, and synthetic ones are usually too simplistic to teach anything useful about messy professional tasks. WorkForge's bet is that grounding every task in real files and checkable facts can resolve that trade-off without faking realism, and the scaling behavior the researchers report across both data volume and task length suggests the approach isn't just a one-off trick.

The benchmark jumps are striking, but GDPVal and APEX Score are still research yardsticks, not a guarantee these agents can handle an actual client deliverable without supervision.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →