AI/ ai agents · agent training · telemetry · benchmarks

New method lets AI agents learn from unlabeled usage logs

TeleTune turns messy, goal-free telemetry into reusable agent skills, beating prior memory methods on two benchmarks with far less validation compute.

A new framework called TeleTune teaches computer-use agents by mining the skills hidden in old, messy usage logs, with no replay required.

Researchers built TeleTune to pull a reusable skill library out of offline telemetry: records of people clicking through software with no stated goal, no way to replay the session, and often several tasks tangled together in one log. Instead of testing a candidate skill update by running it live, TeleTune scores it by how well it predicts the next action in logged trajectories, keeping only edits that improve that held-out accuracy, a proxy the team calls skill-guided progress. The learned skills also index demonstrations by subgoal, so a new task can pull in relevant snippets instead of a full matching trajectory. On the WorkArena and Online-Mind2Web benchmarks, TeleTune beat random retrieval and Agent Workflow Memory, hitting average success rates of 77.1% and 80.6% (6.7 and 7.7 points above the strongest baseline), and it still topped 68.5% success even when the training logs were heavily corrupted.

Agent builders have mostly trained on hand-crafted demonstrations or costly live rollouts, because raw usage logs are too chaotic to learn from directly. TeleTune's offline-only scoring costs 5 to 75 times fewer tokens than validating the same edits with live episodes, which matters for any company hoping to train agents on the exhaust of its own software logs without rerunning every session. That cost gap is the real pitch here, not the accuracy bump.

It is still benchmark-only testing picked by the paper's own authors, so the real test is whether skill-guided progress survives contact with telemetry nobody curated for a publication.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →