AI/ ai-agents · benchmarks · reinforcement-learning · llm-evaluation

New Benchmark Shows No AI Agent Approach Wins Every Task

Agentick pits reinforcement-learning, language-model, and hybrid AI agents against the same 37 tasks, and none of them wins across the board.

A new benchmark called Agentick tested 27 different AI agent setups and found that no single approach beats the rest.

Researchers built Agentick with 37 procedurally generated tasks split into six capability categories, four difficulty levels, and five ways of observing the environment, from plain language to ASCII grids, all wired into one standard interface. They ran 27 agent configurations, including reinforcement learning, large language models, vision-language models, hybrids, and human baselines, through more than 90,000 episodes. GPT-5 mini came out on top overall with an oracle-normalized score of 0.309, meaning even the best agent solved less than a third of what a perfect reference policy would achieve. Older-style reinforcement learning (PPO) still beat every language model on planning and multi-agent tasks.

That matters because it undercuts the assumption that today's large language models have simply absorbed general-purpose reasoning ability. Two other findings stand out: adding a reasoning harness multiplied LLM scores by 3 to 10 times, and stripping tasks down to ASCII characters worked better than describing them in plain English, suggesting current agents lean on format as much as genuine reasoning.

It is a useful reality check for anyone pitching a chatbot as a drop-in replacement for a purpose-built control system: it isn't, at least not yet.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →