A new benchmark called Agentick tested 27 different AI agent setups and found that no single approach beats the rest.
Researchers built Agentick with 37 procedurally generated tasks split into six capability categories, four difficulty levels, and five ways of observing the environment, from plain language to ASCII grids, all wired into one standard interface. They ran 27 agent configurations, including reinforcement learning, large language models, vision-language models, hybrids, and human baselines, through more than 90,000 episodes. GPT-5 mini came out on top overall with an oracle-normalized score of 0.309, meaning even the best agent solved less than a third of what a perfect reference policy would achieve. Older-style reinforcement learning (PPO) still beat every language model on planning and multi-agent tasks.
That matters because it undercuts the assumption that today's large language models have simply absorbed general-purpose reasoning ability. Two other findings stand out: adding a reasoning harness multiplied LLM scores by 3 to 10 times, and stripping tasks down to ASCII characters worked better than describing them in plain English, suggesting current agents lean on format as much as genuine reasoning.
It is a useful reality check for anyone pitching a chatbot as a drop-in replacement for a purpose-built control system: it isn't, at least not yet.