AI agents that learn new skills by trial and error have a new trick: looking inward.
Researchers have introduced Rep2Skill, a framework that lets large language model agents revise their own textual "skills" - reusable instructions for completing tasks - by analyzing their internal model representations, not just the text of what they did. Typically, skill evolution works by having an optimizer read through long execution logs and sparse success-or-failure signals to figure out what went wrong. Rep2Skill instead tracks how an agent's internal representations shift during a task, flags the moments where those internal dynamics diverge from patterns seen in successful runs, and turns that into targeted textual feedback for rewriting the skill. In tests across two agent environments with two open-source LLMs, Rep2Skill beat text-only skill-revision approaches, including in the harder self-evolution setup where the same model has to both execute tasks and critique its own performance, without help from a stronger outside model.
Most agent self-improvement right now treats the model as a black box, judging it only by what it writes down afterward. Rep2Skill's bet is that an agent's hidden representations carry diagnostic information - essentially a readout of where things started going sideways - that text transcripts alone miss. That matters for anyone trying to make agents better at a task without running expensive fine-tuning jobs.
Still, this was validated in just two environments with two open-source models, so whether reading an agent's internal state scales to messier, real-world deployments remains an open question.