Researchers built a benchmark to find out whether AI agents can actually plan a space mission, not just talk about one.
The new AstroAgentBench tests five large language model agent systems across seven task families covering scheduling, observation planning, constellation design, and relay support. Each agent submits a planning artifact - a schedule, a design, an observation sequence - that gets checked by an external verifier for things like timing, geometry, resource limits, and mission value. Researchers ran 35 held-out test cases and compared agent output against scores from task-specific solver references, the purpose-built optimization tools engineers already use for this work. The strongest agent systems matched or beat those solver references on several task families, but weaker systems often failed to produce valid, high-value plans, and even the best agents lost accuracy on tasks heavy on geometry or design.
This is a useful reality check for the idea that general-purpose AI agents can replace specialized mission-planning software. The benchmark's failure analysis is the real finding here: agents mostly stumble either by misreading what a task actually requires, or by building weak solutions once they understand it - and the ones that succeed do so by checking their own work against verifier feedback rather than guessing once and submitting.
In other words, the agents that do best are the ones that know how to grade their own homework before turning it in - which says more about good process than about any particular model being spaceflight-ready.