A new agent harness called Praxa promises to prove, step by step, that an AI agent's actions actually happened - not just that the agent claimed they did.
Praxa builds in five explicit checkpoints: proposal, authority, dispatch, verified external effect, and serving promotion. Researchers ran four separate evaluations to test it. A self-run audit of the codebase passed 1,027 of 1,027 unit tests and 89 of 89 Workerd tests, though the raw transcripts were not independently reproduced. In a Terminal-Bench pilot across 12 tasks, the harness-equipped version tied the baseline at 17 of 36 strict trials while using 37.49% more input tokens and 50.73% more output tokens. A separate development comparison found a Praxa-derived agent matched baseline accuracy on 180 of 180 trials while using 37.11% fewer tokens, 33.84% lower estimated cost, and 11.63% fewer steps.
The interesting part isn't speed, it's the accounting. Most agent frameworks let a model simply claim it finished a task. Praxa forces an external read-back to confirm the effect actually occurred before anything gets marked done, which matters more as companies hand agents real permissions like editing production config or moving money.
The authors are careful to say this proves the bookkeeping works, not that the agent is smarter, safer, or cheaper at scale - a distinction plenty of agent pitch decks skip.