A researcher at Anthropic walked away from his job this week, warning that self-improving AI systems could put humanity at risk.
Jacob Coxon left his role at Anthropic this week, saying he no longer wanted to work inside a system racing toward self-improving AI without brakes. His core worry: AI systems that can improve themselves could outpace human oversight before anyone notices the guardrails have failed. Rather than walking away quietly, Coxon used his exit to call for pacing agreements: commitments between AI labs to slow development in step with each other, rather than one lab restraining itself while a rival keeps sprinting.
This isn't an outsider lobbing criticism from a think tank. It's someone who was inside Anthropic, a company that markets itself as the safety-conscious alternative to other frontier AI developers. Pacing agreements are a coordination problem economists would recognize: no single lab wants to slow down first if it just hands rivals a lead. That Coxon raised this as a condition for leaving, rather than a footnote, suggests the tension between commercial speed and safety promises is getting harder to paper over from the inside.
Labs have talked about coordination before; getting competitors to actually slow down together, when speed is the entire business model, is another matter.