Security/ ai · security · bug bounty · prompt injection

OpenAI Launches Bug Bounty for AI Safety Risks

The program targets AI-specific attack vectors like prompt injection and agentic flaws, formalizing a threat model that's still being written.

OpenAI has launched a bug bounty program specifically for AI safety and abuse risks.

The new track is separate from OpenAI's existing security bounty, which covers conventional software flaws. This one targets a different category: prompt injection, data exfiltration through model interfaces, and vulnerabilities in agentic AI systems that can take autonomous actions on a user's behalf. Researchers who find and responsibly disclose qualifying flaws will receive payouts, though OpenAI did not specify amounts in the announcement.

Bug bounty programs for conventional software have existed for decades, but the AI-specific variant is newer and the threat model is still being mapped. Classifying prompt injection and agentic exploits as bounty-eligible officially acknowledges what security researchers have argued for a while: AI systems carry a distinct attack surface, and it grows as more autonomous agents get deployed on top of foundation models. A structured program should accelerate how quickly those weak points get found and documented.

The cynic's read: it also gives OpenAI a cleaner process for absorbing awkward disclosures before researchers post embarrassing demos publicly.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →