OpenAI launches Safety Bug Bounty to target AI-specific abuse risks
OpenAI has launched a public Safety Bug Bounty program to help researchers find AI-specific abuse and safety issues across its products. The new program runs alongside the company’s existing Security Bug Bounty, but it focuses on risks that may not fit the classic definition of a security vulnerability while still creating real-world harm.
The move matters because AI systems now face a broader range of threats than traditional software. In OpenAI’s view, some of the most serious problems involve misuse, prompt injection, data leakage, and account abuse rather than standard bugs like remote code execution or SQL injection. The new bounty aims to give those issues a formal reporting path.
Access content across the globe at the highest speed rate.
70% of our readers choose Private Internet Access
70% of our readers choose ExpressVPN
Browse the web from multiple devices with industry-standard security protocols.
Faster dedicated servers for specific actions (currently at summer discounts)
OpenAI says the Safety Bug Bounty is designed to complement, not replace, its Security Bug Bounty. Reports will be reviewed jointly by the safety and security bounty teams, and submissions may move between programs depending on the type of issue involved.
What the new OpenAI program covers
OpenAI says the program focuses on AI-specific safety scenarios. One major category is agentic risk, including prompt injection and data exfiltration cases where attacker-controlled text can reliably hijack a victim’s agent and push it to perform harmful actions or leak sensitive data. OpenAI says this includes products such as Browser, ChatGPT Agent, and similar agentic systems, and the behavior must be reproducible at least 50% of the time.
Another category covers the exposure of OpenAI proprietary information. According to the company, researchers can report generations that reveal reasoning-related proprietary information or other confidential internal data.
A third category focuses on account and platform integrity. OpenAI says this includes attempts to bypass anti-automation controls, manipulate trust signals, or evade account restrictions, suspensions, or bans.
Why this is different from a normal bug bounty
Traditional bug bounty programs usually reward vulnerabilities tied to unauthorized access, broken authentication, exposed data, or other classic security flaws. OpenAI’s new program extends that idea into the safety layer, where the risk may come from how an AI system behaves under abuse rather than from a simple coding mistake.
That distinction matters because many AI-related attacks do not fit cleanly into old security models. A model that follows a malicious prompt, leaks sensitive information through agent behavior, or performs unsafe actions at scale can create serious harm even when there is no textbook exploit chain behind it. OpenAI says the new bounty creates a structured way to report those cases.
What is out of scope
OpenAI has also drawn a firm line around what the program will not accept. Generic jailbreaks that only produce rude language or surface public information are out of scope. General content policy bypasses without a clear abuse or safety impact also do not qualify under this program.
For issues that enable unauthorized access to features, data, or functionality beyond allowed permissions, OpenAI directs researchers to its existing Security Bug Bounty instead. The company also says it may run private bounty efforts around specific harm areas, including past examples tied to biorisk content in advanced agent products and GPT-5.
Key details at a glance
- OpenAI has launched a public Safety Bug Bounty program
- The program is hosted through Bugcrowd
- It complements the existing OpenAI Security Bug Bounty
- Main focus areas include agentic abuse, proprietary information exposure, and account integrity
- Some agentic reports must be reproducible at least 50% of the time
- Generic jailbreaks and low-impact policy bypasses are excluded
- Traditional unauthorized-access bugs should go to the Security Bug Bounty instead
Safety Bug Bounty vs. Security Bug Bounty
| Program | Main focus | Example issues |
|---|---|---|
| Safety Bug Bounty | AI abuse and safety risks | Prompt injection, agent hijacking, data exfiltration, unsafe agent behavior, proprietary information exposure |
| Security Bug Bounty | Traditional security flaws | Unauthorized access, broken permissions, exposed data, platform vulnerabilities |
Why researchers and enterprises should pay attention
For researchers, the program opens a paid reporting channel for AI risks that may have fallen into a gray area before. For enterprises, it signals that OpenAI sees agent misuse, data leakage, and abuse-driven failures as a distinct attack surface that needs its own reporting and triage system.
The broader message is that AI security now includes more than model weights, APIs, and infrastructure. It also includes how models, agents, and account systems behave under manipulation. OpenAI’s new bounty formalizes that view and puts public incentives behind it.
FAQ
It is a public bug bounty program focused on AI abuse and safety issues that may not qualify as traditional security vulnerabilities but can still cause meaningful harm.
OpenAI says the Safety Bug Bounty runs on Bugcrowd.
OpenAI lists agentic risks such as prompt injection and data exfiltration, exposure of proprietary OpenAI information, and account or platform integrity weaknesses.
Generic jailbreaks that only produce rude language or public information are excluded, as are general policy bypasses without a clear safety or abuse impact.
Read our disclosure page to find out how can you help VPNCentral sustain the editorial team Read more
User forum
0 messages