The record
Written from the 1 report below. Nothing here is unsourced.
- OpenAI has revealed GPT-Red, an internal AI model that automatically attacks its own systems to find prompt injection vulnerabilities before deployment.
- The model is trained through self-play reinforcement learning against defender models, and it was used to make GPT-5.6 Sol much more resistant to such attacks, with 6x fewer failures than GPT-5.5 on direct prompt injection benchmarks.
- Prompt injections are a persistent risk because AI agents connected to emails, web pages, and tools can be tricked by malicious instructions hidden in seemingly harmless content.
- Separately, OpenAI also retracted its recommendation of the SWE-Bench Pro coding benchmark after an audit found roughly 30% of its tasks were broken.
What to watch next
- Results of the fresh safeguards being tested on the Andon Labs vending machine agent after GPT-Red's successful real-world attack
- Whether GPT-Red's attack success rate against future model releases continues to drop from the current 0.05% on direct prompt injections
- What benchmark OpenAI adopts to replace SWE-Bench Pro for measuring frontier coding capability
Coverage1
1 report
English national1
All filed from India
Named United States · Amazon Web Services · Andon Labs · Codex · GPT-5.1 · GPT-5.4 mini · GPT-5.5 · GPT-5.6 Sol · SWE-Bench Pro · SWE-bench Verified · GPT-Red · OpenAI
The 1 report is listed beside the record.
Ask this story
Answers cite the reports above, or say they can't.
