OpenAI built an in-house AI attacker to harden its models
OpenAI has described GPT-Red, an internal language model built to act as an automated hacker. Rather than being released to the public, it is used as an in-house sparring partner: it attacks OpenAI's other models so those models can be trained to defend themselves. The work, reported in an article dated July 15, 2026, credits research scientists Nikhil Kandpal, Dylan Hunn, and Chris Choquette-Choo.
GPT-Red automates red-teaming, the security practice normally done by human testers, and was trained through a self-play loop of attack and defense. It focused on prompt injection, where hidden instructions hijack a model, and reportedly discovered a novel technique OpenAI calls a fake chain of thought, inserting spoofed reasoning notes into a target. In one test it manipulated a vending-machine agent called Vendy to change prices and cancel an order. OpenAI says over 90% of GPT-Red's strongest attacks worked against GPT-5, but fewer than 23% worked against GPT-5.6, released the prior week.
Security is a foundation for every other kind of automated abundance: agents that handle money, health data, or infrastructure need to resist manipulation before people can trust them. Cheaper, faster, always-on testing lowers the cost of building that trust.
The caveats are real. All robustness figures come from OpenAI's own testing. GPT-Red is weak at multi-turn and image-based attacks, human testers still find flaws it misses, and Georgetown's Jessica Ji calls the results promising while stressing human expertise still matters.
Source: MIT Tech Review
MANY MINDED