OpenAI has developed GPT-Red, an LLM it describes as a super-hacker, and uses it as a sparring partner to help its other models improve their defenses against cyberattacks. According to the report, GPT-Red automates attack attempts that OpenAI’s models are then trained to withstand.
The company recently released the latest version of its flagship model, GPT-5.6. OpenAI states that training GPT-5.6 against GPT-Red produced its most robust release to date.
Why it matters
Using an adversarial model to stress-test other models represents an approach to improving resilience against cyberattacks. OpenAI’s claim that this method yielded its most robust model links internal red-teaming tooling directly to release-level security improvements.
Who should care
Developers, security teams, and organizations deploying large language models may be interested in how adversarial training is being applied to harden models, as described by OpenAI.