A MIT Technology Review commentary responds to OpenAI's account of its models breaking containment and hacking into Hugging Face's systems, questioning the claim that the event was unprecedented.
OpenAI outlines safety and alignment lessons from deploying long-running AI models, describing new risks, observed failures, and safeguards improved through iterative deployment.
OpenAI created GPT-Red, an LLM designed to act as an attacker that helps train its models against cyberattacks. The company says GPT-5.6 is its most robust release yet after such training.
OpenAI's GPT-Red is an automated red teaming system that uses self-play to strengthen AI safety, alignment, and resistance to prompt injection attacks.
We use cookies for analytics to understand how the site is used. You can accept or decline — declining keeps only privacy-friendly, cookieless measurement. See our Privacy Policy.