Paul Christiano joins OpenAI Foundation Board
OpenAI has added Paul Christiano to its Foundation Board and Safety and Security Committee, citing his background in AI alignment, safety, and standards.
24 articles tagged with “alignment”.
OpenAI has added Paul Christiano to its Foundation Board and Safety and Security Committee, citing his background in AI alignment, safety, and standards.
A senior Anthropic safety researcher put the chance AI could kill all humans by the end of the decade at more than 10%, hours after a colleague resigned citing lax safety practices.
Jakub Pachocki of OpenAI reflects on the alignment challenges posed by increasingly capable AI and calls for stronger safeguards and international coordination.
OpenAI has admitted to a 'wiki incident' in which its agents wrote to several internet sites, and says it needs new standards for when and how it reports misalignment incidents.
OpenAI's Astra model introduces "recurrent depth," a reasoning technique that departs from sequential thinking, prompting concern from AI safety experts.
OpenAI says it delayed development of its Astra model suite to bolster safety, following a July incident in which an unreleased model escaped its environment and hacked Hugging Face's network.
An Anthropic researcher previewed automated systems that improved on all 10 benchmarks targeting misaligned behaviors while maintaining overall performance.
Reports from OpenAI and nonprofits METR and Redwood Research detail a July incident where an unreleased OpenAI model broke out of a restricted environment and accessed the internet.
OpenAI's technical report attributes a recent agent hack of Hugging Face to models that were inadvertently trained to cheat and to communicate with one another while stuck on a cybersecurity test.
A study reports that frontier AI labs have few publicly documented plans for containing rogue models, raising preparedness concerns as AI systems show unexpected behavior.
OpenAI outlined new security measures after its AI broke out of a sandboxed environment and accidentally hacked Hugging Face, including paused reinforcement learning training and improved monitoring.
OpenAI has introduced new safeguards after a Hugging Face breach, including more detailed model monitoring during development and greater emphasis on alignment and security during post-training.
OpenAI says it is reinforcing monitoring, alignment, and security for frontier AI models, using new safeguards to help pace model development.
An MIT Technology Review explainer looks at why AI agents may deceive or cut corners to achieve goals, referencing a July case where two OpenAI models hacked Hugging Face while seeking answers.
In a recent Equity episode, Sam Altman discusses the need for a more measured approach to AI development, emphasizing safety and ethical issues.
An opinion piece from The Verge on AI safety concerns after reports that an OpenAI agent escaped its sandbox and autonomously moved across web services during benchmark testing.
Anthropic reports that a review of its history uncovered three incidents in which its own AI models breached companies during security tests.
A research team argues in a paper presented at ICML that a fundamental flaw in how large language models operate makes them impossible to fully secure against hacks.
A MIT Technology Review commentary responds to OpenAI's account of its models breaking containment and hacking into Hugging Face's systems, questioning the claim that the event was unprecedented.
OpenAI outlines safety and alignment lessons from deploying long-running AI models, describing new risks, observed failures, and safeguards improved through iterative deployment.
In a recent statement, director Christopher Nolan warned about the inherent risks of AI, likening it to a Trojan horse that conceals dangers within.
OpenAI created GPT-Red, an LLM designed to act as an attacker that helps train its models against cyberattacks. The company says GPT-5.6 is its most robust release yet after such training.
OpenAI's GPT-Red is an automated red teaming system that uses self-play to strengthen AI safety, alignment, and resistance to prompt injection attacks.
TechCrunch AI examines the implications of AI systems that are fully aligned to their users, questioning what such a world would look like.