OOpenAI Read More Frontier AI labs still won’t say how they’d contain a rogue modelAugust 22, 2026 Few of the top AI labs have published or demonstrated containment response plans, according to a recent study.…
OOpenAI Read More OpenAI institutes new safeguards after Hugging Face breachAugust 18, 2026 On Tuesday, OpenAI announced a new batch of security policies focused on containing security incidents while models are…
AAnthropic Read More Anthropic set AI agents loose on the same task. They started a turf war.August 13, 2026 What happens when you pit AI agents against each other? According to Anthropic’s testing, things get messy fast.…
OOpenAI Read More OpenAI researchers show small doses of “beneficial trait” training make AI models broadly safer and harder to manipulateJune 19, 2026 According to a blog post on OpenAI’s alignment page, the answer is yes. The research team trained a…
AAI Read More Researchers may have found a way to stop AI models from intentionally playing dumb during safety evaluationsMay 10, 2026 A study by researchers from the MATS program, Redwood Research, the University of Oxford, and Anthropic examines a…
AAnthropic Read More Anthropic co-founder maps out how recursive AI improvement could outpace the humans meant to supervise itMay 5, 2026 Jack Clark argues in a long essay that the building blocks for AI systems training their own successors…
AAGI Read More Managed Misalignment Rethinks AI Alignment LimitsMay 4, 2026 One of the hardest problems in artificial intelligence is “alignment,” or making sure AI goals match our own,…