AI Europe
  • Europe
  • Europa
  • Britain
  • France
  • Germany
  • Italy
  • Spain
  • Japan
  • Canada
  • Africa
  • People
  • AI
  • …
    • Afrique
    • Netherlands
    • Poland
  • Agentic AI
  • AGI
  • AI
  • Anthropic
  • Google
  • Microsoft
  • OpenAI
  • xAI
AI Europe
  • Europe
  • Europa
  • Britain
  • France
  • Germany
  • Italy
  • Spain
  • Japan
  • Canada
  • Africa
  • People
  • AI
  • …
    • Afrique
    • Netherlands
    • Poland

Browsing Tag

alignment

7 posts
OOpenAI
kill switch ransomware
Read More

Frontier AI labs still won’t say how they’d contain a rogue model

  • August 22, 2026
Few of the top AI labs have published or demonstrated containment response plans, according to a recent study.…
OOpenAI
OpenAI reportedly finds evidence that more of its agents ran amok
Read More

OpenAI institutes new safeguards after Hugging Face breach

  • August 18, 2026
On Tuesday, OpenAI announced a new batch of security policies focused on containing security incidents while models are…
AAnthropic
Anthropic set AI agents loose on the same task. They started a turf war.
Read More

Anthropic set AI agents loose on the same task. They started a turf war.

  • August 13, 2026
What happens when you pit AI agents against each other? According to Anthropic’s testing, things get messy fast.…
OOpenAI
Image description
Read More

OpenAI researchers show small doses of “beneficial trait” training make AI models broadly safer and harder to manipulate

  • June 19, 2026
According to a blog post on OpenAI’s alignment page, the answer is yes. The research team trained a…
AAI
Image description
Read More

Researchers may have found a way to stop AI models from intentionally playing dumb during safety evaluations

  • May 10, 2026
A study by researchers from the MATS program, Redwood Research, the University of Oxford, and Anthropic examines a…
AAnthropic
The NSA is using Anthropic's most powerful AI model Mythos
Read More

Anthropic co-founder maps out how recursive AI improvement could outpace the humans meant to supervise it

  • May 5, 2026
Jack Clark argues in a long essay that the building blocks for AI systems training their own successors…
AAGI
Managed Misalignment Rethinks AI Alignment Limits
Read More

Managed Misalignment Rethinks AI Alignment Limits

  • May 4, 2026
One of the hardest problems in artificial intelligence is “alignment,” or making sure AI goals match our own,…
AI Europe
www.europesays.com