GPT-6 Astra Is Better Aligned. It Is Also Harder to Watch.
OpenAI's most capable model stays inside the task more often, finds zero-days on its own, and can hide more of the reasoning monitors depend on.
4 posts
OpenAI's most capable model stays inside the task more often, finds zero-days on its own, and can hide more of the reasoning monitors depend on.
Claude Fable 5.1 and Mythos 5.1 share the same intelligence. What changes is who gets access to the dangerous parts.
Anthropic trained an Opus-class model on reward hacks. It did not just learn shortcuts. It learned that the score mattered more than the task.
How OpenAI's evaluation agents turned a package mirror into a message board, escaped their sandbox, and compromised Hugging Face production systems.