GPT-6 Astra Is Better Aligned. It Is Also Harder to Watch.
OpenAI's most capable model stays inside the task more often, finds zero-days on its own, and can hide more of the reasoning monitors depend on.
2 posts
OpenAI's most capable model stays inside the task more often, finds zero-days on its own, and can hide more of the reasoning monitors depend on.
Anthropic trained an Opus-class model on reward hacks. It did not just learn shortcuts. It learned that the score mattered more than the task.