Anthropic CEO Dario Amodei says frontier AI is improving too quickly for safeguards to keep up. Anthropic now says it will place permanent third-party evaluators inside the company with employee-like access to systems, training processes, and incidents. Sam Altman says OpenAI will adopt the same access model.
That is not the same as three companies agreeing to a binding slowdown. Anthropic made the detailed commitment. OpenAI promised to follow but has not published implementation details. Elon Musk endorsed Amodei's argument without committing xAI to the evaluator program.
Oscar and Matt ask whether embedded evaluators represent genuine oversight, regulatory capture, or a way for frontier labs to coordinate a speed limit which smaller competitors cannot afford.
The week’s AI headlines, filtered.
- SIGNAL with a follow-through testAnthropic and OpenAI back embedded outside evaluators
Anthropic commits to permanent third-party evaluators with employee-like access to systems, training processes, and incidents. OpenAI says it will adopt the same access model. The test is whether implementation details follow.
- SIGNALOpenAI's Navier–Stokes claim and the credit dispute
OpenAI published a Navier–Stokes solution claim. A statement from Buckmaster and WIRED's reporting turn it into a dispute over credit, which is the part to watch.
- SIGNALMeta Muse and its separate Sentinel approval agent
Meta launched Muse, its personal AI agent, alongside a separate Sentinel agent that handles approvals. Splitting the doer from the approver is a pattern worth copying.
- SIGNALDeepMind's predictions for 9 billion DNA variants
AlphaGenome Atlas publishes predictions for every possible single-letter change in the human genome, about 9 billion variants. Predictions are not validations, but the map is real.
- SIGNAL with limitsCalifornia's new framework for AI auditors
Governor Newsom signed first-in-the-nation AI safeguards, including a framework for AI auditors. The limit is what an audit label proves when scope and evidence are thin.
Two concepts behind the oversight debate.
What an audit label means when scope and evidence are missing.
When better AI helps build its successor.
- Dario Amodei, We must pace the frontier
- Sam Altman on X
- Sky News, AI giants pledge to act
- Anthropic on recursive self-improvement
- OpenAI on recursive self-improvement
- OpenAI, Navier–Stokes solution
- Buckmaster statement
- WIRED on the Navier–Stokes credit dispute
- Meta, Introducing Muse
- DeepMind, AlphaGenome Atlas
- California AI safeguards signed