OpenAI says it is now working with Anthropic and Google DeepMind on AI safety evaluation - Bloomberg reported the coordination on September 15. The concrete artifact so far is a pilot run this summer in which OpenAI and Anthropic each tested the other's publicly released models with their own internal safety and misalignment evaluations. What does not exist yet is the hard part: shared risk-level mappings, because the labs say apples-to-apples comparison across model families is genuinely difficult.
The pilot is the part worth taking seriously. This summer, each company ran its own safety and misalignment evaluations against the other's publicly released models - cross-examination, not self-report. All three already publish bounded safety cases under their own frameworks: Anthropic's Responsible Scaling Policy, OpenAI's Preparedness Framework, DeepMind's Frontier Safety Framework. The coordination story is that the examiners are starting to trade papers, and the pilot proved the mechanics work.
What the labs conspicuously did not establish is a shared risk-level mapping. OpenAI's own framing is that apples-to-apples comparison across labs and model families is hard - a red line for one lab's eval suite is not automatically a red line for another's. That caveat is honest, and it is also the whole game: thresholds are where safety coordination stops being press-release and starts being constraint. Without a shared scale, each lab still grades its own homework.
The timing lands inside the convergence that built all summer. Amodei's 'pace the frontier' essay made the case that capability is outrunning the ability to check it. Altman told his own staff he was open to slowing cutting-edge work if other labs moved too - the classic arms-race concession that nobody can move first alone. Hassabis has called for international coordination in nearly the same words. Cross-lab evaluation is what that rhetoric looks like at the level where someone actually runs an experiment.
Skeptics will note what this is not. It is not regulation - participation is voluntary and nothing announced so far binds anyone. It is not a pause. And the labs' safety teams still answer to the same executives who ship the models; cross-testing released models is a different and easier problem than inspecting training runs mid-flight, which is where the real contested territory sits. The German government's response to the pace debate - that halting is 'not viable' for Europe and real governance needs Washington and Beijing at the table - also still stands: this covers three American labs, not the frontier.
Still, the direction of travel is the story. Two years ago the labs treated safety cases as marketing adjacent. Last year they published them. This year they started testing each other's released models, and described the coordination in public rather than letting it leak. The next step that matters is a shared scale for what counts as dangerous - the thing they just explicitly declined to promise. Whether coordination gets teeth depends on that map ever existing.
The Teeth Question
What to watch is not the announcement but the first disagreement. Cross-lab evaluation gets its real test the first time one lab's eval flags what the other's cleared - the moment coordination meets competitive incentive. That is also exactly the gap the labs named: without shared risk-level mappings, each lab still grades on its own curve, and a standard nobody shares is a press release with appendices. The second thing to watch is scope: three American labs trading evaluations is not a standard until the labs outside the room - and the governments that actually regulate - treat it as one.
The Takeaways
- OpenAI says it is working with Anthropic and Google DeepMind on cross-lab safety evaluation; a summer pilot had each lab testing the other's publicly released models with its own eval suites.
- It rides the pace-the-frontier convergence - Amodei's essay, Altman's conditional-slowdown, Hassabis's coordination calls - but nothing announced binds anyone.
- The named gap is the real story: no shared risk-level mappings, because apples-to-apples comparison across labs is genuinely hard. Without that scale, each lab still grades its own homework.

