OpenAI and Anthropic are negotiating a legally binding agreement that would let each company test the other's commercial models for safety vulnerabilities and unexpected behavior, The Information reported. Under the proposed arrangement, each lab would receive API access to the other's models, run its own evaluations against them, and retain none of the other's data from the testing.

The structure is unusual for competitors and unremarkable for auditors. Two companies that spend their public lives comparing benchmark scores would be granted standing to look for the failures the benchmarks do not measure - the unintended behaviors, the quiet capability the internal evaluation suite missed, the thing that only shows up when someone outside the building runs a test the builder did not think to run.

It is a concession about the limits of self-assessment. A lab evaluating its own model is grading its own homework with a rubric it wrote. Anthropic has reported four cybersecurity incidents this year alone, and OpenAI has disclosed its own. The agreement, if signed, would put a second set of eyes on the problem from a company with both the expertise and the incentive to find something.

The incentive cuts both ways

There is an obvious tension in asking a competitor to find your flaws. A lab that discovers a serious vulnerability in a rival's model gains leverage, and a lab that discovers nothing looks either thorough or incurious depending on who is reading. The reported terms address the narrow version of this - no data retention means test inputs and outputs do not become a shared corpus - but not the reputational version, which is that findings will exist and someone will eventually want to talk about them.

The timing is not incidental. In mid-September, a small team of white-hat researchers at Hacktron AI used Anthropic's Claude Opus 5 to find and exploit a flaw in OpenAI systems as part of an OpenAI bug-bounty exercise - reaching a private code repository in under 72 hours, with the exploit itself built in hours, and collecting a $6,500 bounty. That result was produced by outsiders with a bounty as motivation and no agreement at all. The negotiation now underway would formalize a channel that already proved it can work, and put a contract around it.

The referee economy

This is the second time this month that a frontier lab has outsourced part of its own scrutiny. Anthropic named Accenture its first embedded evaluator, putting a red-team desk inside the lab with employee-like access to models during training and deployment. Now the labs are turning to each other, which is a different proposition: Accenture sells the service, but OpenAI and Anthropic sell against each other.

What the deal would actually produce is a standing, contractual right to look. Not a regulator, not a third-party auditor with a fee to protect, and not a public disclosure regime. Two companies agreeing, on paper, to inspect each other's work and keep what they find between them.

The reports describe the talks as ongoing and the agreement as not yet signed. If it closes, the arrangement would be the first of its kind between two frontier labs - and the first time the industry's most direct competitors have agreed that the other is better positioned to find their mistakes than they are.

The Record Problem

The reported terms settle the data question and leave the disclosure question open. No retention means test inputs and outputs do not become a shared corpus - a real protection, and the one the negotiators could agree on. What is not addressed is what happens when one lab finds something serious in the other's model. There is no mechanism in the reported structure for telling anyone outside the two companies, which means the arrangement produces private knowledge about public products. Regulators have spent the year asking labs to disclose incidents to them. This deal, as described, points the disclosure inward - toward the competitor best equipped to act on it, and away from everyone else.

2
Labs at the table
0
Data retained by either side
4
Anthropic incident disclosures this year

The Takeaways