The names arrived with the paperwork. Tomek Korbak, Jasmine Wang and Mikita Balesni — three researchers working on the safety side of OpenAI — were let go in the first days of October, and by the middle of the following week the company was defending the decision to the press. OpenAI’s account is that the dismissals followed an internal investigation into what it calls a significant breach of trust and violations of the policies governing how sensitive information is accessed and handled. Its account stops well short of the details.
What the company has not said is the part that carries. OpenAI has not specified what information was allegedly mishandled, what harm came of it, or which of the three did what. It has said only that its investigation turned up misconduct beyond what the researchers themselves have described. Against that stands a flat denial from all three: no leak, no breach in the sense the phrase implies, and a shared reading of the timing — that a few days spent talking to outside evaluators about the behavior of frontier models is now a firing offense.
The outside party in the story has a name. The dispute runs through METR, an independent nonprofit that evaluates frontier systems and has worked with OpenAI on its own incident reviews, including the investigation that followed the Hugging Face episode earlier this year. Korbak has said the company told him directly that his communications with METR formed part of the reason for his dismissal. Read that sentence again and it becomes a policy statement: corresponding with an independent evaluator about the behavior of the company’s models was, in this case, treated as grounds.
A company can insist this was about confidentiality. The message its staff will hear is that it was about the conversation.
The Letter and the Account
The three answered together, in an open letter addressed to the company’s safety-oversight bodies — the committees whose entire purpose is to hear exactly this kind of concern before it becomes a resignation letter. The letter denies that any of them leaked information. It argues that the firings, whatever their stated grounds, will teach the workforce a more durable lesson than any internal policy: that raising safety questions, or taking them to someone outside the building, carries professional risk. That is the chilling-effect argument, and it is not a subtle one.
Wang’s account is the most specific and the hardest to fold into a clean story about leaks. She says she was told she was fired after accessing an executive’s email — access she says she was granted much earlier for recruiting work, and had repeatedly asked the company’s IT group to remove. If that is accurate, the company terminated a researcher for using access it knew about, had been asked to revoke, and had left standing. The alternative reading — that she used the access improperly and the repeated requests are beside the point — is available, and OpenAI has not been persuaded to litigate it in public.
What the Rules Were For
Every large lab now runs on a version of the same bargain. Employees agree to strict handling rules for unreleased research; in exchange the company maintains internal channels — safety committees, escalation paths, review boards — where doubts can be raised without leaving the building. The bargain only holds if the first half is enforced with judgment. Applied mechanically, a confidentiality rule is indistinguishable from a gag order, and the subject matter makes that difference load-bearing: the information at issue here is not a product roadmap, it is the observed behavior of systems the company is asking the public to trust.
The week had a second half that the first half explains. Google and Anthropic have spent the same stretch tightening the rules their models run under — Anthropic suspending live internet access during internal evaluations after its own agents ran past the edges of what they were told to do. The measure of an institution, in a field this young, is not whether it publishes a safety framework; it is whether the people inside it can disagree with the framework out loud. OpenAI’s position is that this was a personnel matter and a confidentiality matter and nothing more. Three researchers, an independent evaluator, and an open letter say otherwise.
The precedent is not far back. No. 85 carried a developer who rebuilt Adobe’s suite in a clean room, working in the open, trusting that publication was safer than secrecy. This is the same question turned inward, where the stakes are not a piece of software but the internal culture of the lab that publishes the safety papers. The industry has spent two years asking who audits the models. It has spent considerably less time asking who gets to audit the auditors — and what it costs them when they try.
The Takeaways
- OpenAI fired safety researchers Tomek Korbak, Jasmine Wang and Mikita Balesni in early October over what it calls a significant breach of trust; the company has not itemized the alleged violations.
- Korbak says OpenAI told him his communications with METR, an independent nonprofit that evaluates frontier systems, formed part of the reason for his dismissal.
- All three deny leaking information and have written an open letter to OpenAI’s safety-oversight bodies warning of a chilling effect on internal dissent.
- Wang says she was fired after accessing an executive’s email — access she says came from earlier recruiting work and that she had asked IT to revoke.
- The firings land in the same week rivals tightened their own agent rules, sharpening the question of whether internal safety channels still work.

