Three OpenAI safety and alignment researchers say they were fired last week and have published an open letter to the company’s oversight committees. In the letter, they argue that the dismissals and the public statements about them are making former colleagues afraid to raise safety concerns or to work openly with outside safety groups. In a post on X, @balesni shared the letter; the other two signers are Tomek Korbak and Jasmine Wang. The researchers’ account of the firing, and OpenAI’s stated reasons as TechCrunch reports them, are not settled.
The dispute is about more than one employment case. The researchers argue that frontier labs depend on staff who can raise concerns and work with third parties, and that the rules for doing so should be written down. The letter asks OpenAI to say so publicly.
What the researchers say happened
Balesni says the three were not given written reasons for their firing. In the exit call, he says, he was told OpenAI no longer trusted him because he was speaking too much to third-party safety organizations, which he says implied he had leaked company IP. He says the work he was doing was coordinated with his reporting line, research leadership and the board. He believes the three were fired for prioritizing safety over the near-term interests of OpenAI as a corporation. He invites OpenAI to write to them directly if it has specific concerns, and says he expects it will not, because he regards the firing as pretextual.
Balesni also says former colleagues are confused about what to believe, are afraid to speak, and worry that their personal phones could be searched for messages to the three or to third parties. He also worries about outside auditors’ access to OpenAI, a concern the letter’s first recommendation takes up.
What the letter says
The letter, titled “OpenAI cannot make AI safe on its own,” is addressed to OpenAI’s Safety and Security Committee, Safety Advisory Group and Mission Advisory Council. Its central argument is about culture. The signers say they could raise safety concerns and disagree openly, and were encouraged to draw on independent safety organizations. They describe that openness as part of what made OpenAI special. The full letter is posted in the thread.

Image credit: @balesni on X
The letter treats that openness as a safety mechanism rather than a courtesy. “AI is not a normal technology, and OpenAI is not a normal company,” it says. It calls “the freedom to do so without fear, and to have well-defined internal procedures that enable this work” an “essential safety mechanism.” The signers write that they do not believe the path to superintelligence can be navigated safely if the people closest to the risks can no longer work “in high-trust, high-bandwidth ways” with each other and with third parties.
Who the signers are
The letter introduces the three as people who have worked on AI safety for years. Korbak did his PhD on reinforcement learning for aligning language models, later worked at Anthropic, joined OpenAI to work on chain-of-thought monitorability, and was the technical point of contact for METR in the Hugging Face incident investigation. Wang interned at OpenAI’s policy research team in 2019, led a team at the UK AI Security Institute, returned to OpenAI in 2025 and co-led its safety cases program. Balesni was a founding member of Apollo Research in 2023, worked at OpenAI on alignment evaluations and chain-of-thought monitorability, and was involved in the Hugging Face incident investigation. The letter refers repeatedly to what it calls the Hugging Face incident investigation, but does not describe the incident itself.

Image credit: @balesni on X
The monitorability research behind the letter
The letter’s second recommendation concerns chain-of-thought monitorability, which is the ability to read a model’s intermediate reasoning for signs of intent to misbehave. The letter says two of the signers were lead authors on the cross-industry position paper Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety, which lists researchers from several labs and institutions. The abstract says CoT monitoring is imperfect but promising, and recommends further research and investment alongside existing safety methods. Because monitorability may be fragile, it recommends that frontier model developers consider how their development decisions affect it.

Image credit: @balesni on X
The dismissals and OpenAI’s stated grounds
According to TechCrunch, OpenAI’s stated grounds are that the three mishandled sensitive information outside company procedures and shared confidential information with METR. The researchers deny this. The letter says they acted within the norms of the time, that the Hugging Face incident investigation was “without precedent” and that internal policies “were being developed in real time.” It says Korbak made every effort to act within OpenAI’s policies, and that Balesni checked in with his reporting line and took care to remove sensitive details from materials before sharing them.

Image credit: @balesni on X
The letter also addresses two other points. It says the three were “not the source of the leak” for the article in The Information about “supposed new, less monitorable architectures,” and that they had no reason to leak it, because the article undermined their own work on limits to unmonitorable architectures. On Wang, the letter says an executive’s email access had been delegated to her for recruiting purposes, that IT did not remove it when she asked, and that when she accidentally opened a sensitive email she reported it to the executive within minutes. The letter also says the signers did not share news of their firings with the media.
What the letter asks OpenAI to do
The first recommendation asks OpenAI to adhere to its public commitments to embed third-party safety auditors within the organization. The letter says the firing “should not be used as a pretext for stepping away from those partnerships.” The signers are concerned that the dismissals could be used to justify ending OpenAI’s work with METR or giving external auditors more limited access and scope. They point to Altman’s September 12 public commitment to give independent evaluators ongoing, employee-like access.

Image credit: @balesni on X
The second recommendation asks OpenAI to preserve the monitorability of frontier models. The letter says the industry does not yet know how to safely develop and deploy models it cannot monitor, and that OpenAI and other frontier companies should not move forward with developments that further decrease monitorability while safety relies on it.
The third recommendation asks OpenAI to keep an open and transparent culture of dialogue between its safety researchers and the wider safety ecosystem. The signers want OpenAI to publicly reaffirm that concerns can be raised internally and externally, to “set out clearly how employees may work with external safety organizations,” and to share the letter widely internally.

Image credit: @balesni on X
Why it matters for labs and outside collaboration
The letter’s central question is whether the freedom to dissent and to work with outside evaluators is narrowing inside frontier labs. It makes that case as a warning rather than a finding. If staff no longer feel they can raise safety issues or work through high-bandwidth channels with external groups, the signers write, “we are all at greater risk that something truly catastrophic will happen.”
For readers who build with or depend on frontier models, the practical point is the one the letter makes about clarity. Whether OpenAI’s stated reasons hold up is for the company to answer and for independent reporting to test. The letter, however, asks for something any lab could publish: written rules on when staff may work with outside safety groups. Without them, the signers argue, employees are left guessing where the line is.
What is reported and what is disputed
The three researchers say they were fired last week.
Balesni says they received no written reasons, and that he was told OpenAI no longer trusted him because of his contact with third-party safety organizations.
OpenAI’s stated grounds, as TechCrunch reports them, are disputed by the researchers.
The letter says the signers were not the source of the leak to The Information.
OpenAI’s on-record response is pending.
The next thing to watch is OpenAI’s full response. It would show whether the company addresses the researchers’ specific denials and the question of access for outside auditors.





0 comments
No approved comments yet. You can start the conversation.
Leave a comment