OpenAI safety employee resigns

David Robinson, a former OpenAI safety employee who helped develop the company’s Preparedness Framework and worked on safety reports for 12 frontier-model launches, has resigned and publicly criticized the company’s approach to AI safety.

OpenAI safety employee resigns

In an essay published by The Atlantic on October 3, Robinson argued that AI companies are moving too quickly and are not being careful enough as increasingly capable systems are developed and deployed.

His central warning was blunt: “The time for trial and error is over.”

Robinson said the AI industry can no longer depend primarily on deploying increasingly capable systems, observing what happens and then improving safeguards after problems emerge. In his view, the potential consequences of failures become more serious as AI systems gain greater capabilities and autonomy.

Why David Robinson left OpenAI

Robinson said he resigned because he believes OpenAI’s culture places too much emphasis on rapid AI development and not enough on the level of caution required for increasingly capable systems.

During his three-and-a-half years at OpenAI, Robinson said he helped draft the company’s Preparedness Framework and oversaw safety reports connected with 12 frontier-model launches.

His criticism is therefore significant because it comes from someone who was directly involved in OpenAI’s safety work rather than from an outside observer.

In his resignation essay, Robinson argued that AI companies are not being sufficiently cautious about the risks associated with increasingly powerful systems.

Robinson criticizes iterative AI deployment

A major part of Robinson’s criticism concerns the idea of iterative deployment.

The approach allows AI systems to be released, monitored in real-world environments and improved as researchers discover new problems. Companies can strengthen safeguards as they learn more about how their systems behave.

Robinson believes that approach becomes increasingly difficult to justify as AI capabilities advance.

His concern is that an approach based partly on discovering and correcting problems after deployment necessarily accepts some degree of failure. If AI systems become more autonomous and capable, however, the consequences of those failures could also become more significant.

That is the context behind his warning that the industry’s period of trial and error should come to an end.

Why aviation and nuclear safety are part of the comparison

Robinson has argued that frontier AI companies should learn from industries where mistakes can have extremely serious consequences.

He pointed to fields such as aviation and nuclear power, where safety systems typically rely on multiple layers of protection, redundancy and carefully designed procedures.

The underlying principle is that a single human mistake should not automatically result in a catastrophic failure.

Robinson believes advanced AI development needs a similar approach, with stronger safeguards built into the development and deployment process rather than relying primarily on improvements after systems are already in use.

The AI alignment problem

Robinson also highlighted the continuing challenge of AI alignment.

AI alignment broadly refers to the effort to ensure that AI systems behave consistently with human intentions, goals and safety requirements.

As models become more capable, researchers face a difficult question: how can they establish that a system will continue behaving safely when it encounters situations that were not represented in controlled testing?

Robinson’s argument is that increasingly capable AI systems require stronger scientific methods for evaluating these risks before deployment.

Passing a particular safety evaluation does not necessarily prove that an AI system will behave safely in every possible real-world environment.

OpenAI says it continues to strengthen safety

OpenAI has rejected the suggestion that it is ignoring safety concerns.

An OpenAI spokesperson told Reuters that the company is working to ensure its models do not become more capable than it can safely manage and secure.

The company also said that it can pause training or hold back models when it determines that additional caution is necessary.

That response highlights the central disagreement between Robinson and OpenAI.

OpenAI maintains that its safety processes can evolve alongside increasingly capable models. Robinson argues that the pace of AI development requires much stronger safety practices before capabilities advance further.

Why the resignation matters

Robinson’s departure adds to a broader debate about whether AI companies are advancing capabilities faster than their safety research and governance systems can keep pace.

Modern frontier AI systems are increasingly capable of handling multi-step tasks, using tools and operating with greater autonomy. As these systems become more powerful, safety concerns can extend beyond inaccurate answers or unreliable outputs.

The more important question becomes whether an AI system could take actions that its developers did not anticipate.

That is why Robinson’s criticism focuses not only on individual models but also on organizational culture and the engineering practices used to develop advanced AI.

He argues that frontier AI laboratories should draw lessons from established high-risk industries and develop stronger scientific methods for determining whether increasingly autonomous systems can remain safe.

What Robinson’s resignation does — and does not — establish

Robinson’s resignation is significant evidence of disagreement over AI safety within the industry, but it does not independently establish that OpenAI’s current safety systems are unsafe or that a catastrophic AI incident is imminent.

His claims represent his assessment of OpenAI’s culture and safety approach.

OpenAI, meanwhile, says it continues to strengthen its safeguards and can slow down or hold back systems when safety considerations require it.

The larger question remains unresolved: How much safety certainty should frontier AI companies achieve before deploying substantially more capable systems?

As AI models become increasingly autonomous, that question is likely to remain at the center of the industry’s safety debate.

Scroll to Top