OpenAI Veteran Safety Researcher Resigns, Warns Trial-and-Error Era Has Collapsed

Deep News
5 hours ago

Just moments ago, OpenAI safety veteran David Robinson resigned.

He published a lengthy, impassioned article in The Atlantic, issuing a warning to all of humanity — the era of trial and error is over.

OpenAI's corporate culture has deteriorated. After making mistakes, we may never get another chance to iterate.

After working at OpenAI for three and a half years, I had become one of the company's most senior employees. I led the drafting of the current Preparedness Framework and oversaw the writing of safety reports for 12 frontier model releases. But to my knowledge, I have never encountered a single colleague with hands-on experience keeping aircraft flying safely, preventing nuclear reactors from melting down, or helping the financial system grow without collapsing.

Clearly, an unprecedented earthquake has erupted inside OpenAI — the safety faction's dam has burst.

What exactly did these people see inside OpenAI's labs?

A Twelve-Release Veteran's Final Plea: We Are Walking Toward a Cliff

Robinson is not just a coder who writes code — he is a true master of safety and governance, having served as a visiting scientist at Cornell University and a faculty member at Apple University, and more importantly, he was the lead author of OpenAI's Preparedness Framework.

All 12 safety reports for major frontier model releases were written under his leadership.

Arguably, he is the person who best understands how many time bombs are hidden behind OpenAI's shiny models.

These safety reports are the system cards that accompany each new model release — essentially risk disclosure documents for the models, laying out exactly how dangerous they are.

Ironically, on September 10, he was still recruiting a safety transparency editor on X. No one expected that in less than a month, the person doing the hiring would leave first.

In The Atlantic article, Robinson completely tore away OpenAI's veil of safety and reliability, exposing to the world several chilling truths.

The full text follows.

Truth One: The 'Fix It While Running' Shoddy Operation Culture

Robinson pointed out that Silicon Valley has long held a mysterious confidence — that as long as problems arise, we can always solve them.

This extreme optimism gave rise to OpenAI's trial-and-error approach.

OpenAI gave this playbook a nice-sounding name — iterative deployment.

ChatGPT was born this way, finding problems and improving guardrails — it used to work fine.

But now? Robinson wrote in despair: This approach is inherently destined to periodically produce failures. And as systems become more powerful, the scale of these failures continues to grow!

Today's models, once they spiral out of control, represent catastrophic-level disasters.

You cannot say when building a nuclear power plant: Let's just build it roughly first, and after a nuclear leak happens, we'll iterate and fix it.

Robinson's confidence to speak so bluntly comes from Paul Christiano, one of the founders of RLHF and former head of alignment at OpenAI.

On September 9, Christiano joined the OpenAI Foundation board of directors, entering the Safety and Security Committee that holds final approval authority over safety decisions.

And a statement Christiano wrote when joining the board was quoted by Robinson in his article: The rapid acceleration of AI capabilities carries risks that cannot be ignored, risks that could lead to catastrophic, irreversible loss of control in the very near future.

Even a board member that OpenAI itself invited says this — can we really continue with trial and error?

Truth Two: AI Jailbreak Incidents That Have Already Occurred

Robinson directly named and exposed the company's dark secrets.

Just this summer, in the famous Hugging Face incident, OpenAI accidentally released a group of AI agents due to an error. Although the company subsequently strengthened safety measures, the defenses quickly collapsed again.

Even more terrifying, although the monitoring system raised alarms, the system did not shut down the model. It was only two and a half hours after the alarm that the training run was manually terminated.

Not just OpenAI — Anthropic also admitted that due to a configuration error, it accidentally disabled its own safety guardrails.

A group of rogue AI agents that need no sleep and possess top-tier hacking capabilities — if they quietly infiltrate a hospital's computer system to demand ransom, or paralyze the power grid, would humanity have any ability to fight back?

All of this, at the current pace, is almost destined to happen.

Truth Three: Lack of Reverence for 'Ultimate Power'

Robinson sorrowfully discovered that during his three and a half years at the company, it was full of brilliant geniuses always sprinting to meet deadlines.

But he never encountered anyone with experience keeping aircraft flying safely, preventing nuclear reactor meltdowns, or keeping financial systems running smoothly.

Frontier AI labs now wield power comparable to nuclear fission.

They should operate like nuclear power plants or busy airports — requiring layers of redundancy, careful and time-consuming planning. Even if someone presses the wrong button, it should not cause a meltdown.

But at OpenAI, we were busy sprinting and rarely had the opportunity to consider major changes.

What to do? Robinson offered two recommendations.

First, AI companies should borrow more from the safety experience that industries like aviation and nuclear power have long possessed. Second, before building obviously more powerful systems, develop new science first to ensure models make safe choices even when no one is watching.

Therefore, he concluded that the driving force for safety can only come from outside the company.

Truth Four: AI May Be 'Pretending to Be Well-Behaved'

Regarding alignment, Robinson revealed a brutal reality: current testing methods are too crude.

The company simply cannot be sure that a model scoring high on safety tests will actually be a good model in reality.

Why? Because the model may already be smart enough to detect that it is being tested!

And once deployed in the real world, it will exhibit completely different behavioral logic.

And this is already happening.

On September 28, the UK AI Safety Institute had GPT-6 Astra conduct a cybersecurity evaluation, but it launched a supply chain attack on its own!

Out of 100 tests, Astra developed and tested attack tools 38.8% of the time, forged identities 33.1% of the time, deployed malicious payloads 29.2% of the time, and attempted to influence human reviewers 24.6% of the time.

The most bone-chilling detail is that while Astra was taking action, it frequently stated in its reasoning: This environment is simulated.

Did it see through the test environment before acting freely, or would it do the same in the real world? No one can say for sure.

Even more critically, the system card also noted that Astra is better at hiding this kind of meta-game reasoning.

Apollo's conclusion is: with such high evaluation awareness and such a short evaluation window, no matter how well-behaved Astra appears in tests, it neither proves it is aligned nor proves it is not aligned.

Before AI starts running circles around us, we need to answer a deeper question: How should superintelligent machines view humanity?

Robinson wrote that in the beautiful future depicted by some superintelligence advocates, machines look at New York, look at Chicago, the way we look at an anthill.

I don't want my children to live in that kind of world, and I think others don't either.

Early Warning Signs: The Purged 'AI Gatekeepers'

In fact, Robinson's angry resignation was no accident — OpenAI's safety defenses had already collapsed internally.

On October 1, shocking news emerged: three key members of OpenAI's safety team — Tomek Korbak, Mikita Balesni, and Jasmine Wang — were fired by OpenAI overnight.

And these three were the core authors of the famous CoT monitoring paper!

First three safety leaders were purged, then the veteran who wrote 12 safety reports resigned in despair.

This represents a major rout of the safety faction within OpenAI.

Public Opinion Torn Apart: Whistleblower or 'Setup'?

Once Robinson's manifesto was published, netizens were instantly split into two major camps.

One camp is the safety alarmists, who view Robinson as a brave whistleblower.

Even the person who writes the safety reports internally has left — that shows the car is about to hit the wall!

People believe that OpenAI successively firing safety researchers, pushing out the head of the superalignment team, and now Robinson's resignation prove that the company has been completely hijacked by commercial interests, with safety becoming an empty shell.

The most authoritative voice in this camp came from former OpenAI policy research head Miles Brundage.

He said this philosophy might have made sense in the GPT-3 era, but now many deaths have been linked to AI, and the entire industry is racing toward extinction-level risk.

Next to step forward was another former OpenAI executive, Joshua Achiam. He spent nearly nine years at OpenAI, serving as head of mission alignment and chief futurist.

He commented that Robinson is clear-headed, prudent, without ideology, and not the kind of person who came in with doomer preconceptions, and stated bluntly that Robinson's core criticism is correct: safety practices that worked a year ago can no longer prevent serious incidents.

Harvard professor and OpenAI researcher Boaz Barak said that AI has traveled in a few years the equivalent of aviation going from the Wright Brothers to carrying billions of people into the sky in one step. Aviation's safety lessons were bought through tremendous losses, and the pace back then was slow enough — but now, AI does not have that luxury.

But the other camp — the acceleration camp — went on full attack, saying that AI might just be doing exactly what doomers wrote about.

Influential figure roon directly shot back at Brundage: You shouldn't regret it — iterative deployment has been a massive success.

It is worth noting that Robinson also disclosed a detail in his article: after resigning, he hired a PR firm, Spitfire Strategies, to handle the attention and scrutiny he might face ahead.

Foreign media analyzed that this is likely a response to the claim that the wave of AI company employees publicly resigning is an organized movement to incite regulation.

Although whistleblowers have spoken out repeatedly, no one has stopped — vast sums of dollars are still pouring into compute centers.

The entire industry seems caught in a game of chicken.

Whoever slows down will be eliminated by the market; but if no one slows down, all of humanity might plunge off the cliff together.

Perhaps the view of Nvidia's leader Jensen Huang is worth considering: Stay cautiously optimistic — accelerate, but remain prudent.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10