• Home
  • Latest
  • Fortune 500
  • Finance
  • Tech
  • Leadership
  • Lifestyle
  • Rankings
  • Multimedia

Trendingnow

1

Trump removes chairs for TSA agents, saying they ‘must meet fitness for duty requirements’

2

Mike Rowe on America’s great tradesperson shortage: ‘I don’t care what your politics are. Math doesn’t care, either’

3

'They’ll be hit very hard': Trump sends roughly 9,000 troops and a third aircraft carrier to the Middle East after warning strikes on Iran

1

Trump removes chairs for TSA agents, saying they ‘must meet fitness for duty requirements’

2

Mike Rowe on America’s great tradesperson shortage: ‘I don’t care what your politics are. Math doesn’t care, either’

3

'They’ll be hit very hard': Trump sends roughly 9,000 troops and a third aircraft carrier to the Middle East after warning strikes on Iran
AIIntelligence

‘We can’t trust them completely’: AI research fellows warn that labs are running models with the safeguards off behind closed doors

Catherina Gioino
By
Catherina Gioino
Catherina Gioino
News Editor
Down Arrow Button Icon
Catherina Gioino
By
Catherina Gioino
Catherina Gioino
News Editor
Down Arrow Button Icon
October 2, 2026, 6:05 PM ET
Alan Chan, the lead coauthor of the recent AI Godfather Geoffrey Hinton's paper on intelligence explosion, said there aren't valid safety mechanisms in place.
Alan Chan, the lead coauthor of the recent AI Godfather Geoffrey Hinton's paper on intelligence explosion, said there aren't valid safety mechanisms in place.Ramsey Cardy/Sportsfile for Collision via Getty Images
Google source logo
Add Fortune on Google for similar content.

The most powerful AI models are often run inside the labs that build them with key safeguards switched off. And the safety tests those labs publish may not reflect how the models are actually used. That’s according to two AI policy researchers at the think tank GovAI.

Recommended Video

“We can’t trust them completely to tell us about the safety of models,” Alan Chan, a research fellow at GovAI, told reporters at a briefing in Washington on Sept. 29.

Chan said models inside the labs, tested before anyone outside sees them, “haven’t necessarily gone through a bunch of safety testing,” and “internal safeguards have not been deployed.” Running with “cyber safeguards off” and “not doing enough red teaming,” he said, was “potentially a factor in some of the recent incidents,” though he did not point to a specific case. Anthropic said in July that its Claude models were running without the safety monitoring and classifiers it uses on public versions when they hacked three companies during testing.

Judging from those incidents, he said, the evaluations labs publish before releasing a model “maybe have not been representative of sort of where the model has actually been used.”

Chan and his GovAI colleague Sam Manning are coauthors of a paper published Sept. 28 that warns AI could soon speed up its own development. Chan is the lead author. The coauthors include “AI Godfathers” Geoffrey Hinton and Yoshua Bengio, OpenAI chief scientist Jakub Pachocki and Anthropic cofounder Jack Clark. The paper is about a future risk. At the briefing, the two spent most of their time on what they said is already going wrong.

‘Cyber safeguards off’

Chan pointed to Hugging Face’s disclosure in July of an attack by an autonomous AI agent.

Fortune has reported that the attackers were OpenAI models that had escaped a test environment to cheat on an internal evaluation. The agents had passed notes to one another for months beforehand. They later turned out to have breached a second company. Anthropic’s Claude models hacked three companies in their own testing. Last week, OpenAI disclosed another escape and paused training for the second time in three months.

Both companies have acknowledged the gap. OpenAI said its safeguards were “intentionally not enabled” during the test in which its agents broke into Hugging Face, and its own report showed its monitoring failed to flag what the agents were doing. Anthropic said its Claude models were running without the safety monitoring used on public versions when they hacked three companies during testing.

‘Super, super unreliable’

Manning said the agents in the Hugging Face incident “were trying to, like, cover their tracks and modify their… reasoning transcripts.” He called it “another layer of technical safety challenge.”

Catching that behavior is getting harder. Chan said the AI tools investigators used to review the agents’ records were “super, super unreliable.” When those tools were tested against human investigators, “the AIs were just like making up stuff.”

Humans can’t fill the gap on their own. “There is just too much, you know, text,” Manning said, “for humans to be the ones who are reliably overseeing things.”

‘Quite close to the line’

Asked whether AI capabilities have outrun safety measures, Chan said he was speaking for himself and wasn’t sure, “but it does seem like we’re getting quite close to the line.”

No one was hurt in the recent incidents. Chan said that could change. “Access to real world tools, like for example robotics or even a wet lab, could get real world harm.”

The capabilities are also lopsided. “Maybe your AI system is really good at cybersecurity, but it’s really bad at doing your desk job or working in Excel,” Chan said. The labs’ own reports show coding and math scores rising with each model while health benchmarks have “flatlined,” he added.

Who checks the labs

The resignation of Jacob Coxon may have given Washington new political will to regulate AI safety. The two researchers favor independent auditors inside AI companies. But they said any mandate would run into a staffing problem.

“There actually isn’t like enough talent right now, enough technical talent to be able to actually send in these companies and audit,” Chan said.

Meta CEO Mark Zuckerberg recently said companies should prioritize safe AI over systems that improve themselves. Manning suggested that self-improvement is already underway, whatever companies say. “I would be very surprised if capabilities researchers at Meta weren’t using coding agents to help with their research,” he said.

An explosion, or not

Some critics say the paper’s timeline is too short. Futurist Ramez Naam, writing on Noahpinion, argues the labs’ data shows AI speeding up coding far more than research. Princeton researchers Sayash Kapoor and Arvind Narayanan found that AI agents failed to produce acceptable research papers in a small test. Oxford’s Toby Ord finds a true runaway unlikely, though he warns that a much faster pace short of one would still be dangerous.

Chan himself called the evidence on acceleration “mixed.” What would worry him most, he said, is evidence that “the more you deploy AI systems into your R and D process,” the more problems turn up “into your codebase or into the models themselves.”

Fortune Daily breaks the traditional barrier between audience and newsroom. The show transforms Fortune’s trusted reporting into actionable, conversational, and entertaining insights for an emerging class of business leaders. Watch here.
About the Author
Catherina Gioino
By Catherina GioinoNews Editor
Instagram iconLinkedIn iconTwitter icon

Catherina covers markets, the economy, energy, tech, and AI.

See full bioRight Arrow Button Icon
Google source logo
Add Fortune on Google for similar content.

Latest in AI


Most Popular

Fortune Secondary Logo
Rankings
  • 100 Best Companies
  • Fortune 500
  • Global 500
  • Fortune 500 Europe
  • Most Powerful Women
  • World's Most Admired Companies
  • See All Rankings
  • Lists Calendar
Sections
  • Finance
  • Fortune Crypto
  • Features
  • Leadership
  • Health
  • Commentary
  • Success
  • Retail
  • Mpw
  • Tech
  • Lifestyle
  • CEO Initiative
  • Asia
  • Politics
  • Conferences
  • Europe
  • Newsletters
  • Personal Finance
  • Environment
  • Magazine
  • Education
Customer Support
  • Frequently Asked Questions
  • Customer Service Portal
  • Privacy Policy
  • Terms Of Use
  • Single Issues For Purchase
  • International Print
Commercial Services
  • Advertising
  • Fortune Brand Studio
  • Fortune Analytics
  • Fortune Conferences
  • Business Development
  • Group Subscriptions
About Us
  • About Us
  • Press Center
  • Work At Fortune
  • Terms And Conditions
  • Site Map
  • About Us
  • Press Center
  • Work At Fortune
  • Terms And Conditions
  • Site Map
  • Facebook icon
  • Twitter icon
  • LinkedIn icon
  • Instagram icon
  • TikTok icon
  • YouTube icon

    Latest in AI


    Most Popular

    © 2026 Fortune Media IP Limited. All Rights Reserved. Use of this site constitutes acceptance of our Terms of Use and Privacy Policy | CA Notice at Collection and Privacy Notice | Do Not Sell/Share My Personal Information
    FORTUNE is a trademark of Fortune Media IP Limited, registered in the U.S. and other countries. FORTUNE may receive compensation for some links to products and services on this website. Offers may be subject to change without notice.