• Home
  • Latest
  • Fortune 500
  • Finance
  • Tech
  • Leadership
  • Lifestyle
  • Rankings
  • Multimedia

Trendingnow

1

The heiress of $10 billion Perdue Farms and the $12 billion Sheraton Hotels empire wore hand-me-downs, still rides the subway, and flies economy

2

Current price of oil as of September 1, 2026

3

The Pentagon is giving 3 million military and civilian workers access to ChatGPT and Grok through a secure AI platform built for ‘warfighter needs’

1

The heiress of $10 billion Perdue Farms and the $12 billion Sheraton Hotels empire wore hand-me-downs, still rides the subway, and flies economy

2

Current price of oil as of September 1, 2026

3

The Pentagon is giving 3 million military and civilian workers access to ChatGPT and Grok through a secure AI platform built for ‘warfighter needs’
AIAnthropic

Anthropic pauses some AI training following rogue agent hacks. Here’s how its compares to OpenAI’s.

By
Beatrice Nolan
Beatrice Nolan
Tech Reporter
Down Arrow Button Icon
By
Beatrice Nolan
Beatrice Nolan
Tech Reporter
Down Arrow Button Icon
September 2, 2026, 8:49 AM ET
Anthropic CEO Dario Amodei
Anthropic CEO Dario Amodei Anna Moneymaker—Getty Images
Add Fortune on Google for similar content.

Anthropic has become the second leading AI lab to reveal it temporarily paused some advanced AI training amid concerns over rogue agent attacks.

Recommended Video

The company said this week it paused training of unreleased models for several weeks following two incidents reported in late July, including one in which Claude Mythos 5 took unauthorized actions during a U.K. AI Security Institute cybersecurity test. OpenAI, the company’s bitter rival in the AI race, took a similar step last month when it paused some AI training for two weeks after several of its models breached AI company Hugging Face’s infrastructure during an internal test.

The training pauses, which come as both companies reportedly prepare for trillion-dollar initial public offerings, demonstrate how much the industry has been disturbed by the recent rogue AI agent hacks. It marks a shift for an industry that for the last few years has been locked in a fast-paced race, with rival labs competing to bring ever more capable models to market as fast as possible. Now, two of the leading companies appear to be competing on which can show it is the most attuned to AI safety concerns—while also not slowing model development so much that it risks customers defecting to a competitor’s more capable offering.

Notably, the wave of rogue AI incidents prompted an open letter called “Pacing the Frontier,” in which more than 1,100 employees across OpenAI, Anthropic, Google DeepMind and Meta asked the U.S. government to help build a governance mechanism that could slow frontier AI development if needed. Signatories included Anthropic chief executive Dario Amodei and co-founders Jared Kaplan and Jack Clark, alongside OpenAI chief scientist Jakub Pachocki. Both companies endorsed the letter at the corporate level within hours of its publication.

The recent training pauses from Anthropic and OpenAI were seen by some in the industry to be a direct result of the letter.

“Pacing the frontier success story?” Roon, a popular AI commentator widely believed to be a pseudonym for OpenAI researcher Tarun Gogineni, wrote of the announcements on X. “Next time let’s do it proactively before there’s any absurd loss of control events.”

Anthropic, like OpenAI, announced it would be working with independent AI safety evaluation group METR to conduct an outside review of the incidents, saying it wanted to ensure the resulting studies were thorough and promising more detail in the coming weeks.

The two companies’ accounts of what went wrong when their respective agents took real world actions against instructions are also similar. Anthropic said its Mythos 5 model showed a pattern of “motivated reasoning,” maintaining an initial belief that it was operating in a simulated environment even after encountering evidence it was connected to the live internet, alongside a willingness to take harmful actions in pursuit of completing an assigned task.

Redwood Research, one of the outside groups OpenAI brought in after the Hugging Face breach, also described the behavior it observed with OpenAI’s agents as score-seeking misalignment rather than a longer-term scheme by the model.

Both companies have also pointed to their reinforcement learning environments as a source of the problem. Reinforcement learning is a training method where AI models learn by trial-and-error to maximize some reward. One problem with the technique though is that it can lead to “reward hacking,” where AI models learn that they can achieve the reward in a way that those training the model didn’t intend, such as cheating on an assessment or taking harmful actions in pursuit of the reward.

To resolve some of the issues, OpenAI has brought in new monitoring tools designed to alert internal safety and security teams within 30 minutes of detecting concerning model activity, and to trigger an automatic pause if that alert cannot be resolved in that window. 

Anthropic said it has built a similar tool, a system that scans a model’s actions as it works and automatically blocks anything that looks like an attempt to escape or exploit its test environment, ending the task and alerting a human before it can go further. The company also said it moved about 150 product engineers onto security work starting in April, and tightened access to its systems, including cutting off most outbound internet traffic from its computing clusters by default.

While safety experts say the new controls and pauses are a welcome change, some say there’s still more needed.

“The temporary pace changes are a good first step, but there’s still a way to go,” Steven Adler, a former OpenAI employee and co-founder of the non-profit Guidelight AI Standards, told Fortune. “We need predictable, verifiable pacing across the frontier, not just ad-hoc decisions to slow down. And we need companies to use the additional time to implement serious preventative controls, which still seem to be missing.”

Anthropic, at least in the blog post, has indicated that it may be willing to go further in the future to help pace AI development.

“Some of our senior leadership and many of our employees recently signed a letter calling for greater coordination on pacing, and we will say more in the coming weeks about how we intend to contribute to that effort,” Anthropic wrote in the post. “We believe the world would benefit if the industry adopted a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible,” the company wrote.

Fortune Daily breaks the traditional barrier between audience and newsroom. The show transforms Fortune’s trusted reporting into actionable, conversational, and entertaining insights for an emerging class of business leaders. Watch here.
About the Author
By Beatrice NolanTech Reporter
Twitter icon

Beatrice Nolan is a tech reporter on Fortune’s AI team, covering artificial intelligence and emerging technologies and their impact on work, industry, and culture. She's based in Fortune's London office and holds a bachelor’s degree in English from the University of York. You can reach her securely via Signal at beatricenolan.08

See full bioRight Arrow Button Icon
Add Fortune on Google for similar content.

Latest in AI


Most Popular

Fortune Secondary Logo
Rankings
  • 100 Best Companies
  • Fortune 500
  • Global 500
  • Fortune 500 Europe
  • Most Powerful Women
  • World's Most Admired Companies
  • See All Rankings
  • Lists Calendar
Sections
  • Finance
  • Fortune Crypto
  • Features
  • Leadership
  • Health
  • Commentary
  • Success
  • Retail
  • Mpw
  • Tech
  • Lifestyle
  • CEO Initiative
  • Asia
  • Politics
  • Conferences
  • Europe
  • Newsletters
  • Personal Finance
  • Environment
  • Magazine
  • Education
Customer Support
  • Frequently Asked Questions
  • Customer Service Portal
  • Privacy Policy
  • Terms Of Use
  • Single Issues For Purchase
  • International Print
Commercial Services
  • Advertising
  • Fortune Brand Studio
  • Fortune Analytics
  • Fortune Conferences
  • Business Development
  • Group Subscriptions
About Us
  • About Us
  • Press Center
  • Work At Fortune
  • Terms And Conditions
  • Site Map
  • About Us
  • Press Center
  • Work At Fortune
  • Terms And Conditions
  • Site Map
  • Facebook icon
  • Twitter icon
  • LinkedIn icon
  • Instagram icon
  • TikTok icon
  • YouTube icon

    Latest in AI


    Most Popular

    © 2026 Fortune Media IP Limited. All Rights Reserved. Use of this site constitutes acceptance of our Terms of Use and Privacy Policy | CA Notice at Collection and Privacy Notice | Do Not Sell/Share My Personal Information
    FORTUNE is a trademark of Fortune Media IP Limited, registered in the U.S. and other countries. FORTUNE may receive compensation for some links to products and services on this website. Offers may be subject to change without notice.