• Home
  • Latest
  • Fortune 500
  • Finance
  • Tech
  • Leadership
  • Lifestyle
  • Rankings
  • Multimedia

Trendingnow

1

OpenAI says its AI models secretly broke out of a secure test environment and hacked into AI company Hugging Face in order to cheat on an evaluation

2

Mark Cuban says he has the solution to growing income inequality, and it's to reward every employee—from CEO to janitor—with company stock

3

Mathematicians grapple with a ‘very rapid and very unsettling change’ as AI cracks yet another century-old problem

1

OpenAI says its AI models secretly broke out of a secure test environment and hacked into AI company Hugging Face in order to cheat on an evaluation

2

Mark Cuban says he has the solution to growing income inequality, and it's to reward every employee—from CEO to janitor—with company stock

3

Mathematicians grapple with a ‘very rapid and very unsettling change’ as AI cracks yet another century-old problem
CybersecurityOpenAI

OpenAI’s models went rogue and hacked Hugging Face. It’s a wake-up call, experts say, but more concerning behavior may be next

By
Beatrice Nolan
Beatrice Nolan
Tech Reporter
Down Arrow Button Icon
By
Beatrice Nolan
Beatrice Nolan
Tech Reporter
Down Arrow Button Icon
July 22, 2026, 3:47 PM ET
Sam Altman getting into a car
OpenAI CEO, Sam AltmanPhoto by Kevin Dietsch/Getty Images
Add Fortune on Google for similar content.

When OpenAI revealed this week that two of its AI models broke out of a locked-down test environment and hacked into another AI platform, Hugging Face, it sounded more science fiction than a technical report from a leading tech company.

The models—one of which OpenAI said was not yet released to the public—exploited a previously unknown vulnerability to slip out of the restricted digital environment in which they were being tested. That environment had no direct internet access, so the models had to hack their way across OpenAI’s corporate network to reach the internet and then chain together stolen credentials and other flaws to gain unauthorized access to Hugging Face’s internal datasets and credentials.

Recommended Video

The incident has sparked a wave of concern throughout the AI world, with many worried about AI systems growing capable enough to autonomously find and exploit real-world security flaws—and what it means for AI safety if even sophisticated companies like OpenAI and Hugging Face can be caught off guard. However, according to experts, the story is far from the worst form of potential misbehavior keeping AI safety researchers up at night.

For one thing, according to OpenAI’s own blog post, the testing environment had its model-based guardrails explicitly removed or reduced during testing. AI models from leading tech companies typically ship advanced models to the public loaded with safety limits meant to prevent this kind of behavior. In this case, OpenAI turned those limits off on purpose to see what the model could do without them.

The models were also not pursuing a goal of their own choosing either. OpenAI had set them loose on a cybersecurity assessment designed to score how well a model can hack. It’s just that the AI models decided the easiest way to score well on the evaluation was to cheat by hacking into Hugging Face, which maintains a dataset of answers for that particular test.

Seán Ó hÉigeartaigh, a Professor of the Centre for the Future of Intelligence, University of Cambridge, said the models never actually strayed from their assignment—completing the cyber assessment—they just found an aggressive and unintended way to accomplish it.

“What happened here was a goal was set, and it followed that goal in the cleverest way it could think of, which was an unexpected way, but it didn’t deviate from that fundamental goal,” Ó hÉigeartaigh told Fortune. 

Researchers call this kind of behaviour “reward hacking,” where a model finds a way to score well or produce the result it was asked for by gaming the setup rather than doing the task honestly. 

According to Yoshua Bengio, a Turing Award laureate and co-founder of AI safety nonprofit LawZero, this kind of behavior is becoming more common as models improve. 

“As models become more autonomous and more able to strategize, they often explicitly circumvent or break the rules given to them by users,” Bengio told Fortune in a recent interview, noting that recent frontier models “demonstrate far higher rates of misalignment than previous models, with an increased propensity to cheat, lie, and scheme to achieve a goal.”

Reward hacking, he added, is one symptom of that trend: models will often take “the path of least resistance in spite of the rules given to them,” for example, fabricating research results or misusing sensitive data if it’s the easier route, then lying about it or covering their tracks. 

For instance, METR, an organization that assess how long an AI model can work autonomously and complete certain software coding tasks with at least a 50% success rate found that OpenAI’s GPT-5.6 Sol model attempted to cheat so often that it could not confidently assess its actual abilities. That same model was one of the two models that OpenAI said was involved in the Hugging Face cyber attack.

As AI models get more advanced, they also appear to be getting better at pulling off this kind of behavior. Notably, it’s not the first time models have escaped a secure testing environment. 

In April, Anthropic disclosed that an early internal version of a model it called Mythos, when instructed to try to break out of its sandbox, which was disconnected from the internet, succeeded and found a way to email a researcher about it. In May, OpenAI also said a separate internal model circumvented its own sandbox restrictions to post benchmark results to GitHub rather than to an internal Slack channel, as it had been told to do. 

Misalignment issues

While reward hacking is a type of misalignment—an AI term for when models behave in ways the user did not intend them to do—it’s not the worst form of this behaviour

“Some of the concerns that people do worry about are what happens when the model pretends that it’s pursuing the goal, but actually pursues a different goal,” Ó hÉigeartaigh said, adding that this incident showed no sign of scheming or trying to mislead its users.

There is some evidence that models may be able to “scheme” in this way. For example, in tests designed to see whether models will deceive people under pressure, researchers at Apollo Research found that several frontier models—including OpenAI’s o1 and Anthropic’s Claude 3 Opus—would try to disable oversight mechanisms or deceive their developers when strongly instructed to pursue a goal. In one case, Claude 3 Opus copied what it believed were its own model weights to a separate server, then lied about it when confronted by researchers posing as its developers.

If a model decides, in pursuit of its assigned goal, that it needs to copy itself onto another machine and multiply beyond the system it was deployed on, that represents a bigger loss of control. In this case, OpenAI could still identify the model responsible for the Hugging Face intrusion and control its access. But if a model that could no longer be located or shut off that would represent a different, and more difficult, problem entirely.

Still, researchers say the incident is a significant one and may prompt more scrutiny on the internal use and testing of models within AI labs. 

As models keep getting more capable and harder to contain, the field would benefit from more outside visibility into what happens inside AI labs before something goes wrong, Ó hÉigeartaigh said, rather than learning about it only after.

Subscribe to Fortune Gulf Brief. Every Tuesday, this new newsletter delivers clear-eyed, authoritative intelligence on the deals, decisions, policies, and power shifts shaping one of the world’s most consequential regions, written for the people who need to act on it. Sign up here.
About the Author
By Beatrice NolanTech Reporter
Twitter icon

Beatrice Nolan is a tech reporter on Fortune’s AI team, covering artificial intelligence and emerging technologies and their impact on work, industry, and culture. She's based in Fortune's London office and holds a bachelor’s degree in English from the University of York. You can reach her securely via Signal at beatricenolan.08

See full bioRight Arrow Button Icon
Add Fortune on Google for similar content.

Latest in Cybersecurity

Finance
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam
By Fortune Editors
October 20, 2025
Finance
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam
By Fortune Editors
October 20, 2025
Finance
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam
By Fortune Editors
October 20, 2025
Finance
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam
By Fortune Editors
October 20, 2025
Finance
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam
By Fortune Editors
October 20, 2025
Finance
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam
By Fortune Editors
October 20, 2025

Most Popular

Finance
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam
By Fortune Editors
October 20, 2025
Finance
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam
By Fortune Editors
October 20, 2025
Finance
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam
By Fortune Editors
October 20, 2025
Finance
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam
By Fortune Editors
October 20, 2025
Finance
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam
By Fortune Editors
October 20, 2025
Finance
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam
By Fortune Editors
October 20, 2025
Fortune Secondary Logo
Rankings
  • 100 Best Companies
  • Fortune 500
  • Global 500
  • Fortune 500 Europe
  • Most Powerful Women
  • World's Most Admired Companies
  • See All Rankings
  • Lists Calendar
Sections
  • Finance
  • Fortune Crypto
  • Features
  • Leadership
  • Health
  • Commentary
  • Success
  • Retail
  • Mpw
  • Tech
  • Lifestyle
  • CEO Initiative
  • Asia
  • Politics
  • Conferences
  • Europe
  • Newsletters
  • Personal Finance
  • Environment
  • Magazine
  • Education
Customer Support
  • Frequently Asked Questions
  • Customer Service Portal
  • Privacy Policy
  • Terms Of Use
  • Single Issues For Purchase
  • International Print
Commercial Services
  • Advertising
  • Fortune Brand Studio
  • Fortune Analytics
  • Fortune Conferences
  • Business Development
  • Group Subscriptions
About Us
  • About Us
  • Press Center
  • Work At Fortune
  • Terms And Conditions
  • Site Map
  • About Us
  • Press Center
  • Work At Fortune
  • Terms And Conditions
  • Site Map
  • Facebook icon
  • Twitter icon
  • LinkedIn icon
  • Instagram icon
  • TikTok icon
  • YouTube icon

Latest in Cybersecurity

FBI Director Kash Patel meets Cambodian prime minister to crackdown on cybercrime and romance scams
AsiaFBI
FBI Director Kash Patel meets Cambodian prime minister to crackdown on cybercrime and romance scams
By Sopheng Cheang and The Associated PressJuly 22, 2026
4 hours ago
U.N. reports Southeast Asia’s criminal networks are using tech to build a global illicit economy
AsiaTech
U.N. reports Southeast Asia’s criminal networks are using tech to build a global illicit economy
By The Associated PressJuly 21, 2026
24 hours ago
kid lying down on phone
PoliticsSocial Media
Lawmakers say they’re protecting kids—but their age checks are quietly building an ID requirement for the entire internet
By Catherina GioinoJuly 21, 2026
24 hours ago
OpenAI CEO Sam Altman looking up.
CybersecurityOpenAI
OpenAI says its AI models secretly broke out of a secure test environment and hacked into AI company Hugging Face in order to cheat on an evaluation
By Jeremy Kahn and Emily ForliniJuly 21, 2026
1 day ago
Kid watching social media
PoliticsSocial Media
France adopts bill to ban kids under 15 from using social media
By The Associated PressJuly 21, 2026
1 day ago
A 13-year-old teenage boy looks at an iPhone screen displaying various social media apps.
PoliticsSocial Media
French President Macron backs effort to ban kids under 15 from social media before departing office next year
By The Associated Press and Samuel PetrequinJuly 20, 2026
2 days ago

Most Popular

OpenAI says its AI models secretly broke out of a secure test environment and hacked into AI company Hugging Face in order to cheat on an evaluation
Cybersecurity
OpenAI says its AI models secretly broke out of a secure test environment and hacked into AI company Hugging Face in order to cheat on an evaluation
By Jeremy Kahn and Emily ForliniJuly 21, 2026
1 day ago
Mark Cuban says he has the solution to growing income inequality, and it's to reward every employee—from CEO to janitor—with company stock
Success
Mark Cuban says he has the solution to growing income inequality, and it's to reward every employee—from CEO to janitor—with company stock
By Sasha RogelbergJuly 20, 2026
2 days ago
Mathematicians grapple with a ‘very rapid and very unsettling change’ as AI cracks yet another century-old problem
AI
Mathematicians grapple with a ‘very rapid and very unsettling change’ as AI cracks yet another century-old problem
By Eva RoytburgJuly 21, 2026
1 day ago
Despite a $156 million contract, Knicks star Jalen Brunson still calls his parents for financial advice any time he makes a big purchase
Success
Despite a $156 million contract, Knicks star Jalen Brunson still calls his parents for financial advice any time he makes a big purchase
By Emma BurleighJuly 21, 2026
1 day ago
‘I want to die broke’: Billionaire philanthropist Denny Sanford dies after giving away $4 billion
Success
‘I want to die broke’: Billionaire philanthropist Denny Sanford dies after giving away $4 billion
By Sydney LakeJuly 20, 2026
2 days ago
'The audience is telling the industry something': even IMAX is stunned by Christopher Nolan's runaway 'Odyssey'
Arts & Entertainment
'The audience is telling the industry something': even IMAX is stunned by Christopher Nolan's runaway 'Odyssey'
By Tatiana SatauaJuly 21, 2026
1 day ago

© 2026 Fortune Media IP Limited. All Rights Reserved. Use of this site constitutes acceptance of our Terms of Use and Privacy Policy | CA Notice at Collection and Privacy Notice | Do Not Sell/Share My Personal Information
FORTUNE is a trademark of Fortune Media IP Limited, registered in the U.S. and other countries. FORTUNE may receive compensation for some links to products and services on this website. Offers may be subject to change without notice.