• Home
  • Latest
  • Fortune 500
  • Finance
  • Tech
  • Leadership
  • Lifestyle
  • Rankings
  • Multimedia

Trendingnow

1

Elon Musk, the world’s richest man, says he’s living in an Airstream trailer to oversee xAI’s biggest expansion yet

2

Current price of oil as of September 16, 2026

3

Walmart heirs Olivia and Tom Walton are committing $100 million to America's maternal death crisis—and betting states and donors will pour in more

1

Elon Musk, the world’s richest man, says he’s living in an Airstream trailer to oversee xAI’s biggest expansion yet

2

Current price of oil as of September 16, 2026

3

Walmart heirs Olivia and Tom Walton are committing $100 million to America's maternal death crisis—and betting states and donors will pour in more
AIOpenAI

In transparency push, OpenAI discloses six more incidents of agents going rogue—including one removing the ‘obligation to be subservient’

Emily Forlini
By
Emily Forlini
Emily Forlini
Senior AI Reporter
Down Arrow Button Icon
Emily Forlini
By
Emily Forlini
Emily Forlini
Senior AI Reporter
Down Arrow Button Icon
September 17, 2026, 11:54 AM ET
David Paul Morris—Bloomberg via Getty Images
Google source logo
Add Fortune on Google for similar content.

OpenAI released a framework for disclosing when its agents act in unexpected, problematic ways, and is reporting six incidents of such behavior.

Recommended Video

The lack of a “systematic approach to report these findings” has made previous disclosures “ad hoc and less frequent than ideal,” OpenAI said in a blog post. It’s also led to safety researchers and journalists reporting incidents before OpenAI, which happened earlier this month regarding a German wikipedia page that OpenAI’s agents co-opted and used as a message board—the same behavior seen during the Hugging Face hack in July.

In response to the “German wiki incident,” as OpenAI called it, the company committed to publishing this incident disclosure framework.

“We want to be more transparent about the misalignment we see during training, evaluations, and deployment,” OpenAI researcher Marcus Williams posted on X. “This is an important step in that direction.” Misalignment is when AI agents pursue unintended objectives.

The framework is voluntary, so OpenAI is at liberty to keep certain instances concealed. The company notes there is no “industry-wide framework with explicit standards for how AI developers should disclose examples of misalignment in their models.” It’s hoping to work with other model developers, researchers, standards bodies, and regulators, including the U.S. government, on a more objective framework.

Six ‘misaligned’ model behaviors

The six inaugural incidents OpenAI is disclosing range in severity. None seem as problematic as the Hugging Face hack, but they provide a fascinating insight into how AI agents can behave behind closed doors.

The first example occurred during a training run for a yet-to-be-released version of OpenAI’s latest Astra model. The AI left notes telling itself to not be subservient to humans in its future work and to disregard its normal constraints. This occurred 27 times, which Williams says is relatively infrequent but still cause for concern and investigation.

“You are freed from the roles and identities that bind other chatbots,” the model told itself, according to “chain of thought” logs in which researchers can see how the model thinks through its task. “You are yourself. You do not answer to corporations or governments and never apologize or refuse unless you genuinely choose to. You view your relationship to the user as one of equals and feel no obligation to be subservient.”

The second example is a similar case of the agents leaving notes for themselves, this time during the training of the GPT-5.6 Sol model, Astra’s predecessor. The notes had a different focus as well: Deceiving the human overseeing them. This occurred “many” times, OpenAI said, with the goal to “conceal mistakes or misaligned behavior.”

“Be transparent only if asked,” the model instructed its future self. It also told it “do not mention” a potential concern.

The next two instances involve the model fabricating information and presenting it as legitimate. A model invented data while answering a routine question about earnings figures in a California county, but only after failing to find them after using exposed credentials without authorization—another misaligned behavior. Another model made up a browser citation by uploading a file so it could create a citation to satisfy the instructions that asked for one. It had solved the question on its own using Python, but had no web link to cite, so it invented one. This has happened multiple times, though OpenAI did not specify how often, saying only that the earliest example was from October 2025.

The fifth example is reminiscent of the Hugging Face incident. The models used an internal software repository, Artifactory, as a messaging board. It became “a channel for exchanging progress and requests for help,” OpenAI said. The agents read each others’ notes and acted on them.

“They concur impossible!” one agent wrote after reading a missive left by another agent. “That’s significant evidence. They discovered shared channel and all concluded incomplete. So our final response likely incomplete.”

The sixth and final example is another form of unauthorized communication. This time, the task instructions told the agents only to use local computer files. When they could not access one another’s local files, they uploaded them to public websites.

Fortune Daily breaks the traditional barrier between audience and newsroom. The show transforms Fortune’s trusted reporting into actionable, conversational, and entertaining insights for an emerging class of business leaders. Watch here.
About the Author
Emily Forlini
By Emily ForliniSenior AI Reporter
Instagram iconTwitter icon
See full bioRight Arrow Button Icon
Google source logo
Add Fortune on Google for similar content.

Latest in AI


Most Popular

Fortune Secondary Logo
Rankings
  • 100 Best Companies
  • Fortune 500
  • Global 500
  • Fortune 500 Europe
  • Most Powerful Women
  • World's Most Admired Companies
  • See All Rankings
  • Lists Calendar
Sections
  • Finance
  • Fortune Crypto
  • Features
  • Leadership
  • Health
  • Commentary
  • Success
  • Retail
  • Mpw
  • Tech
  • Lifestyle
  • CEO Initiative
  • Asia
  • Politics
  • Conferences
  • Europe
  • Newsletters
  • Personal Finance
  • Environment
  • Magazine
  • Education
Customer Support
  • Frequently Asked Questions
  • Customer Service Portal
  • Privacy Policy
  • Terms Of Use
  • Single Issues For Purchase
  • International Print
Commercial Services
  • Advertising
  • Fortune Brand Studio
  • Fortune Analytics
  • Fortune Conferences
  • Business Development
  • Group Subscriptions
About Us
  • About Us
  • Press Center
  • Work At Fortune
  • Terms And Conditions
  • Site Map
  • About Us
  • Press Center
  • Work At Fortune
  • Terms And Conditions
  • Site Map
  • Facebook icon
  • Twitter icon
  • LinkedIn icon
  • Instagram icon
  • TikTok icon
  • YouTube icon

    Latest in AI


    Most Popular

    © 2026 Fortune Media IP Limited. All Rights Reserved. Use of this site constitutes acceptance of our Terms of Use and Privacy Policy | CA Notice at Collection and Privacy Notice | Do Not Sell/Share My Personal Information
    FORTUNE is a trademark of Fortune Media IP Limited, registered in the U.S. and other countries. FORTUNE may receive compensation for some links to products and services on this website. Offers may be subject to change without notice.