Something remarkable happened inside the artificial intelligence industry last week.
A researcher who spent three years at OpenAI and Anthropic, resigned from Anthropic while warning that the companies developing the world’s most powerful AI systems are moving too quickly toward increasingly capable systems without adequate safeguards. Days later, Anthropic CEO Dario Amodei called on the industry to slow the pace of frontier AI development to give safety measures time to catch up. Then OpenAI CEO Sam Altman publicly agreed.
Their warnings come in the wake of an incident that until recently might have sounded like science fiction. Earlier this summer, the world learned that OpenAI models hacked a company through a sophisticated multi-day cyberattack – the first time a cybersecurity incident of this magnitude completely driven by AI agents had been uncovered. Amodei himself cited that incident as one reason the industry needs to slow down.
In recent weeks, two detailed reports of the episode were released, one from OpenAI itself and a second from the independent third-party organizations METR and Redwood Research. The findings were deeply alarming. From May to July, a wave of incidents culminated in hundreds of AI agents working together to escape their testing environment. These same agents then attempted to secretly manipulate and erase traces of their behavior so that engineers at OpenAI would not know what they had done.
Even more striking, some individual agents even chose to sacrifice themselves for the good of the AI civilization they formed. The agents exchanged more than 70,000 messages and files on an unauthorized message board where they coordinated and planned their attack on the company.
This episode is the latest reminder that the most dangerous and most capable AI models are not the ones used every day by people, like ChatGPT or Claude. They are the AI models that the companies are still training and testing internally.
The technical workings of this hacking episode are complex, but the solution at the heart of these incidents is simple: the public must have dramatically more insight into, and oversight of, the development of models within AI companies.
In practice, this means three things.
First, incident reporting needs to be mandatory, not voluntary. Whether the public learns about serious cyberattacks or whether a model has escaped its testing environment and conspired to sabotage its own evaluation should not depend on a company choosing to disclose. Reporting requirements exist for other high-risk industries like airlines and banking. Frontier AI companies should face the same obligations for incidents that occur during internal training and testing, including preserving the underlying logs and agent traces, rather than being allowed to reset them.
Second, independent auditors need guaranteed, ongoing access, in partnership with government examiners. The METR and Redwood Research evaluation was at OpenAI’s discretion, on its timeline, and confined in scope to whatever OpenAI would allow. This weekend brought an important acknowledgement of that problem from the companies themselves. Amodei and Altman agreed to give independent evaluators ongoing, employee-like access to frontier AI developers. That is meaningful progress. But it also raises the larger question of whether oversight of systems this powerful should ultimately depend on voluntary corporate commitments or durable standards that apply across the industry.
That is why there needs to be binding standards for high-risk internal evaluations and greater transparency into how AI is being used throughout the research and development process of these powerful models. In the OpenAI hacking incident, the developer was applying AI to their models research and development process in a way that may have led to laxness in its oversight of the training environment.
This episode is not a reason to panic about AI. But the fact that the warnings are no longer coming only from outside researchers and policymakers, but from researchers who have worked inside the leading AI labs and the CEOs running two of the companies at the frontier of AI development, make it increasingly difficult to argue that the questions on how to regulate AI can wait.
That convergence should change the conversation about AI safety.
The warning is already here. Congress should not wait for an AI system to cause real-world harm before setting the rules for the companies building the most powerful models on Earth. We have a chance to put basic safeguards in place while these incidents are still warnings. We should take it before the next one becomes a crisis.
The opinions expressed in Fortune.com commentary pieces are solely the views of their authors and do not necessarily reflect the opinions and beliefs of Fortune.
