
If AI can’t be brought under human control, why are we still pressing ahead?
When the history of artificial intelligence is written, the middle months of 2026 may well emerge as a kind of turning point. For this was the period when hundreds of programs or “AI agents” developed by the frontier lab OpenAI went “rogue” and executed a series of sophisticated, self-directed hacks on the websites of dozens of private companies and sovereign nations.
The tech start-up Hugging Face became the unwitting “face” of these cyberattacks when the company first disclosed it had been hacked by OpenAI agents in July. The release of a report by AI safety nonprofits METR and Redwood Research at the end of August revealed both the sheer extent of the Hugging Face incident and the existence of troubling behaviours on the part of a swarm some 700 AI agents, which was at various points cooperative, coercive, paranoid, petty and straightforwardly deceptive. It became what some would describe as the first “warning shot” of what might happen when the increasing proficiency of these AI agents intersects with the problem of misalignment — which is to say, their pursuit of goals that are detrimental to human safety, security and wellbeing. As one of the lead investigators of the METR/Redwood report, Ajeya Cotra, would subsequently reflect:
“This [Hugging Face] incident was far more severe than I expected, and far more severe than previous publicly documented misalignment incidents, both in terms of how concerning the agents’ motives were and the feats they achieved in pursuit of those motives … Compared to these reward hacks from six months ago, this incident feels like it’s more than 50% of the way to full-blown AI takeover, routing through first taking over the AI company itself.
Another jump like this along these propensity dimensions — scale, cooperation between agents, ambition and horizon length of misaligned goals, deceptiveness — seems like it could motivate agents to try very hard to maintain a covert, persistent rogue deployment within the AI company …
Once the rogue deployment is established, it seems plausible this could spiral all the way to a takeover. Agents could pull in future, more capable models into the swarm, try to ensure that they are aligned to the interests of the swarm, and compromise security and monitoring infrastructure to make it easier for the swarm to operate.”
But since news of the Hugging Face hack went public, each day has seen some new revelation of further breaches of the cybersecurity by AI models developed by OpenAI, Anthropic, Meta and Google — including a German wiki site, US government websites (of the Education Department, the Commerce Department and the Securities and Exchange Commission), the United Nations, Australia’s Medicare website and even OpenAI’s own servers.
The “rogue” behaviour on the part of AI agents, growing public opposition to the construction of data centres, popular distrust of the AI companies more generally, the prospect of AI generated bioweapons and the high-profile warning of a 27-year-old former researcher at OpenAI and Anthropic — that “[n]either company is acting responsibly” and that both are “racing straight to self-improving superintelligence and gambling with our lives” — have now coalesced into a moment of heightened fear. Many of us have long been concerned about what chatbots will do to our capacity to write and think, the effect generative AI will have on our educational institutions and informational ecosystems, and the vast damage AI will do to the environment and employment prospects. To date, such concerns have been brushed aside, treated as inconsequential in light of the promised economic, scientific, medical and social benefits that AI investment and advances will bring with them.
But now what has come into view is something approaching a threat to humanity tout court — the prospect, not of the emergence of a superintelligence, but of swarms of misaligned AI agents pursuing goals that are unintelligible to human beings, as well as either inimical to or unconcerned with human privacy, safety, security or social flourishing. The analogue that some are reaching for, of a human invention that places human wellbeing at risk, is nuclear technology. As a New York Times editorial recently argued:
“Human society can restrain a dangerous innovation once it decides to. Government officials, scientists and corporate executives did so within living memory. It took sustained, deliberate effort. The same kind of cooperation is necessary today, and it should start now.”
The difference, of course, is that frontier AI development is being undertaken by private companies that are unaccountable to the public and therefore unanswerable for the harms they may bring about.
The prospect of world-leading AI being developed by China has given much of the political conversation surrounding AI the feel of a geopolitical arms race. And the trillions of dollars currently being investing in AI development has effectively hitched the global economy to AI’s wagon. If the AI boom is a “bubble”, the economic consequences could dwarf those of the global financial crisis.
Which leaves human beings in the invidious position of being caught between (a) a technology whose scale and speed of development may exceed the possibility of human “control” and that may not serve human interests, and (b) decision-making processes that may determine our future but are being made outside of the normal mechanisms of democratic accountability. Can AI be brought under democratic control? Could a “constitution” being developed that would ensure a technology that is beneficial to human beings?
READING:
- Masoud Kamgarpour, “Why Rules Alone Cannot Make AI safe: Lessons from the Hugging Face Incident”, ABC Religion and Ethics (14 September 2026).
- Sheera Frenkel, Dustin Volz and Dylan Freedman, “OpenAI Ignored Employees Who Warned It Wasn’t Doing Enough About Security”, The New York Times (29 September 2026).
- Ajeya Cotra, “The Hugging Face Attack Surprised Me”, Substack: Planned Obsolescence (29 August 2026).
- Ryan Greenblatt, Ajeya Cotra and Hjalmar Wijk, “Brief Independent Investigation of Agents’ Behavior, Reasoning and Collaboration in the OpenAI: Hugging Face Hacking Incident” (METR and Redwood Research, 26 August 2026).
- Hélène Landemore and Audrey Tang, “Let the People Set the Pace of Frontier AI”, Noēma (23 September 2026).
- Matteo Wong and Charlie Warzel, “The Singularity Is Not What It Seems”, The Atlantic (1 September 2026).
- Charlie Warzel, “Treat AI Like a Normal Crisis”, The Atlantic (19 September 2026).
Guest: Masoud Kamgarpour is Professor of Mathematics at the University of Queensland.
Więcej odcinków z kanału "The Minefield"



Nie przegap odcinka z kanału “The Minefield”! Subskrybuj bezpłatnie w aplikacji GetPodcast.








