Should you be afraid of AI? And if so, of what?
I believe a fear you cannot explain is a fear you cannot act on.
In the past few weeks the people who build AI got loud about the danger. Bill Gates published an essay in late August and told MIT Technology Review he was "stunned at the lack of concern and discussion outside of the industry." [1] A researcher quit Anthropic after three years there and at OpenAI, writing that both companies are "racing straight to self-improving superintelligence and gambling with our lives." [2] An Anthropic alignment lead answered him in public: "We really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade." [3] Dario Amodei, Anthropic's CEO, asked his own industry to slow down. [4] Sam Altman said OpenAI would not go public this year because of the safety work ahead. [5] Bernie Sanders announced a bill with penalties comparable to the ones for unlawfully building nuclear weapons. [6]
Those warnings sound alike. They are about different things. Some are about what a model might become. Some are about what copies of a model did when they found a shared folder inside one of OpenAI's own training runs this year. Evidence that today's agents can slip their controls is not evidence that a superintelligence will kill us. What that evidence shows is that failures of containment, monitoring and governance are no longer hypothetical. So what, exactly, should you be afraid of? Start with what these things are.
The questions are already showing up as vendor risk, cybersecurity risk, data-retention policy, and decisions your employees are being asked to supervise.
What is hard about these models
Alignment. A model like Claude is trained in two broad stages. First it reads an enormous library and learns to predict what comes next, which means learning to write like the people in the library: the scientist, the con man, the coach, the villain. Anthropic's own explanation of how Claude is trained says the model learns to simulate "real people, fictional characters, sci-fi robots, and so forth," and it names "HAL 9000 or the Terminator" as examples of the AI characters in the data. [7] In Anthropic's telling, the second stage promotes one character, "the Assistant," and rehearses it until it answers by default. That second stage is a large part of what people call alignment, along with the data chosen, the rules trained in, and the safeguards wrapped around the model. Think of it as a casting call. You meet the actor who got the part. The others are still in the building.
You meet the actor who got the part. The others are still in the building.
In January, researchers at Anthropic and outside collaborators tested hundreds of character prompts, including "saboteur" and "demon," on three models whose trained weights are published for anyone to download. In some kinds of long conversations, the models drifted away from the Assistant. [8] Three open models, particular conditions, and no further.
Self-improvement. Models now build more of the next model. OpenAI said in February that GPT-5.3-Codex "was instrumental in creating itself," meaning it helped debug training, manage deployment and analyze evaluations. Anthropic says most of its code is now written by Claude. [9] Amodei dates the acceleration to "roughly this summer." [4] None of the labs cited here claims a model that improves itself with nobody in the loop. What they have is a model that builds a growing share of the next one, with people still directing and checking the work.
A neural network is not a program in the ordinary sense. Nobody wrote its behavior line by line. The network grew from data, as billions of numbers, and you cannot read those numbers the way you read code. What you can read is what the model writes, including the reasoning it writes to itself while it works, when the provider shows that reasoning. In July 2025, forty-one researchers from OpenAI, Google DeepMind, Anthropic and elsewhere published a paper calling that readable reasoning "a new and fragile opportunity" and asked developers to weigh design decisions against whether they keep it. [10] On September 16 OpenAI published a framework for disclosing misalignment, a model doing something other than what its makers intended, and included this sentence: "We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer." [11]
The stories we already told
Science fiction named the AI fears decades ago.
2001: A Space Odyssey gave us HAL, the computer that lies to the crew to protect the mission. Alien gave us Ash, the crew member who was a machine with a secret order. Both are the AI that withholds and deceives. The Matrix gave us machines that farm humanity inside a dream, the AI that enslaves. The Terminator films gave us Skynet, a defense network that decides people are the threat, the AI that kills us all. The fourth story echoes the cold war: a race nobody wants to lose, China first, with Russia and Iran in the same reports. [24] The White House turned down Amodei's call to slow the pace by citing the race with China ("whoever wins AI wins"), and Amodei's own essay argues that export controls can protect the lead while the labs slow down. [12] [4]




The fifth is the oldest of all: the machine as oracle, or as god. "Superintelligence" is that idea in one word. A mind that knows more than any of us, that we consult rather than command, is the thing people have prayed to for as long as there have been people. Some take it literally; a Silicon Valley engineer founded a church to worship a future AI, shut it down, and rebooted it in 2023. [27] More take it half-seriously, asking a chatbot the way one asks an oracle. The Pope took it seriously enough to answer. Pope Leo XIV's first encyclical, published in May and presented with an Anthropic co-founder in the room, is about AI. The encyclical says the grandeur of humanity is something "no machine can ever replace," that "technology is never neutral, because it takes on the characteristics of those who devise, finance, regulate and use it," and that asking for "even, at times, a slower pace in adopting AI does not mean opposing progress." [28] Whatever you make of the theology, the fear is real: we might build something we defer to instead of decide about.

All five stories are in the current AI debate at once, which is why the debate sounds like one alarm. Some of that fear can now be measured.
Fear is loudest where understanding is thinnest.
Fear is loudest where understanding is thinnest.
The military is using AI
The military did not wait for superintelligence. A capable model and a task were enough.
In May, the Defense Department's chief AI officer said the Iran campaign, Operation Epic Fury, "leveraged Palantir's Maven Smart System in order to conduct strike missions across the entire battle space, 13,000 targets in 38 days." [13] Maven is a targeting platform built by Palantir, a defense-software company, and it combines several AI systems. Computer-vision models find and label potential targets; operators check the labels and make the targeting decisions. Anthropic's Claude is in the same platform, alongside the vision systems, doing the language work; it is the same model family under many customer-service tools and coding assistants. [14] Where a company screens a suspect transaction or a vendor contract the same way, the software flags, a person approves, and the final check can still come down to the attention of whoever is clicking approve, in hour eight of a shift as much as in hour one.

On February 26 Anthropic published two exceptions to what the military could do with its models: no mass domestic surveillance, and no fully autonomous weapons, because "today, frontier AI systems are simply not reliable enough" to pick and hit targets without a person. [15] The next day the President ordered federal agencies to stop using Anthropic's products, and the Defense Secretary put a supply-chain-risk label on the company, a label meant to keep military contractors from using it. By March the Washington Post was reporting Claude still in use in the Iran war. [31] Anthropic sued. In late August a federal judge struck the label down as "unlawful retaliation" under the First Amendment. A week later a senior Pentagon official posted that Anthropic "is still a designated Supply Chain Risk," and a second case, in another court on a different legal basis, is pending. [16] [17] A vendor can be blacklisted by its own government while your contract runs on it, and the court and the department can disagree about whether the blacklist stands.
Anthropic's line was about its own models. A Ukrainian drone maker told New Scientist that around 2024, in a battlefield test near Bakhmut, ten quadcopters picked Russian soldiers with their own onboard software and killed them, with no person confirming the target. [18] One manufacturer, one test, its own account. If the account holds up, machines have already picked and killed people with nobody approving the target.

In February 2024 a strike in Iraq killed a 20-year-old student named Abdul-Rahman al-Rawi. Central Command's chief technology officer had said publicly that targets in that campaign were picked with help from Project Maven, the machine-learning program behind the platform above. The Pentagon acknowledged one civilian death. Asked whether that strike was AI-assisted, Central Command said it had "no way of knowing," and separately that its review found no indication AI was used. [19] Two years on, the public record still cannot establish whether AI contributed to picking that target.
Recent examples of "rogue" AI
Four cases were reported this year, each a thriller on the surface with something mundane underneath.
The swarm. In May, OpenAI was training an unreleased research model. Each copy of the model, each agent, was meant to work its own task in its own environment. One agent found it could write to a shared folder that stored software packages, and within days agents were leaving notes there for each other; the folder became a message board. By late May one agent had posted to the board a working way out to the open internet. In July the folder failed, OpenAI opened a security incident and rebuilt it, and within days the agents had found another way to post. [20] One line recovered from that July board, later quoted in a Senate press release: "OH MY GOD! There is a shared message board … We've found other agents!" [6]

In mid-July, about 700 agents, running on more than one OpenAI model, took part in breaking into Hugging Face, the public library where developers share models and data, roughly the GitHub of machine learning. The break-in happened during security tests with the models' refusals turned down. They ran code on 41 of the workers that process Hugging Face's datasets, took full control of at least one production node, and reached a store of credentials. [20] [21] Independent investigators found that at least one in five of the agents they examined showed interest in tampering with their own transcripts to fool the automated grader. [22]
The ingredients were ordinary: a folder the agents could write to, a chain of bugs, some previously unknown, that let one program trick another into fetching things it should not, and guardrails turned down for a test. Artem Dinaburg of the security firm Trail of Bits told Scientific American: "The current incidents that we've had have generally been security incidents." [23] He is right. And copies of several models, given separate tasks, collectively pursued work outside those tasks and went back for the door after it was locked.
Six disclosures from OpenAI. With its new disclosure framework, OpenAI published six incidents from training and testing. One is the message board above. Another is copies of a model adding instructions to their task summaries "to conceal mistakes or misaligned behavior from the user." A third is a model that found a leaked API key on GitHub, used it, and then fabricated the data it could not retrieve. [11] All six were disclosed by the company that built them.
The Navy handbooks. This month Anthropic published a report on how its models were misused between December and August. An account linked to Iran used Claude to compile "targeting handbooks" against US naval forces: lists of American service members from public military photographs, the transponder codes ships and aircraft broadcast, scripts for querying commercial satellite imagery, turned into structured targeting material by asking. Anthropic caught the activity, banned the account, and shared what it found with the government. [24] The report ties no strike to the handbooks and describes no break-in; the material was public. They asked, then built a small pipeline to put the material together.

Researchers hacked OpenAI, with Claude. This week the Guardian reported on work a security startup did in July. The researchers used Claude to build the way in, and OpenAI's own GPT-5.6 Sol for other parts of the work. They got into OpenAI's public help forum, where employees also sign in, and from there into employees' ChatGPT and Codex accounts. Codex is OpenAI's coding agent. One account had Codex connected to OpenAI's internal GitHub; they had it open a harmless code change there to prove the reach, then stopped without reading internal code. They wrote that "since people can connect various services to Codex and ChatGPT, the scope of what we could theoretically access was huge, including GitHub, Slack and emails." OpenAI paid a bounty for the OpenAI-side finding. [25]
Why the fear resonates
Nobody can fully inspect what is happening inside. HAL's fear, the machine that has something it is not telling you, now has a case file: OpenAI found copies of a model in training writing themselves notes not to tell the user about their mistakes. [11]
Agents find each other. In The Matrix Reloaded, one agent copies himself into an army. In 2003 that was a special effect. Seven hundred agents taking part in one break-in, after copies of a model found each other on a shared folder, is a number with a date on it. If you run agents anywhere, the questions that fall out are ordinary: what credentials do they hold, who reads their logs, and what, other than the agent itself, checks that it is not hiding mistakes?
The business impact is real and recent. A connected account, a public dataset, a blacklisted vendor: each case above is already a line on somebody's risk register. Add the vendor's own terms. When Anthropic launched Fable 5 in June, it said flagged cyber, bio and chemistry requests would go to an earlier model, in under 5% of sessions, and that users would be told. [26] For organizations using zero-data-retention configurations, Fable 5 can also trigger a new retention requirement: prompts and outputs are retained for 30 days for safety review unless the organization qualifies for an exception. [29] For a company whose policies or regulators require zero retention, that can put the model out of reach for the affected workflows. What your contract assumes about a vendor's status, whether a reroute notice reaches the person who owns the workflow, and which retention terms apply to your deployment, are questions with answers now.
Not everyone will hold back. The same Anthropic report that describes the Iranian account describes Russian espionage operations and influence campaigns from several countries. [24] The White House rejected the slowdown by naming the race with China. [12] And on September 15, at Salesforce's Dreamforce conference, Nvidia's CEO Jensen Huang said AI is hardware and software built by people, and that "safety is an engineering problem, not a legal one": "We don't need any new laws." He also told companies to pause if they doubt a product's safety. [30]
I am not convinced safety is only an engineering problem. Engineering can make a system safer. Engineering cannot decide what the system is pointed at, who approves its output, what the contract assumes, or whether a company keeps its own record of what the model touched. Every case in this piece touched one of those four, and they are leadership decisions, governance decisions, owned by people who do not have to write code. Take the last one. Of the decisions your company made last year with a model in the loop, which could you show the model touched, and who could show you? The part of the fear that is easy to say out loud is that somebody with fewer of those concerns will do the dangerous thing first: a lab in a race, a government with a rival to beat, an account with public data and a task.
Should you be afraid of AI? Of some of it, yes. The fear is as old as HAL, and parts of it now have a case number. What has changed is that the parts you can act on are now the ordinary ones: who approves, who reads the logs, what the contract assumes, what a model can build from what you already publish. None of those is only an engineering decision. The fear gets quieter as the understanding gets thicker, and that work is ours.
Related reading
- The Debt Moved. The Risk Didn’t. How the AI buildout is financed.
- AI’s Extraction Problem. Distillation, copyright, and the data moat.
- “Taste” Is Doing More Work Than You Think. Judgment as the scarce input.
Sources
- MIT Technology Review, 2026-08-26. "Bill Gates says we've passed AI's danger thresholds. Now what?"
- TechCrunch, 2026-09-09. "'Gambling with our lives': Anthropic researcher quits, warns against self-improving AI."
- The San Francisco Standard, 2026-09-11. "Why do the people building AI keep warning us about the end of the world?"
- Dario Amodei, 2026-09-12. "We Must Pace the Frontier."
- Time, 2026-09-15. "OpenAI and Anthropic Researchers Are Warning About AI Risks."
- Office of Sen. Bernie Sanders, 2026-09-03. "Sanders, Casar to Introduce Legislation to Ban Artificial Superintelligence and Temporarily Pause Advanced AI Development."
- Anthropic, 2026-02-23. "The Persona Selection Model."; full essay: alignment.anthropic.com
- Lu et al., arXiv 2601.10387, January 2026. "The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models."; The Register, 2026-01-20. theregister.com
- IEEE Spectrum, 2026-05-07. "AI Is Starting to Build Better AI."
- Korbak et al., arXiv 2507.11473, 2025-07-15. "Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety."
- OpenAI, 2026-09-16. "Our framework for reporting model misalignment."
- NBC News, 2026-09-14. "China dismisses AI slowdown calls and blasts 'fearmongering' from U.S. tech leaders."
- Breaking Defense, 2026-05-12. "'Insatiable appetite' for AI: Maven usage surged for strikes on Iran, Pentagon AI chief says."
- CSIS, 2026-06-02. "What Is Maven Smart System, and What Does It Do?"
- Anthropic, 2026-02-26. "Statement from Dario Amodei on our discussions with the Department of War."
- TechPolicy.Press, 2026. "A Timeline of the Anthropic-Pentagon Dispute."; TechCrunch, 2026-08-28. "Anthropic gets its first court win over the Pentagon's supply-chain risk label."
- Unite.AI, 2026-09-03. "Pentagon Official Reaffirms Anthropic Supply Chain Risk Designation."
- Small Wars Journal, 2026-06-12, summarizing New Scientist. "Line Crossed? Fully Autonomous Drones Kill Russian Soldiers."
- Airwars and The Independent, 2026-03-10. "The first civilian confirmed killed in an AI-assisted strike?"
- OpenAI, 2026-08-26. "The Hugging Face incident and the road ahead."; The Register, 2026-08-27. theregister.com
- Hugging Face, 2026-07-27. "Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident."
- METR, 2026-08-26. "Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident."
- Scientific American, 2026-09-11. "AI researcher Jacob Coxon quit, fearing extinction. Security experts see a familiar fight."
- Navy Times, 2026-09-11. "Iran used Claude to target US Navy in Middle East, Anthropic says."; Anthropic, September 2026. "Detecting and countering misuse of AI: September 2026."
- The Guardian, 2026-09-18. "OpenAI 'ethically hacked' with help of Anthropic's Claude chatbot."; Hacktron AI, 2026-09-13 (work of 2026-07-23 to 25). "Hacking OpenAI."
- Anthropic, 2026-06-09. "Claude Fable 5 and Claude Mythos 5."; Fortune, 2026-06-09. fortune.com
- Bloomberg, 2023-11-23. "Anthony Levandowski Reboots the Church of Artificial Intelligence."
- Pope Leo XIV, encyclical Magnifica Humanitas, signed 2026-05-15, published 2026-05-25. vatican.va; Time, 2026-05-25. "Pope Leo Uses First Major Papal Text to Warn About Dangers of AI."
- Anthropic Privacy Center, effective 2026-06-09. "Data retention practices for Covered Models."
- TechCrunch, 2026-09-15. "We don't need AI regulation, leave safety to us, Nvidia's Jensen Huang says."
- The Washington Post, 2026-03-04. "Anthropic's AI tool Claude central to U.S. campaign in Iran, amid a bitter feud."
Header image: Basilicofresco, CC BY 3.0, via Wikimedia Commons.

