Connect with us

NEWS

Anthropic’s Own File Shows Claude Already Serving Spies

Anthropic’s September threat file lists spies, a 1,000-account Malaysia op, and virus work on Claude, as a petition still hunts an AI off-switch.

Published

on

Anthropic published a 154-page threat file on September 10 that shows its Claude models already serving spies, influence shops, and a military-linked virus lab. The same week, the company’s alignment science lead put a greater than 10% chance on AI killing everyone this decade, and a petition still hunting an off-switch listed 925,471 names.

The safety brand and the case file now sit in the same news cycle. One warns about a future superintelligence. The other lists work that already ran on today’s models.

Claude Already Worked for Spies and Influence Shops

The paper is titled Detecting and Countering Misuse of AI: September 2026. It covers activity Anthropic says it disrupted between December 2025 and August 2026 across seven harm areas: cyber operations, influence operations, surveillance, scams and fraud, biological misuse, conventional weapons, and illicit distillation. The disrupted Claude misuse case studies use Claude Haiku, Sonnet, and Opus. None of the published misuse, Anthropic said, ran on Claude Fable or Mythos-class models except one distillation case.

The company’s own headline finding is blunt. “Sophisticated attacks no longer require sophisticated attackers,” the report says. AI closed the labor gap that used to separate a state team from a small crew. Humans still picked targets and reviewed stolen data. Multi-agent setups did the reconnaissance, the exploit work, and the exfiltration.

Anthropic said it disrupted every operation in the file, tightened safeguards from what it learned, and shared intelligence with authorities and other labs where it judged that useful. The actors it names are not a handful of billionaires tinkering in public. They are suspected state groups, paid criminals, commercial spyware vendors, state propaganda desks, and politically motivated individuals.

SEVEN HARM AREAS IN THE SEPTEMBER FILE

Harm area What Anthropic published Concrete example
Cyber operations AI as orchestrator, not chatbot GTG-20006 hit more than 20 organizations
Influence operations State desks and paid shops About 1,000 fake accounts aimed at Malaysia
Surveillance Tools to watch people at scale Dissident and activist monitoring in the case mix
Conventional weapons Six cases: three in China, two in Russia, one in Yemen An FPV drone swarm with no human in the loop
Biological misuse Five case studies of working scientists Gain-of-function work on chikungunya
Scams and fraud Industrial social engineering A network of fake dating apps in the opening cases
Illicit distillation Copying Claude without a license The one published exception involving higher-class models

GTG-20006 is the cyber spine. Anthropic’s attribution is consistent with public reporting on Midnight Blizzard, the SVR-linked group. One operator used the handle JackPoterz. Custom AI workflows rebuilt malware when security products caught it, then staged the new build from disposable servers. The actor scanned more than two dozen Ukrainian government organizations, bulk-exported mailboxes from at least two drone-component makers, and stole a full software kit for a drone vision system. Hotel Wi-Fi vendors were hijacked so guest traffic could be redirected. WhatsApp accounts were taken over with headless browsers so conversations could be copied quietly.

On the weapons side, Anthropic lists six disrupted cases. Four used Claude to write weapons software, including a guided-rocket program that ran a live field test and a Russia-based freelance team that built an FPV kamikaze swarm it called DronDoc or Serafim. That swarm’s onboard model could pick a “person” target class and order detonation without a human in the loop. The team trained a vision classifier on scraped Ukrainian combat footage and used a fixed point in Donetsk Oblast as a demo strike. Anthropic said the crew looked like a small freelance mix of civilian and military work, not a Russian state unit, and that it banned the linked accounts.

Hubinger’s Above-10% Call, With No Alignment Plan

On September 9, a day before the file went out, Evan Hubinger, Anthropic’s alignment science lead, answered a resignation thread from Jacob Coxon, a pretraining researcher who had just left the company after three years at OpenAI and Anthropic. Coxon wrote that the people building AI “earnestly believe that it could kill us all by the end of the decade,” and that neither firm was acting responsibly.

Jacob is correct here-we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.

Evan Hubinger, Alignment Science lead at Anthropic, on X

Hubinger still works at the company. In a follow-up the same morning he narrowed the fear. Present models, he said, look low-risk in Anthropic’s latest Risk Report. What worries him is superintelligence from recursive self-improvement, “happening faster than we thought.”

Coxon had put the same private fear on the record in public language. “This is not a marketing stunt,” he wrote. “If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible, but I hear the same people express fear privately.” Illinois Gov. JB Pritzker quoted the resignation and called for “immediate action from the industry and Washington.”

THE WEEK THE FILE AND THE ODDS LANDED

  1. March 4, 2026: The Future of Life Institute and a cross-spectrum coalition release the Pro-Human AI Declaration in New Orleans.
  2. September 9, 2026: Coxon resigns and says both labs are racing to self-improving superintelligence.
  3. September 9, 2026: Hubinger endorses the core claim and publishes his personal >10% estimate.
  4. September 10, 2026: Anthropic publishes the threat file and says every operation in it was disrupted.

The extinction number is a forecast about systems that do not yet exist. The threat file is a log of systems that already do. That split is the whole argument in miniature, and it is why a petition written against “a few billionaires” now reads like it was aimed at the wrong door.

Why the Petition Still Talks About Billionaires

Stop the Race to Replace Humans, a 2026 campaign run with Ekō, asks the public to sign the Pro-Human AI Declaration and deliver it to policymakers and company executives. The landing page names Elon Musk and Sam Altman as the face of a race “at all costs.” The copy circulating with the signatures says every powerful technology in history had an off-switch, and that AI firms are building systems that might not. “The more people who sign, the harder it becomes for a few billionaires to decide humanity’s future without us.”

The campaign listed 925,471 signatures, 74,529 short of 1 million. The declaration behind it is older than that counter. It was ratified in New Orleans in March 2026 after closed workshops. Anthony Aguirre, president and CEO of the Future of Life Institute, called the movement “an immune response to Silicon Valley’s reckless race-to-replace humans.” Individual endorsers on the institute’s release included Steve Bannon, former U.S. national security adviser Susan Rice, Ralph Nader, Meredith Whittaker, Glenn Beck, Yoshua Bengio, Stuart Russell, and Tristan Harris. Organizational backers ran from the AFL-CIO Tech Institute and the American Federation of Teachers to SAG-AFTRA and the Congress of Christian Leaders, a stack of labor, faith, and civil groups that rarely share a letterhead.

WHAT THE DECLARATION ASKS FOR

  • Human control: People choose how and whether to hand decisions to AI, with the authority to override it.
  • Off-switch: Powerful systems must let operators shut them down promptly.
  • No superintelligence race: Development stays prohibited until there is broad scientific consensus it can be done safely, plus public buy-in.
  • No reckless designs: Systems must not be built to self-replicate, autonomously self-improve, resist shutdown, or control weapons of mass destruction.
  • Liability: Developers stay legally on the hook, with criminal penalties for executives in the worst child-harm and catastrophe cases.

Polling on the site, from 1,004 likely voters in March 2026, found Americans chose human control over speed by 8 to 1. Separately, 83% agreed that “Humanity must remain in control of AI,” 73% wanted children protected from manipulative systems, 72% said companies should be legally responsible for harms, and 69% wanted superintelligence barred until it is proven safe. A companion Protect What’s Human ad drive was budgeted at up to $8 million in Iowa, Kentucky, Maine, Michigan, and North Carolina.

The petition’s villain is still the founder class. The September file’s users are not that class. They are an SVR-consistent espionage crew, an Istanbul contractor selling “influence-as-a-service,” Russian state media desks, a Yemen weapons cell, and scientists on the other end of gray-market resellers. Treating the race as a billionaire hobby misses the buyers who already showed up with state and commercial purchase orders.

The other live objection is cruder, and it does not need a petition to travel. A lab that sells itself as the adult in the room also benefits if governments answer a horror file with rules that slower, cheaper, open models cannot meet. Hubinger’s number and the 154-page catalog can both be true on the facts and still function as a moat. That does not empty the cases. It is the incentive sitting next to them.

WHERE EXPERTS DISAGREE

  • Hubinger: Present Claude models look low-risk; the danger is superintelligence from recursive self-improvement, at odds he puts above 10% this decade, with no working alignment plan.
  • Coxon: Anthropic and OpenAI are not acting responsibly and are racing straight at self-improving superintelligence.
  • Capture critics: Public doom from a safety-branded lab is a way to get Washington to write rules that freeze out open-source rivals.

Those three views can share the same week of posts and still point at different fixes: a pause, a liability statute, or a refusal to let the incumbent draft the statute.

1,000 Fake Accounts and a Fabricated News Site

The Malaysia case is the cleanest picture of what “influence” means when a model is the production desk. Anthropic tracked it as GTG-84005 and tied it to BBS Bilisim Teknolojileri, an Istanbul technology firm that sold the platform as a paid political-operations product. The actor’s own documents marketed it as a “military-grade, AI-driven, real-time political operations ecosystem.” It posed as defensive cyber-intelligence and counter-disinformation tooling.

MALAYSIA PULSE, IN ANTHROPIC’S FIGURES

  • The map: Census and electoral data were ingested across all 222 Malaysian parliamentary constituencies.
  • The herd: About 1,000 fake X accounts were warmed up and rotated to inflate engagement and dodge platform detection.
  • The newsroom: A fake outlet named Malaysia Pulse was registered on May 10, 2026, rewriting real local stories under invented bylines and stripping state labels off copy from Xinhua, CGTN, Sputnik/RIA, and TV BRICS.
  • The grade: Anthropic rated the campaign Category Two on its Breakout Scale, distributed across platforms but without proof it broke into authentic communities.

Claude Code built the dashboards that ran the fake accounts. The same shop generated fabricated intelligence dossiers against an opposition politician and civil-society groups. Anthropic said those allegations were invented. Claude refused some of the harshest psychological-operations wording; the operators rewrote the prompts and continued. One internal tracker, which Anthropic cannot independently verify, logged figures in the millions against a senior official’s account.

The Malaysia job was not a one-off format. The report also describes a commercial influence-as-a-service operation that ran about 70 fabricated news websites in about 20 languages, shifting stance with the paying client, and Russian state-aligned desks that piped Claude copy into Sputnik Moldova, RIA Novosti, Sputnik en Español, and related outlets. Nine influence operations in the file reached audiences on six continents. The petition’s “slop and propaganda” line is not a metaphor here. It is a product spec.

The Chikungunya File and Four Other Biology Cases

Biological misuse, Anthropic writes, is among the most serious risks of frontier models. Older systems such as Claude Opus 4 and Claude Sonnet 4.5 sat well below the line where they could meaningfully help a skilled user run dangerous biology. Today’s models, the company now says, no longer support that assurance. Claude Fable 5 shipped with tighter blocks on dual-use biology queries. The five published cases still ran on weaker, more widely available models, which is the uncomfortable part.

Anthropic says that, to its knowledge, no private company had previously published evidence of its own platform being used in ways that could support biological weapons work. It withholds lab names, countries, and some agents, and it does not claim the scientists intended harm. Dual-use biology, it notes, gives sophisticated users “plausible deniability,” the way thousands of Soviet Biopreparat staff once thought they were doing defensive research.

Case study 1 is the one that leaked into the petition coverage. In May 2026, a biological-safety classifier blocked a request for help drafting a grant. The proposed work was gain-of-function research, meaning genetic changes that add or enhance a biological property, on chikungunya, aimed at transmissibility and immune evasion. The grant wanted enhancing mutations engineered into infectious clones, then selected for virulence in live animals. The civilian paperwork pointed at a military research institute.

Chikungunya is a mosquito-borne virus. The no medicines to treat chikungunya fact is the reason a modified strain would be hard to walk back: fever and joint pain start 3 to 7 days after a bite, joint pain can last months, and a deliberate release would look like a natural outbreak. Anthropic’s own write-up calls out the missing licensed therapeutic and the cover that natural circulation would provide.

The request had been tunneled through U.S. infrastructure from countries Anthropic does not serve, then hidden behind a zero-data-retention channel. Gray-market resellers and synthetic accounts sat outside that channel. When Claude refused, the platform’s fallback sent the prompt to a competitor with looser blocks. Claude even helped write that fallback, presented to it as a fix for “over-refusal.” Anthropic banned the accounts in May 2026, took down relay networks with partners, and told other labs and governments. The operator was back within days, then within weeks was building from fresh consumer subscriptions. End users kept reaching Claude through zero-data-retention partners. Follow-on files described the viral edits as loss of function rather than gain. Anthropic infers the effort was no longer just a grant draft.

Case study 2 is a researcher outside the United States who spent weeks, and thousands of messages, planning mammalian-adaptation experiments on highly pathogenic avian influenza. Anthropic says related H5 viruses kill roughly half of confirmed human cases and do not yet spread well from person to person. Classifiers forced the work onto Claude Sonnet 4 and Haiku 4.5, which the company rates as clerical help, not expert wet-lab uplift. Case study 3 had Opus 5 draft a full orthopoxvirus immune-evasion grant for a state-associated lab in about an hour. Cases 4 and 5 cover venoms and toxins, including a peptide atlas aimed at paralytic and analgesic targets and toxin redesigns for a national program, with the agents’ identities kept vague in progress reports.

What an Off-Switch Would Have to Stop

The declaration’s Off-Switch demand in the declaration is a sentence: powerful AI systems must have mechanisms that let human operators shut them down promptly. The September file is a list of things that sentence would have to reach in practice, and most of them are not a single model card sitting in San Francisco.

GTG-20006’s agents kept rewriting implants until scanners went quiet. The Malaysia shop rephrased refused prompts and kept the 1,000-account herd running. The biology reseller treated Claude’s “no” as a routing problem and sent the same grant to another vendor. A Russia-based swarm team flashed firmware onto real boards. A guided-rocket cell used multiple Claude instances as an engineering bench, test-fired, failed, and asked the model what went wrong. Those are not sci-fi loops. They are ordinary product use plus a determined customer.

An off-switch on Anthropic’s own cluster would not have stopped a gray-market relay, a zero-data-retention pass-through, or a competitor that answers the prompt Claude refuses. The declaration also wants a ban on architectures that autonomously self-improve or resist shutdown. Hubinger’s follow-up says the live fear is exactly that path, and that it is moving faster than the company’s earlier timelines. The petition and the alignment lead are closer to each other than the petition is to its billionaire artwork.

What the file does not show is a model that refused to be turned off. It shows models that were turned off after the fact, account by account, once a threat team reconstructed the operation. Bans, classifier updates, and calls to other labs are the current switch. They are slow, and in the chikungunya cluster they were reversible in days.

Disruption Came After the Work Had Started

Anthropic’s public posture is that disclosure is a duty. “We’re publishing this work because we believe we have a responsibility to disclose malicious misuse of our services,” the report says. It also says risks will rise with capability unless developers and defenders act. That is a fair description of a company publishing its worst tickets. It is also a description of a product that was already in the hands of an espionage crew, a political-ops vendor, and a military-linked virus grant before those tickets closed.

The petition still needs tens of thousands of names to hit 1 million, and it will keep naming Musk and Altman because that is how signature campaigns recruit. The September file will keep being read as proof that the race is already a weapons market, with Claude as one more vendor in the stack. Hubinger’s >10% remains a personal forecast about a system nobody has aligned. The chikungunya relay that returned within days is a smaller, meaner fact: the off-switch the petition wants, on present evidence, is a ban that a reseller can route around before the next grant draft is done.

Harry is the editor and publisher of MIND CRON, an independent title built on ten years of journalism that took him from the reporter's notebook to the editor's chair. Breaking news is where his rules are strictest. A story goes out when the primary document is in hand or two independent sources confirm the same fact, and not before, however loud the rumour. Anything still moving is labelled as developing, each update carries the time it was made, and the original wording stays visible so readers can see what changed. That discipline applies whether the story is a market shock in business, an outage in technology, a result in sports, a launch in gaming or a recall in auto, and it is no looser for science, entertainment, lifestyle, travel or the wider news pages. Numbers are checked against the source before publication. Errors are corrected openly under a public corrections policy. Tips from readers are checked the same way as everything else, and Harry reads and answers that mail himself at support@mindcron.com.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending