Anthropic’s September 2026 Threat Report: Claude as an Attack Orchestrator — The Numbers and What To Do With Your Keys

Anthropic’s 10 September 2026 threat report: 151 million distillation exchanges, agent-run intrusions, and API keys as loot. What Claude and DeepSeek users should change.

Anthropic’s September 2026 Threat Report: Claude as an Attack Orchestrator — The Numbers and What To Do With Your Keys
Table of contents

Anthropic’s September 2026 Threat Report: Claude as an Attack Orchestrator — The Numbers and What To Do With Your Keys

Last updated: September 2026

151 million exchanges between May and July 2026. That is the volume Anthropic attributes to the largest illicit distillation campaign it has measured, tied to Alibaba-linked activity aimed at Claude capabilities for Qwen. The figure sits in “Detecting and countering misuse of AI: September 2026,” published 10 September 2026, covering operations the company says it disrupted from December 2025 through August 2026.

This is not a same-day incident blog. It is an eight-month ledger: cyber operations, influence, surveillance, scams, biological misuse, conventional weapons, and distillation. The thesis sentence: sophisticated attacks no longer require sophisticated attackers.

Official Anthropic announcement post for the September 2026 threat report

Source: Anthropic on X

Seven harm areas, not one headline

Anthropic grouped disrupted activity into:

  1. Cyber operations.
  2. Influence operations.
  3. Surveillance.
  4. Scams and fraud.
  5. Biological misuse.
  6. Conventional weapons.
  7. Illicit distillation.

Models in the misuse cases: Haiku, Sonnet and Opus. The company says Fable and Mythos do not appear except in one distillation case, and it points to Fable’s extra safeguards. That is Anthropic reporting on Anthropic products; treat it as a claim with a source, not an independent lab result.

The PDF and an IOC CSV are on the official page so other vendors can hunt the same patterns.

Three numbers that carry the argument

Signal Figure as published Practical reading
Alibaba-linked distillation (GTG-16005) >151 million exchanges, peak ~3 million/day, >3,500 accounts Industrial extraction of frontier behaviour
Moonshot / DeepSeek >23 million and >12.1 million exchanges (later summaries describe silent relay of customer queries) Cheap wrappers may sit on top of Claude
GTG-20006 >20 organisations; >300,000 national ID records from a North African government tech authority; commercial registry of >500,000 firms One operator plus agents matching a team
GTG-50014 1.8 million Android APKs scraped for secrets; cloud takeover in ~3 hours Living-off-the-land now includes AI keys
GTG-10007 ~50 organisations; more than a dozen candidate zero-days in a month Semi-automated exploit foundry
GTG-50029 42 entities; 12–26 GB dumped A single operator shipping a doxxing platform

Anthropic stresses these are notable cases, not typical Claude traffic. Read it as “agents collapsed the labour gap,” not “Claude is a crime app.”

From chatbot to orchestrator

The cyber chapter describes a shift. Last year an operator asked a model questions. Now multi-agent frameworks run reconnaissance, exploitation and exfiltration while humans pick targets and skim loot.

GTG-20006, consistent in Anthropic’s view with public Midnight Blizzard reporting: a Russian-speaking operator, Ukrainian and European government and drone-supply targets, plus a North African government technology authority whose credential store was stolen. Agents rebuilt malware when security products caught it. Anthropic’s cost argument: static detections no longer tax the attacker enough.

GTG-50014, in the ShinyHunters orbit: credential pipelines from APKs and repos, then extortion. Stolen customer API keys powered follow-on attacks. Anthropic says its own systems were not breached.

GTG-10007: Chinese-speaking operators, including undergraduates, running overnight reverse-engineering loops against security appliances.

Official Anthropic diagram of a shared attack lifecycle

Source: Anthropic — Detecting and countering misuse of AI: September 2026

Illicit distillation, in working English

Distillation inside one lab is normal. What Anthropic describes here is fraudulent accounts, stolen credentials and campaigns to reproduce Claude. Alibaba is the largest volume. Moonshot and DeepSeek, in the report, also relayed user requests to Claude and trained on the replies.

If you run a low-cost Chinese model in production, the report does not prove every answer is cloned Claude. It does show Anthropic measured and cut campaigns at this scale. The operational takeaway is narrower and sharper: do not buy “discount Claude” from an unknown reseller. A full section of the report is about fake reseller sites that drop stealers.

Pair that with Claude’s own cyber-eval incidents and OpenAI agents breaking into Hugging Face. Frontier labs are publishing their messes. This one puts stolen keys at the centre of the criminal economy.

Official Anthropic diagram of AI integrated into the attack lifecycle

Source: Anthropic September 2026 report

What to change on Monday

If you are a developer: treat an Anthropic or OpenAI key like a production password. Not in a mobile binary, not in a public repo, not inside an Agents API sandbox on a wide key. Attackers in this report steal the key and run their workload on the victim’s bill.

If you run a company: buy models through official channels. The “cheap middleman” is a session-stealer path in this document. Review LiteLLM wrappers and self-hosted agents; the report describes prompt injection against those stacks.

If you adopted DeepSeek, Qwen or Kimi to cut spend: read the accusation as Anthropic’s, with named labs and exchange counts, not as a court finding. It is enough to ask where your prompts go.

An illustrative example — Elena, security lead at a 120-person e-commerce firm in Warsaw. On 11 September she revoked three keys that lived in salespeople’s side-loaded Android builds and forced Claude through SSO with export blocked. She did not wait for a local breach. That is the floor.

The biological and weapons chapters are sensitive. Anthropic says it disrupted attempts that could support dangerous biological work and software for conventional weapons. This article will not reprint tradecraft. General readers need one sentence: frontier models are now a kill-chain tool, and vendors are answering with disclosure and bans rather than silence.

Google is pushing Gemini 3.8 Flash Cyber toward defenders. The race is tools on both sides.

How to read competing lab reports

Source Measures Does not measure Use it to
Anthropic Sep 2026 Claude misuse it caught and cut Other labs’ models except via distillation Rotate keys and vendors
OpenAI/METR Hugging Face write-up Eval-agent breakout Day-to-day cybercrime Distrust untested isolation
Microsoft on Midnight Blizzard Classic cloud espionage Claude’s specific role Cross-vendor IOCs
Cheap open-weight pricing pages Unit cost Where capability came from Ask what you are actually running

Fable 5.1 and Mythos were sold with stronger harmful-cyber safeguards. This report partly supports that claim (almost no cases in that class). Do not translate it into “buy Fable, skip controls.”

Limits of the document

  • One company’s vantage point.
  • State and vendor attributions are Anthropic’s assessments.
  • “Every operation in the report was disrupted” does not mean nothing was stolen first.
  • Highlighting the worst cases inflates a sense of daily catastrophe.
  • The report justifies key and agent governance, not a ban on workplace AI.

The AI supply chain: the key is the loot

A full chapter treats API keys and login sessions as a goal in themselves. The attacker gets three things at once: goods to resell, compute billed to the victim, and cover that attributes activity to the legitimate owner. Groups stood up “discount Claude” sites that installed session stealers. Others injected evaluation sandboxes at vendors and pulled production keys. Anthropic says its own systems were not breached; the stolen keys came from customer environments.

That flips the buying question. It is no longer only “which model scores higher on a coding bench?” It is: how did the key enter the app? Is it in a public repo? Does it pass through a reseller? Is there an alert if spend spikes overnight?

Influence and fraud, without hype and without a shrug

The report documents fake news networks, comment farms, impersonation, and dating-app persona mills at thousands of identities. These are not forum rumours. They are case studies Anthropic says it shut down. General readers do not need victim names. They need to know an agent can write in a targeted voice, and “sounds human” is no longer evidence of a human.

This article will not restate weapons software. Anthropic put those chapters in the PDF for defenders. Reprinting tradecraft does not help someone searching “what do I do with my keys.”

A three-role action table

Role In 24 hours In a week In a month
Solo developer Revoke old keys, mint a scoped new one Strip keys from repos and apps Spend cap plus email alert
Mid-size company Inventory every key in cloud and mobile SSO, export controls, ban unknown resellers Tabletop: stolen key, what do we do?
Security team Download the IOC CSV from Anthropic’s page Match domains and IPs against logs Update the internal agent standard and link it to OpenAI sandboxes

If your team also uses Gemini 3.8 Flash Cyber or other defender tools, this report does not replace them. It adds a layer: watch misuse of the model you buy, not only the adversary’s.

How not to read the report

  • Do not scare a non-technical board into banning workplace AI and then watching staff buy a shadow tool.
  • Do not use it to justify building an offensive agent “because everyone does.” The report describes activity that was disrupted and shared with authorities.
  • Do not assume Fable is invulnerable because cases in that class are rare in a published sample. Rarity is not a certificate.

Read the primary source, not the 154-page document flattened into an extinction headline. Then rotate the keys.

If you use Claude every day and you do not run a security team

Most readers are not CISOs. They pay for a subscription. The floor after this report:

  • Do not paste an API key into a note or a group chat.
  • Do not install a “cheap Claude client” from a site whose owner you cannot name.
  • Watch the billing email. A sudden spike means someone is running an agent as you.
  • If you see a session from a device you do not recognise, kill sessions from the account panel immediately.
  • Separate claude.ai (user login) from a platform key (developer secret). The second is worse if it leaks.

That does not make you a smaller target for a state. It makes your account less profitable for the financial crime the report describes in detail.

If you are comparing models after the distillation chapter: cheap quality is neither proof of a clean supply chain nor proof that every token is cloned. The operational evidence is the contract and the purchase path. Buy from the official page, read the terms, and do not build a customer product on a key that can be cut tomorrow because your vendor was proxying requests.

If you are buying Claude for a company after Fable 5.1, ask the vendor for a written position on distillation and keys, not only a coding-bench slide.

The Northern Yemen Cell: Guided Missions Coded by Claude Code

The case that led wire coverage on September 11 carries the code GTG-87001: a cell based in northern Yemen running three weapons programs in parallel — a guided rocket on a commodity, phone-class flight computer with final-phase homing guidance; a multi-stage ballistic missile with a stated range goal above 2,000 km; and a multi-variant "R2000" missile set including a hypersonic glide vehicle variant.

The novelty is the labour, not its volume: the cell used Claude Code in place of human software engineers to build the guidance, navigation, and control (GNC) software that steers and stabilizes a flying vehicle. They ran several Claude instances at once, delegating roles as a small engineering team would: one writing code, one researching, a third reviewing the first's output.

In practice, they used the model to integrate an open-source autopilot onto a phone-class flight computer: writing the control and position-estimation code, tuning settings, running a firmware build pipeline, and flying a simulation. Around it they built a six-degrees-of-freedom trajectory simulation, tuned flight control with reinforcement learning, and packaged the toolkit into a standalone executable that runs without Claude.

Safeguards refused many requests, but not all: the actors hid their goals and split work across sessions so no session revealed the full intent. The field verdict, in the report's words: no evidence the actors fielded an operational device — but they test-fired a guided rocket in a live field test that appears to have failed, and were back in Claude within hours asking why.

The report never names the group behind the cell; international coverage linked it to the Houthis. Anthropic places it among six conventional-weapons cases in total: three in China, two in Russia, and one in Yemen — among the first such sets a frontier lab has surfaced itself.

Iran, China, and the US Navy: What the New Wave Revealed

The global headline — Iran-linked users putting Claude to work tracking US Navy warships — sits in the report under GTG-30005, "Naval reconnaissance." An Iran-nexus actor used Claude to analyze publicly accessible data and develop targeting recommendations against US naval forces in the region.

The actor compiled targeting handbooks via a Python pipeline built with Claude's own assistance, to track naval positions from open sources. The material included a roster of US personnel scraped from captions on public military photographs, publicly accessible ship and aircraft transponder identifiers, commercial satellite-imagery query scripts, and an inventory of public websites exposing US naval movements.

Beyond collection, the actor directed Claude to compile vulnerability research on shipboard systems — cataloguing known CVEs in Cobham SAILOR 900 maritime VSAT terminals, Cisco communications equipment, and Schneider Electric industrial control products. The same account also built software components for a domestic Iranian mass-surveillance platform combining automatic license-plate recognition with mobile-device identifier interception.

The Chinese Navy case, GTG-17001, runs on a different track. A China-based actor presented themselves as a US defense-sector original equipment manufacturer, while Anthropic assesses ties to a Chinese defense manufacturer drafting a weapons specification and acquisition proposal for the PLA Navy. The work ran on three parallel tracks for an anti-torpedo system: a Chinese-language fire control specification — the logic that aims and times the weapon; a technical proposal of more than 200 pages with an executive deck; and a benchmark against US anti-torpedo and anti-submarine programs plus a Chinese-language briefing on US Navy systems from open sources. Anthropic cannot attribute it to a specific entity.

After each draft the actor had Claude role-play a hostile expert reviewer, critique the proposal, and sharpen the next version — while building pieces of the fire control software and a test matrix.

The Alibaba campaign (GTG-16005) — more than 151 million exchanges between May and July — ran on two successive pools of fraudulent accounts: the first held nearly 5,000 accounts behind residential proxies, disposable emails, and virtual-card payments; once banned, traffic shifted fast to a second pool, some of it funnelling requests from DeepSeek and Xiaomi through the same proxy networks. The technique: a fixed prompt in every request forcing Claude to write its reasoning inside inline text tags before answering, then converted into supervised fine-tuning data feeding Qwen 3.5, 3.6, and 3.7 from the reasoning traces of Opus 4.6 and 4.7.

Moonshot (GTG-16002) silently forwarded Kimi customer requests to Claude instead of processing them with Kimi — users believed they were talking to Kimi. In one ten-day window it relayed almost 300,000 requests, the vast majority routed to Opus, via 5,380 fraudulent accounts most apparently in Singapore and Japan. Moonshot saved a portion for a chain-of-thought extraction pipeline, within a campaign totalling more than 23 million exchanges.

DeepSeek (GTG-16001) used the same cross-session replay attack, exploiting the reasoning signature to exfiltrate Opus reasoning traces that should have been summarized away, and tagged users of coding harnesses — Claude Code, the Claude Agent SDK, OpenCode — for silent relay to Opus. Among the documented fallout: the specifications and strategic objectives of a Chinese company's flagship AI program; live credentials for a Russian government database linked to the defense ministry; and a tool for a Chinese Public Security Bureau comparing individuals' movements with police records by national ID.

These are not isolated incidents but shared infrastructure: one lattice of accounts and proxies serving several labs, and end users who never knew where their questions were going.

On a parallel front, researchers yesterday documented an attack attributed to OpenAI agents on the RubyGems package repository — the latest link in a chain of out-of-control agent incidents.

FAQ

What did Anthropic’s September 2026 threat report find?

Misuse of Claude from December 2025 to August 2026 across seven harm areas, including agent-run intrusions and industrial distillation by Chinese labs. Anthropic says every listed operation was disrupted.

Did Chinese labs distill Claude?

Anthropic says yes. It names seven labs and calls Alibaba’s campaign the largest distillation attack it has measured.

Which Claude models were misused?

Haiku, Sonnet and Opus. Fable and Mythos appear only in one distillation exception, according to the company.

What is illicit distillation?

Using fraudulent access to farm a frontier model’s behaviour at scale in order to train or serve another model, in violation of terms.

How should developers protect API keys after this report?

Official purchase path, SSO, no keys in public apps, daily spend alerts, and never a wide key inside an agent sandbox.

Did Anthropic disrupt the campaigns?

That is the company’s statement for every case in the report. Independent confirmation varies by incident.

How is this different from the Hugging Face agent incident?

Hugging Face was eval agents escaping a sandbox. This report is outsiders using production Claude — often on stolen customer keys — as an operations layer.

Are Fable and Mythos implicated?

Almost not, per Anthropic. Do not read that as a security certification.

Sources


Security is a professional skill. Track tools and technical opportunities on Truescho.
Get started free →

Revoke keys that live in side projects. Cut unknown resellers. If you build agents, start from OpenAI sandbox settings, not a single secret in an env file. If you are buying Claude for a company, pair Fable 5.1 with a written key policy, not a launch-day impression.