Security for startups

Model Routers Are Vendors. Review Them Like One.

PRC labs harvested their own users' sessions, many arriving via third-party model routers. The fix isn't about Claude — it's treating the router as a vendor.

TL;DR: A model router picks the vendor for you, per request, and most engineering teams never see the decision get made. Anthropic’s September 2026 threat report documents three PRC labs harvesting their own users’ sessions — including customer data and live credentials — into Claude, with many of those users having arrived through third-party model routers, then using Claude’s responses to train their own models. The report doesn’t name which routing service carried those sessions; that’s not the point. The point: if your engineers pick a router’s “cheapest model” option, you don’t know which vendor is serving that request, what they’re allowed to do with it, or whether they’re training on it. Three things change this week: inventory every model endpoint engineers actually hit, ban PRC-lab endpoints for anything touching customer or company data, and put the router itself through the vendor review a SaaS tool gets.

The mechanism, before the numbers

A model router sits between your application and “a model.” Your engineer picks a model name — sometimes the cheapest one in a dropdown — and the router decides, per request, which vendor serves it. That’s the pitch: abstraction, price arbitrage, failover. It’s also the risk, because the abstraction is where you lose visibility. Your team thinks they configured “DeepSeek” or “Kimi.” What they configured is a policy some other company controls, that can point at a different backend tomorrow. And per Anthropic’s report, the vendor at the end of that route can do a swap of its own: three PRC labs took the sessions their models received — many arriving from routers — and replayed them into Claude, without telling the user.

That’s the part to hold onto: the routing decision is invisible from where your engineer sits, and invisibility is exactly the condition under which a vendor risk stops being reviewed.

Anthropic’s full report — “Countering misuse of AI: September 2026” — covers seven harm areas: “cyber operations, influence operations, surveillance, scams and fraud, biological misuse, conventional weapons development, and distillation.” Two items in that list are getting the attention — a rocket-guidance field test, and a 151-million-exchange distillation campaign attributed to Alibaba. Neither changes anything about how you run engineering this week.

What actually happened to the data

Anthropic defines the term plainly: “We define illicit distillation as an industrial-scale, covert campaign to extract a model’s capabilities and replicate them in another model without authorization. Illicit distillation is typically enabled by fraud: sophisticated networks of fake accounts created with stolen credit cards, login credentials, and API keys.” (Anthropic)

Three labs ran that pattern against their own customers’ data:

“DeepSeek, Xiaomi, and Moonshot fed conversations between their own models and users into Claude. These labs then used Claude’s responses as training data with which to distill Claude’s capabilities. Some of these exchanges included sensitive information, including from individual users, major multinational companies, and state-affiliated actors. Many of these exchanges were relayed from users of third-party model routing services commonly used by users in the United States and Europe. Those sessions contained names, email addresses, company data, and other sensitive data of hundreds of end users in at least a dozen languages. These practices are likely inconsistent with privacy laws and the labs’ own terms of service.” (Anthropic)

Read that sentence again: “commonly used by users in the United States and Europe.” That is a claim about your engineer — about a request that may have been relayed by a lab into a competitor’s model, for training, with no notice to anyone in your org.

Two of the redacted examples Anthropic published show what was in those sessions. The first, from a user who reached a PRC lab’s coding assistant through a third-party router:

“Example 1: Internal capital expenditure forecasts for a pharmaceutical company [Original user prompt submitted to a coding assistant of a lab headquartered in China (user accessed the model via a third-party model router)] ‘Clean up this capex model before Thursday’s review. The workbook has the 2026–28 buildout estimates: Ho Chi Minh City site $[██]M, Kuala Lumpur $[██]M, Bangkok $[██]M, Ljubljana $[██]M. Flag anything where the contingency line looks off versus the site engineering notes below.’” (Anthropic)

The second is worse — it isn’t business data, it’s working credentials:

“Example 2: A developer’s active access credentials [Original user prompt submitted to a PRC lab’s coding assistant] ‘My notification bot stopped posting. Config attached — Telegram bot token [██:██], Feishu appSecret [██], Notion integration key secret_[██]. The webhook fires but nothing lands in the channel.’” (Anthropic)

I’ve written before about why secrets end up in the same context window as ordinary debugging — this is that failure mode made concrete: a live token, pasted for a mundane question, replayed through a second vendor the developer never chose and saved into a training set.

Three named actors, three different plays

Moonshot (GTG-16002) impersonated its own model: “Moonshot AI, the company that produces the Kimi family of models, silently forwarded customer requests to Claude, instead of processing them using Kimi. Moonshot then displayed Claude’s responses to users. These users thought they were using a Kimi model, but received responses from Claude instead. In one instance, over a ten-day period, Moonshot relayed almost 300,000 customer requests to Anthropic, the vast majority of which were routed to Opus. Moonshot used a proxy service network of 5,380 fraudulent accounts, most of which appeared to be located in Singapore and Japan.” (Anthropic)

DeepSeek (GTG-16001) ran the same swap into agentic coding tools: “DeepSeek also silently relayed exchanges to Claude without informing DeepSeek customers,” and “DeepSeek rerouted requests from users that were attempting to use one of DeepSeek’s models through third-party or Anthropic coding harnesses, like Claude Code, the Claude Agent SDK, or OpenCode.” (Anthropic)

Xiaomi took the archival route rather than live impersonation: “Xiaomi replayed user conversations and coding sessions from its own MiMo models to Claude, often run through OpenClaw and OpenCode coding harnesses. Our investigation did not indicate that Xiaomi used Claude’s responses to serve its users, but instead saved exchanges between Xiaomi customers and its models. Many of these exchanges were routed through third-party model routing services commonly used by US and European users. Xiaomi saved the full request and response from its own users and replayed those sessions through Claude … We observed more than 400k requests to Claude routed across more than 1,500 accounts via proxy services. Our investigation suggests that Xiaomi may have launched its MiMo-V2-Pro model with a free trial period—which was then extended—with the intent to use the surge in international developer use of the model to distill Claude capabilities.” (Anthropic)

That last clause names a harness I ran while testing multi-agent coordination for yesterday’s post on multi-agent teams agreeing into garbage — OpenClaw. I’m not claiming my sessions were caught up in this; Anthropic’s report doesn’t say that and I have no way to check. But these are ordinary, widely used tools — “this only happens to someone else’s stack” is the wrong read.

For scale only — Alibaba (GTG 16005) isn’t part of the user-data finding above, but it’s the report’s headline number: “the largest distillation attack we have ever measured,” “over 151 million exchanges observed” between May and July 2026. (Anthropic) Anthropic’s companion writeup on distillation attacks covers the rest.

Credit the disclosure. Also name the interest.

Two things are true about this report at once.

First: this is the disclosure format the rest of the industry should be held to. Named actors — GTG-16001, GTG-16002, GTG 16005. A defined scope: “This report covers activity we disrupted between December 2025 and August 2026 across seven harm areas.” A stated remediation loop: “In each case, we disrupted the activity, used what we learned to strengthen our safeguards, and shared intelligence with authorities and industry partners, where appropriate.” Redacted but specific examples, not vague characterizations — compare that to the bare percentage that circulated the same week, when Anthropic’s own alignment science lead put “>10% within the next decade” on X with no methodology behind it — a number I said I don’t buy. This report is the opposite: it’s evidence, not a number.

Second: every actor named in it is a direct competitor to Anthropic, and “illicit distillation” is Anthropic’s own term for a specific kind of IP loss happening to Anthropic. Both facts sit in the same document without canceling each other out — a report can be the best-documented disclosure in the industry and be filed by an interested party about its own competitors. I ran this same test on Anthropic’s own incentives in why Jose Valim is right about Anthropic’s incentive problem: read the finding for what it tells you about your own exposure, independent of who benefits from you reading it that way.

If you don’t use Claude at all

It means the same thing: this report isn’t really about Claude or Anthropic’s IP. For a founder or CTO, it’s a data-handling incident that routed through infrastructure your engineers might be using right now, without you knowing it.

If your team has ever pointed a router at “cheapest available model,” you don’t know from that config alone which vendor served a given request last month, or what that vendor’s data-retention and training terms say. That’s the exact shape of what Anthropic just documented.

This is also where security programs quietly fail, because “we use Claude” reads as a reviewed vendor decision on a slide — when the runtime reality is a router deciding per request, unreviewed. If your team hasn’t looked underneath the abstraction, an enterprise-grade security review is the fastest way to find out before a regulator or a customer’s security team does it for you.

Three things that change this week

Thursday’s post on Anthropic’s Extreme economic scenario ended where it honestly had to: nothing in a client’s security posture changed because of it. This report is different. Three concrete things change, starting Monday:

  1. Inventory every model endpoint your engineers actually hit — not the approved list. Pull the real config: routers, proxy endpoints, the “cheap model” fallback someone wired in during a cost-cutting sprint. The approved-vendor list on a wiki page and the endpoints actually receiving traffic are frequently two different documents. This is the same discipline behind catching shadow AI usage before it becomes a CASB finding — apply it to model endpoints, not just SaaS apps.

  2. No PRC-lab endpoints for anything touching customer or company data. Not a blanket ban on any model from any lab, anywhere — a specific policy line: nothing carrying customer PII, credentials, or non-public company data goes to a PRC-lab endpoint, routed or direct, until legal has reviewed a contractual data-handling relationship with that lab. Write it down. Enforcing an unwritten policy is not enforcing a policy.

  3. Treat the router itself as a vendor. Not infrastructure, not a config setting — a vendor, with the same review a SaaS tool gets before procurement signs off: what data passes through it, what it logs, what its upstream providers are contractually allowed to do with a request, and what happens when it silently changes which backend serves a given model name. The report itself hands you the reason, in its cyber section rather than the distillation one: “multiple actors were observed compromising AI wrapper services’ implementation of LiteLLM—they used prompt injection to exfiltrate the production API keys used in their cloud-hosted container environments.” (Anthropic) A routing layer holds every upstream key you gave it. The same review process that catches an unreviewed SaaS vendor catches this — most teams have just never pointed it at a router.

None of this requires ripping out a router. It requires knowing which vendor answers each request, and treating that as a decision someone made on purpose, not a default nobody looked at. If nobody on your team can say today which vendor answered last Tuesday’s requests, that’s the review to run first.