Jose Valim Is Right: Anthropic's Incentive Problem

José Valim says Anthropic hasn't separated its security case from its commercial interest. He's right — and it's every frontier lab's problem.

TL;DR: Anthropic published its position on open-weights models on July 27 — and contrary to what the resulting news cycle implied, it does not call for a ban. It asks for chip export controls, restrictions on industrial-scale distillation by authoritarian states, and mandatory pre-release safety testing that Anthropic says “would need to be global” to work. José Valim, the creator of Elixir, responded on X with a critique sharper than most of the reaction I’ve seen: you can’t ask for global cooperation on safety testing while arguing your own country should pursue policies that preserve its strategic advantage, and you can’t expect the public to read your safety argument as neutral when it also protects your business. Valim’s still right, on both points, and it’s worth being precise about why.

What Anthropic actually said

Start with the part most of the commentary skipped: what the post says, not what the internet decided it said.

On July 24, Nvidia CEO Jensen Huang posted an open letter urging Washington not to restrict downloadable AI models. It went out with 25 signatures and doubled to 50 within a day, including OpenAI, Google, AMD, Cisco, GitHub, and Block. Two names were conspicuously missing from every version of the list: Amazon and Anthropic. Given that Amazon is Anthropic’s largest investor, the pairing of absences read, to a lot of people watching, like a signal — and the obvious inference was that Anthropic wants open weights restricted or banned.

Three days later, Anthropic answered directly — in a post signed by Dario Amodei personally — and the actual position is narrower than the inference. The post is unambiguous: “Anthropic has never advocated for a ban on open-weights models.” Amodei states it twice, opening and closing: “Protectionist bans would not address my most serious national security concerns,” and, in the summary, “we have not and are not advocating for a ban on open-weights models as a category.” That’s a real position, stated plainly, worth taking at face value rather than assuming the worst version of it because a competitor’s letter didn’t have their name on it.

What the post does ask for is three specific things. First, chip export controls — don’t sell powerful chips or chipmaking equipment to China, and close the smuggling routes around existing controls. Second, restrictions on industrial-scale distillation operations that let authoritarian states cheaply extract capability from frontier models without doing the underlying research. Third, and the strongest of the three on the merits: mandatory pre-release safety testing for cyber, biological, and alignment risks, applied to all sufficiently capable models — open and closed — testing that Anthropic says “would need to be global, which means even the CCP would need to be on board.” The named threats are authoritarian states achieving durable military or surveillance superiority, and models capable enough to meaningfully assist cyberattacks or biological misuse.

None of that is a ban. It’s a specific, arguable policy ask — the kind of position I’d have written myself if I were running a lab that had spent five years building a safety-first brand and then watched a competitor’s open letter imply the opposite of its actual position without asking first.

Where Valim’s critique lands

José Valim read the same post and posted a response that doesn’t dispute any of the facts above. He grants the underlying security concerns are “solid.” His objection is structural, and it’s really two objections wearing one tweet.

The first is a game-theory point, and it’s the cleaner of the two. In his words: the post “tries to have it both ways: it calls for global cooperation on mandatory safety testing while simultaneously arguing that the US should pursue policies to preserve its strategic and economic advantage in AI.” His conclusion: “If states are expected to act in their own national interest, it is unclear why global powers would voluntarily participate while their own technological and strategic ambitions are constrained.”

Sit with that, because it’s not an abstract objection — it describes how the ask is actually structured. Chip export controls and distillation restrictions read, in Valim’s framing, as measures to preserve US advantage over China. Mandatory testing, in Anthropic’s own words, is a regime that “would need to be global, which means even the CCP would need to be on board.” Put those next to each other and read them the way a state actor would: one policy keeps you behind, the other asks you to accept a constraint on your own model releases in the name of shared safety. Why would the same government sign up for the second while the first is explicitly designed to keep it out of the race the second claims to be making safer for everyone? A regime you’re simultaneously trying to out-compete and cooperate with has no clean reason to pick cooperation on terms written by the side asking to stay ahead.

Anthropic isn’t blind to this. The post offers an answer: cooperation on bioweapons may be possible “because it is in China’s interest too.” That’s true, and it’s the best version of the case — but it’s load-bearing for exactly one of the three named risks. Nothing in mutual bio-interest explains why a state would accept pre-release gating on cyber capability or alignment while chip controls are explicitly designed to keep it a generation behind. The answer covers the narrowest threat and leaves the asymmetry Valim named untouched.

You can’t structure two of your three asks around keeping a rival behind and expect that rival to volunteer for the third.

That’s not a rejection of mandatory testing as an idea — it’s the strongest of the three recommendations precisely because it’s the one that, done right, reduces risk regardless of who’s ahead. It’s also the one most undermined by sitting next to the other two. If Anthropic wants global participation in a testing regime, the pitch has to be made on its own terms, separated cleanly from the strategic-advantage framing, or a rational state reads the whole package as a play for advantage with a safety label on it.

Valim’s second point is about perception rather than logic, and it’s the harder one to argue with because it isn’t really about Anthropic’s actual position at all: “the public perception is that Anthropic has not done enough to distinguish its security arguments from policies that also serve its commercial interests. As a result, its safety agenda is unlikely to be perceived as economically neutral.”

That’s a claim about how the argument lands, not whether it’s true. And on the evidence of the last week — a missing signature read as a smoking gun, a clarifying post that had to explicitly say “we have never advocated for a ban” because people had already concluded otherwise — it’s hard to argue the perception isn’t exactly what he says it is.

Why this critique is inconvenient for me

I want to be honest about where I’m standing before I say more, because it changes how much this critique should weigh.

I’m not a neutral observer of either side of this. Claude Code is the tool I use to write and ship code on this site — it’s in my daily workflow, not a vendor I’m evaluating from a distance. I’ve written about what a CLAUDE.md file looks like after fifty real commits and how TDD actually works with Claude Code in an Elixir codebase, because that’s genuinely how I build things now. And the backend I run this site on, and recommend to clients, is Elixir — Valim’s language. If there’s a reader who came in rooting for Anthropic to have a clean answer here, it’s roughly me.

That’s exactly why the critique is worth taking seriously instead of waving off. It’s not coming from a rival lab’s PR account or a competitor’s cheering section — it’s coming from someone whose credibility runs the opposite direction of “wants to see Anthropic look bad.” A builder with no stake in tearing the argument down reading it and still concluding it’s unlikely to be perceived as economically neutral is a stronger signal than the same critique from someone with an obvious axe to grind.

It’s not an Anthropic problem — it’s every lab’s problem

Here’s the part that keeps this from being a hit piece, and the part I think matters most: Valim’s second point isn’t really an indictment of Anthropic specifically. It’s a structural feature of lab-led safety advocacy generally. Every frontier lab’s safety position, traced far enough, ends up disadvantaging somebody else’s business model. A closed-weights lab’s positions on model access constrain open-source labs. Anthropic’s positions on testing and export controls constrain open-weights players and, more distantly, Chinese labs. A lab arguing for stricter oversight of “sufficiently capable models” is, not coincidentally, usually a lab already ahead on capability. This is the same skepticism I apply to any vendor’s stated policy — what does the person making the argument need to be true, the same test I ran on a vendor’s benchmark claim two days before this one — and does the argument cost them anything, or only someone else?

That question is the actual test for whether a safety position is economically neutral, and it’s worth running Anthropic’s three asks through it individually rather than treating the post as one bundle.

Chip export controls on China cost Anthropic nothing — they constrain a jurisdiction Anthropic doesn’t sell into and don’t touch its own model releases. Distillation restrictions cost it a little, not much: Anthropic commits to identifying and banning its own paying accounts caught doing this, but concedes those accounts are usually only identifiable after substantial distillation has already happened — a real cost, just a small and lagging one. Mandatory pre-release safety testing is the one asymmetric case, and to Anthropic’s credit, it’s the one that does cost something real: the post is explicit that testing applies to “all sufficiently capable models, open and closed” — meaning Anthropic’s own frontier releases go through the same gate it’s asking regulators to impose on everyone else. That’s a position that constrains the author, not just the target — the one piece of the three-part ask that clears the neutrality bar Valim is describing. The other two cost Anthropic little to nothing, and the post never draws that distinction itself. That’s the actual gap: not dishonesty, but a failure to separate the ask that’s genuinely self-limiting from the two that aren’t, at a moment when the audience is primed to read all three as the same kind of move.

The fix is structural, not rhetorical

I don’t think Anthropic is wrong about the underlying risk. Chip smuggling to sanctioned states is a real problem with real precedent. Industrial-scale distillation of frontier capability by state actors is a plausible, under-discussed threat. And mandatory pre-release testing, applied evenly, is close to the least controversial safety idea in the entire AI policy conversation — it’s hard to find a serious critic of testing capable models before release, on principle.

The problem isn’t the substance. It’s that the substance arrived bundled with two asks that visibly serve US strategic interest, from a company that also didn’t sign a competitor’s letter the same week, in an industry that has learned to price “safety argument” and “business interest” as correlated rather than independent. No amount of restating “we’ve never advocated for a ban” fixes that, because the framing problem was never about the ban claim — it’s about which asks get bundled together and who’s making them.

If Anthropic wants the testing recommendation read as neutral, the fix isn’t a better press statement. It’s separating the self-limiting ask from the two that aren’t, making that separation explicit rather than implicit, and probably having the testing case made by a coalition or standards body rather than the lab that benefits most from being first through the gate. Solid concerns, credible in isolation. Undermined by delivery. And delivery, unlike the physics of chip fabs and model weights, is entirely something a lab controls.

If you’re building a startup that has to make its own claims to investors, customers, or regulators credible — where the same “does this cost us anything” test gets applied to your own security and compliance posture — that’s exactly the kind of scrutiny I help founders get ahead of before someone else runs the test on them. Let’s talk.