Anthropic's Extreme Scenario Has a Precondition
Anthropic published an economic upside scenario and an alignment risk admission the same week. The upside needs the thing nobody has a plan for.
TL;DR: Anthropic’s Economics team published a scenario explorer this week putting an “Extreme” 2030 outcome (US GDP up 32.4%) on the table, and named its engine: “likely driven by recursively self-improving AI systems.” The same week, Anthropic’s alignment science lead said publicly that Anthropic does not yet have a plan to align superintelligence and is “not clearly on track to.” The upside case and the unsolved problem share one precondition. I’m skeptical of the specific number attached to the risk this week. It’s unfalsifiable, and it flatters the speaker either way it lands. What doesn’t change: a pre-Series-A founder’s security posture, this week, because of any of it.
Two documents, same company, same week
In September 2026, Anthropic’s Economics team published v1.0 of an Econ Scenario Explorer, built off their technical report Economic Scenarios for Transformative AI (Korinek, Jones, Sacher, Cotter, McCrory, 2026). It models three 2030 outcomes for the US economy at 2025 price levels: Modest (+1.6%, $34.1T), Substantial (+8.3%, $36.3T), and Extreme (+32.4%, $44.4T).
Only one of those three names its own engine. The page’s verbatim description of the Extreme case: “in the extreme scenario, AI drives a completely transformed, unprecedented economy, likely driven by recursively self-improving AI systems and a faster rate of AI adoption.” Not “faster models.” Adoption speed is in there too, but it’s paired with recursive self-improvement, and only in this scenario. That pairing is the thing that gets you from Substantial to Extreme.
On September 9, on X, Jacob Coxon, a pretraining researcher with three years across OpenAI and Anthropic, announced he was resigning, writing that “neither company is acting responsibly,” that they “are racing straight to self-improving superintelligence and gambling with our lives,” that “the people building AI earnestly believe that it could kill us all by the end of the decade,” and that “no other human activity poses this level of danger.” He gave no percentage. TechCrunch’s coverage and TechSpot’s both quote him making no numerical claim, and Anthropic “did not immediately return a request for comment” per TechCrunch.
Evan Hubinger — Anthropic’s alignment science lead, still at the company — replied on X the same day. His words, not Coxon’s: “Jacob is correct here – we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.” And: “I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”
That’s the joint. The Economics team’s highest-growth scenario names recursive self-improvement as its likely engine. Anthropic’s own alignment lead says the company building that engine doesn’t have a plan to align what comes out of it. Read the two documents in either order and you land in the same place: the upside case has a precondition, and the precondition is exactly the part nobody claims to have solved. That’s the structure. The number Hubinger attached to it is a separate matter, and I don’t buy it.
What the econ page is, and isn’t
It’s worth being precise about what the scenario explorer actually is, because it’s easy to read more certainty into it than it contains. It’s a task-bundle model: a structured way of asking what happens to output, wages, and labor’s share of income at three different automation intensities. The page itself carries the caveat: “Like every model, it is a stark simplification of a complex reality.” It assigns no probability to any of the three scenarios. The word “probability” doesn’t appear on the page. (A reader poll further down does show “around 10% of respondents have views in line with the extreme scenario” — that’s a self-selected visitor survey, not a model output, and conflating the two would be its own kind of sloppy reading.)
What the page does give you, concretely, if you take it as a planning input rather than a prophecy: in the Substantial scenario, wages for knowledge workers are essentially flat; in the Extreme scenario, they fall by more than 10% by 2030. Labor’s share of the economy runs 59.4% in Modest, 56.1% in Substantial (capital gaining 3.9 points), down to 45.2% in Extreme (capital gaining 14.8 points). Those are directional numbers about where value accrues if automation intensity keeps climbing — genuinely useful for thinking about pricing, hiring mix, and where margin comes from three years out.
None of that requires believing the Extreme scenario is likely, or even plausible. It requires believing automation intensity is a spectrum you’re somewhere on, and that “somewhere on it” is worth planning around. That’s a different claim than “AI could kill all humans, >10% within the decade” — and the page’s own writers keep those two claims apart, even in the same week Hubinger didn’t.
Why I don’t buy the number
I don’t have a p(doom) of my own, and I’m not going to manufacture one for this post. What I have is a problem with the specific number that got attached to the risk this week, and it’s a structural problem, not a disagreement about the underlying danger.
“>10% within the next decade” is unfalsifiable on any timeline that matters to a reader. Nobody collects on that bet in 2036 in a way that updates anyone’s prior — there’s no observation between now and then that confirms or denies it, only the outcome itself, which by construction happens at most once. And it’s self-serving in both directions at once: if you’re building the thing you’re warning about, a stated double-digit chance of catastrophe reads as candor and caution to one audience, and as proof the thing you’re building is important enough to bet a decade of civilization on to another. The number does work for the speaker whether the reader hears it as an alarm or as a pitch. I’ve made a version of this argument before, about Anthropic specifically not having separated its security case from its commercial interest — this is the week the two cases got published side by side, on the record, by the company itself, rather than inferred from the outside.
The number does work for the speaker whether the reader hears it as an alarm or as a pitch.
None of this is a claim that the underlying risk is zero, or that alignment is solved, or that Coxon or Hubinger are wrong to be worried. It’s narrower than that: a specific percentage, offered without a falsification path, attached to a decade-out event, from someone whose employer’s valuation benefits from the stakes reading as maximal — that number is not evidence I’d put weight on either way, and I’d say the same about a competitor’s engineer publishing the inverse claim with equal confidence.
What changes in a client’s security posture this week?
Nothing. That’s the honest, slightly unsatisfying answer, and it’s worth saying plainly rather than dressing it up.
I do fractional CTO and vCISO work for pre-Series-A AI startups, and the controls I’d tell a founder to have in place this week — model access scoping, data retention limits, who can push to prod, what gets logged, what a vendor’s own security posture looks like before you build on top of it — don’t move because a researcher resigned or because an alignment lead put a number on a decade-out risk in a reply thread. Those controls were never priced off p(doom). They were priced off the actual attack surface a seed-stage company has right now: too much standing access, no incident response plan, vendor dependencies nobody’s actually read the terms on. That surface doesn’t get bigger or smaller because Anthropic published two documents on the same day.
Where the econ page earns a place in a founder’s actual planning is the labor-share and wage numbers, used as a spectrum rather than a prophecy: if automation intensity in your market keeps climbing, what does that do to your hiring mix, your pricing, your margin structure over the next three years? That’s a real question with a real range of answers, and it’s the kind of question I’d rather spend a planning session on than a decade-out casualty estimate nobody can falsify. If you’re modeling team shape rather than headcount, I’ve argued the effect runs the other way: AI surfaces the backlog and teams staff up to absorb it. That’s the tension worth sitting with. The econ page’s Extreme case moves 14.8 points of income from labor to capital, and a growing team is not what that looks like.
The future is not predetermined
That’s the econ page’s own line, not mine: “The future is not predetermined. Ultimately, what the economy looks like in 2030 depends on many factors, like what AI can do, and how companies and workers choose to adopt it.” Read it the way Anthropic’s economists apparently meant it, as a hedge on their own model, and it’s a fair one; three scenarios, no probabilities attached, a disclaimer about stark simplification. Read it the way an operator should read it, and it’s not a hedge at all. It’s the whole argument. The two documents this company published this week don’t converge on a forecast. They converge on the fact that the outcome is still being decided by what gets built and how it gets adopted — which is the one part of this story a founder actually has a hand in, this week and every week after it.
What that means for your own stack and controls is a narrower question than what it means for the industry, and it’s the one actually worth answering this week. If you want it answered against your actual stack rather than the news cycle, that’s the work I do.