Nobody Got Replaced. Agents Got Added.

McKinsey didn't swap 25,000 people for AI agents. What its CEO actually said is more useful to founders — and changes the question to ask before hiring.

TL;DR: The claim making the rounds — that McKinsey replaced 25,000 of its 60,000 employees with AI agents — is not what happened. On the HBR IdeaCast in January, Bob Sternfels put the firm at 60,000 total: about 40,000 humans plus roughly 20,000 agents, a number he revised upward to about 25,000 at CES days later. The agents were added on top. The human headcount did fall — from around 45,000 at the end of 2023 to about 40,000 — but that took two years and McKinsey attributes it to attrition and performance management. Meanwhile the buried number is the one founders should care about: client-facing roles are up about 25%. This isn’t a story about replacement. It’s a story about reallocation, and the question it should provoke in your next planning meeting isn’t “how many people can I cut” — it’s “how many agents can one of my people actually supervise before the quality drops.”

What was actually said

The viral version is clean and wrong: 60,000 employees, 25,000 replaced by AI. It’s a good number. It fits in a headline, it confirms something people already half-believe, and it takes about four minutes to check.

Here’s the actual quote, from Sternfels on the IdeaCast on January 6:

“I now update this almost every month, but my latest answer to you would be 60,000, but it’s 40,000 humans and 20,000 agents.”

Sixty thousand is humans plus agents. It’s a headcount figure that includes software. The 40,000 humans didn’t get reduced to arrive at that number — the 20,000 agents got stacked on top of them to produce it. A day later at CES, Sternfels put the agent count closer to 25,000, and a McKinsey spokesperson confirmed the higher figure as the current one. Eighteen months before that, the firm was running a few thousand.

So where does a real reduction show up? It does exist, and it’s worth stating precisely because the honest version is still significant: McKinsey went from roughly 45,000 people at the end of 2023 to about 40,000 — a drop of more than 10% over about eighteen months, the largest in the firm’s history. The firm rejects the layoff framing, attributing the decline to natural attrition plus its normal performance-management process, alongside a 2023 restructuring that eliminated around 1,400 back-office roles.

That’s about 5,000 people over two years, during a consulting downturn that followed five years of near two-thirds headcount growth. Real, consequential, and roughly a fifth the size of the number that went viral — and driven substantially by a market cycle, not a model swap.

The number they cut out is the one that matters

Here’s what bothers me about the popular retelling, and it’s not just that the figure was wrong. It’s that the correction usually stops at “actually, nobody got replaced,” which leaves the most useful finding on the floor.

At CES, Sternfels described the shift as “25 squared.” Client-facing consulting roles up about 25%. Non-client-facing roles down about 25%. And output from that shrinking non-client-facing group up about 10%.

Read that again with a founder’s eyes. The part that got amplified was the shrinking group. The part that got dropped was that the client-facing side grew by a quarter. This is not a firm getting smaller. It’s a firm moving people from the back office to the front and using agents to hold the back-office output up while it does.

And then the sentence that should actually rattle anyone building a company, which Sternfels said in the same breath:

“Our model has always been synonymous that growth only occurs with total head count growth. Now it’s actually splitting.”

The head of a firm whose entire business model was billable humans just said the link between headcount and growth is coming apart at his own company. That’s the finding. It survived the fact-check, it’s on the record, and it’s more disruptive than the fake version — because the fake version says “fire people,” which is a one-time event, and the real version says “your unit of capacity changed,” which is a permanent structural fact you now have to plan around.

The fabricated number told founders to cut. The real number tells them to reallocate — and reallocation is a much harder thing to get right.

Why the wrong number traveled

It’s worth being honest about why a claim like this spreads, because the mechanism will produce the next one too, and you’ll be the target of that one as well.

Numbers about AI displacement are not neutral facts moving through the world on their own merit. They’re capital-allocation instruments. Whether they’re true is a secondary property.

Consider the incentives. A large company that has committed enormous sums to hardware, data centers, and model access needs that spend to look like it’s working. There’s a shape this takes that I’ve written about before: AI made tokens cheap and it’s making hardware costly, and the companies deepest into that spend are the ones with the strongest reason to announce that it’s paying off. Announcing a large headcount reduction attributed to AI does several things at once — it signals cost discipline to the market, it justifies the capital expenditure, and it positions the firm as ahead. Whether the reduction was actually caused by AI or by a demand slump that would have happened anyway is not a distinction the press release is built to make.

Now add the investor layer. Venture capital is structurally biased toward sweeping change — the entire return model depends on finding the thing that resets an industry, so a claim that an entire labor category has been automated is exactly the kind of story that attracts capital toward the companies telling it. That’s not a conspiracy, it’s just what the incentive gradient looks like. A dramatic claim gets amplified because amplification serves the amplifier.

So when a number like this crosses my feed, my procedure is short and it’s mostly one question, asked before anything else: what does the person saying this need to be true?

That’s the same test I ran on a vendor’s 96% security benchmark and on the way Anthropic’s safety asks were bundled, and it’s the test that would have caught this one immediately — because the claim was being repeated most enthusiastically by people selling AI transformation to executives. After that question, the rest is mechanical: find the primary source, not the article about the source. Read the actual transcript or watch the actual talk. Check whether the eye-catching number is a total, a delta, or a projection, because those get swapped constantly. And check the timeline — “replaced 25,000 people” and “grew to 25,000 agents over two years” describe completely different events, and the difference lives entirely in a verb.

Four minutes. It’s not investigative journalism. It’s just not taking the headline’s word for it. And it matters, because the last time everyone repeated a confident AI story without checking the mechanics, the story was Amazon’s, and the details were considerably less flattering than the summary.

Addition is the pattern that works

Now the part I actually believe, which is the reason the McKinsey numbers are interesting rather than just misreported.

The addition pattern is the correct one. Not because it’s gentler, but because it’s the one that compounds.

The company that trains its existing staff to direct a growing number of agents ends up with people who can each carry vastly more scope. The company that cuts staff and hands the remaining work to agents ends up with fewer people, each of whom is now responsible for reviewing more output than they can actually review, and no one left with the context to catch the things that go wrong. One of these is a capability build. The other is a cost-cutting exercise wearing a technology costume, and I think it’s short-sighted in a way that will be expensive to unwind. The teams that cut deepest will be hiring back in eighteen months, at worse terms, for people who now have to reconstruct institutional knowledge that walked out the door.

There’s a second-order problem underneath this that I keep coming back to: agent-heavy development increases your blast radius and your velocity simultaneously. You can do more things, faster, in more places. That means more needs reviewing, not less — and the reviewing is the part that requires judgment, context, and someone who will be there when it breaks. I’ve argued that we’re about to stop producing senior engineers precisely because the pipeline that made them ran through the work agents now do. Cutting the humans who do the reviewing, to pay for the agents who generate the things needing review, is the specific move that breaks this.

The question to actually put in the board deck

So here’s the replacement for “how many people can AI let us cut.”

What’s your human-to-agent supervision ratio, and what’s your evidence for it?

I’ll give you my numbers, from running this daily, so you have something concrete to argue with.

For menial work — organizing files, summarizing notes, mechanical scans, the grunt tier — I’ll run five to ten agents at once without much strain. The work is checkable at a glance and the failure modes are boring.

For actual engineering, writing and changing code that has to work, it’s five to six at a time. Sometimes up to ten, depending on how good the test coverage is on what they’re touching — strong tests raise the ceiling because the tests do part of the supervision for me. Push past that and the quality doesn’t degrade gracefully, it degrades in a specific way: I stop reading the diffs properly and start skimming them, and skimming a diff is functionally the same as not reviewing it. The agents don’t get worse. My attention does.

It’s a spectrum, and it moves. Some days each agent needs more babysitting and the number drops. Some tasks are well-fenced enough that it climbs. Anyone who gives you a fixed universal ratio is selling something.

For an outside data point: Cherny’s setup at Anthropic runs roughly five terminal sessions against separate worktrees plus five to ten cloud sessions, with sub-agents fanning out underneath. Different tooling, different scale of delegation, but the number of things one experienced person actively steers lands in a strikingly similar place.

Now hold that against 25,000 agents and 40,000 humans and do the arithmetic. That ratio is well under one agent per person, which is exactly why McKinsey’s version is coherent — Sternfels has framed the near-term goal as every employee being enabled by at least one agent. It’s an augmentation ratio, not a replacement ratio. If your plan involves anything like five or ten agents per remaining human, you are not proposing the McKinsey strategy. You’re proposing something nobody has demonstrated, and the binding constraint won’t be model capability. It’ll be how many diffs a tired person will actually read on a Thursday afternoon.

That constraint is real, and it’s the one worth designing around. You can push it — better tests, tighter task scoping, agents that review other agents, structures inside your agent fleet that mirror the org structures companies already use for humans. I think that’s where this goes, and I think one person eventually supervises far more than six. But every one of those structures is itself something a human has to build, maintain, and debug. The supervision doesn’t disappear. It changes shape and moves up a level.

The thing that doesn’t scale

One more, and it’s the piece I think gets left out of every one of these headcount conversations.

An agent has no motivation. It doesn’t need to feed anybody. It isn’t trying to make rent, or get the promotion, or avoid being the person who broke production in front of the whole team. It has no stake in whether the company exists next year. It does the work it was handed because it was told to, and that’s the entire depth of it.

Every reason anything gets done in your company is a human reason. Someone cares about being good at this. Someone doesn’t want to let their team down. Someone has ambitions that require this project to succeed. That layer is not a soft benefit sitting on top of the real work — it’s the thing that generates the direction the agents then execute. You can add 25,000 agents and get more throughput. You cannot add 25,000 agents and get more intent, and intent is the scarcer input.

That’s a bigger argument than fits at the end of this post, and I’ll make it properly on its own. But it belongs in the frame here, because the headcount question is usually posed as if people and agents were the same kind of thing in different quantities. They aren’t. One of them supplies the reason.

Before your next hire decision

Three things, concretely.

Check the number that’s driving the decision. Open the primary source. Ask what the person saying it needs to be true. If a board deck’s thesis rests on a statistic somebody screenshotted, that’s not a thesis.

State your supervision ratio explicitly and defend it with something. Not a vendor’s claim — your own observation of when your people stop reading diffs carefully. If nobody on your team can tell you that number, you don’t yet know your actual capacity, and you certainly shouldn’t be sizing headcount against it.

And ask Sternfels’s question rather than the fake one. Not “how many people can we replace,” but “which roles move toward the customer, and what holds up the work they’re moving away from.” That’s the split he was describing, and it’s a reallocation problem — which is harder than a cutting problem, and considerably more likely to still be working in two years. I’ve made a version of this argument before about what AI actually does to the size of the firm, and everything since has pushed me further in the same direction.

If you’re a founder trying to size a team against agent-heavy delivery — and trying to work out which of the numbers landing in your inbox are load-bearing and which are marketing — that’s exactly the kind of question I help founders get right before it’s baked into a plan. Let’s talk.