The Office Mandate Is a Measurement Failure

Return-to-office mandates are what companies reach for when they can't measure output. Agents just destroyed the last proxies that were limping along.

TL;DR: Nobody mandates presence when they can see output. The four-day mandate is a symptom of a measurement system that died, and the proxies it was built on — commits, pull requests, hours visible at a desk — were already bad before agents started producing most of the volume. Now they’re worse than useless, because the person doing the best work may be the one with the fewest keystrokes. The replacement isn’t a better activity metric. It’s the goal: did we ship what we set out to ship, and was the goal ambitious enough to be worth hitting. And every honest case for being in the same room — onboarding a junior, standing up a new project, responding to an incident — is a scheduled event with a start and an end date. Fund those out of the rent you stop paying.

The mandate is a symptom

Here’s the thing that gives it away: no one issues a mandate about the work.

They issue it about the location. Four days a week, badge in, be seen. Then everyone drives in, sits down, puts on noise-cancelling headphones, and joins the same video calls they would have joined from a kitchen table. The mandate produced attendance. It did not produce a single additional decision, shipped feature, or resolved incident.

That’s not an accident of implementation. It’s what the policy is actually for. A company that can look at an engineer and say “here’s what you delivered this quarter, here’s what it was worth” has no reason to care where the chair was. A company that can’t say that has exactly one signal left, and it’s a body in a building. The mandate isn’t a collaboration strategy. It’s an admission, written in real estate.

I want to be clear about my standing here before I go further: I’ve never issued a return-to-office mandate and I’ve never worked under one. Every company I’ve worked for has been remote, and several of them were very good at it. Some of those companies later failed or cut headcount — and in none of those cases was distributed work the reason. The market moved, or the model didn’t work, or the funding stopped. Nobody’s post-mortem said “we should have been in a room.” That’s an observation from where I’ve stood, not a study, and you should weigh it that way.

What I can speak to directly is the measurement problem, because it’s the thing I’ve had to solve on every engagement I run.

Agents destroyed the proxies

The old proxies were never good. Lines of code, commit count, pull requests merged, hours visible — every engineering leader who’s thought about it for ten minutes knows these measure typing, not value. We kept using them anyway, because they were cheap and they correlated with effort badly but nonzero.

Agents broke the correlation.

Start with the volume problem. In six months of building with agents, I put 4,154 commits into a codebase that grew to about 1.5 million lines — and I said in that post, and I’ll repeat here, that the line count includes generated code, vendor code, scaffolding, and configuration. It isn’t 1.5 million lines of artisan craft. It’s surface area. That’s the honest framing, and it’s exactly why the metric is now dangerous: the number went up by an order of magnitude and my hands did not get faster. If you’re ranking engineers by output volume in 2026, you’re ranking their tooling.

Then there’s the shape of the day, which is stranger and matters more. Running agents well is mostly waiting and steering. You set one going, it works, it comes back with a question, you answer it, you check whether it drifted, you send it back. Between those moments there’s real downtime — enough that you can hold several in flight at once, and enough that a lot of the supervision doesn’t need a desk at all. Boris Cherny, who created Claude Code, runs multiple sessions in parallel across terminal worktrees and cloud sessions and starts a batch of them from his phone in the morning, checking in through the day. He’s reported shipping dozens of pull requests a day this way and not hand-writing code at all in 2026.

Whatever you think of that as a way to work, notice what it does to the mandate’s premise. One of the most productive engineers in this industry does a meaningful share of his highest-leverage work from a phone, between other things. There is no desk-hours metric that captures him. There’s no badge reader that would tell you he was working. A supervisor watching the floor would conclude he’d checked out.

If your measurement system can’t distinguish the most productive engineer in the building from someone who’s checked out, the system is broken — and putting everyone in the building doesn’t fix it.

This is the part the office argument never survives. The mandate assumes presence is evidence. Agents made presence and evidence fully independent variables.

What I actually measure

On distributed and embedded engagements, the thing I ship is goals.

That’s the unit. Not velocity, not story points, not a dashboard of activity. We decide what we’re setting out to accomplish in a defined window, we say out loud what “done” looks like, and at the end we answer two questions:

Did we accomplish the thing we set out to do? Binary, or close to it. It shipped, or it shipped partially and here’s what’s left, or it didn’t and here’s what we learned. This is unambiguous in a way no activity metric ever is, and it’s the only question a client has ever actually cared about.

Was the goal ambitious enough? This is the one people skip, and it’s where the real signal lives. A team that hits every goal on time might be executing beautifully or might be setting targets it can clear without stretching. A team that misses might be sandbagged by something real or might be reaching correctly and learning fast. You cannot interpret the first question without the second. Two quarters of clean hits with no near-misses is not a sign of health — it’s a sign the goals are too easy, which is a management failure, not a team failure.

Underneath that, the smaller numbers still have a use, but only in aggregate and only as texture. Commit counts, client acceptance, escaped defects, cycle time — in aggregate, across a team, over a quarter, those paint a picture. What they don’t do is judge a person. A low number doesn’t mean unproductive. A high number doesn’t mean productive. I’ve seen the most valuable contribution in a month be a conversation that killed a feature nobody should have built, which shows up in every metric as a person who did nothing that month. Proving the return on the work is a different discipline than counting the activity, and it’s the one worth building.

This is also why the honest version of estimating agent-heavy work has gotten harder rather than easier — I wrote about what happens to client estimates when agents do the building, and the short version is that the throughput went up while the predictability didn’t. Measure the goal. The activity underneath it stopped meaning what it used to mean.

There’s a related failure mode worth naming: if you can’t measure output and you know it, the performance review becomes theater — a ritual for producing a rating rather than a mechanism for finding out what happened. The office mandate is that same instinct, expressed as a lease.

Presence is an event budget, not a lease

Now the part where I’ll argue against the strongest version of the other side, because there is one and it deserves better than a dismissal.

Some things genuinely go better in a room. I believe that. What I don’t believe is that any of them require a building you pay for twelve months a year.

Onboarding juniors. Real. Getting someone new up to speed is faster with a whiteboard, a shared table, and the ability to interrupt. It’s also bounded. That’s a week, maybe two, on location, at the start. It is not a permanent seating arrangement, and treating it as justification for one is a category error — you’re using a two-week need to buy a two-year lease. This matters more now, not less, because onboarding into an agent-heavy codebase is a different job than onboarding used to be, and the supervision that makes it work is about reading someone’s reasoning, not watching them type.

Incident response. Partly real, and usually stated wrong. What incident response needs is response distance — someone who can be hands-on inside an acceptable window. That’s a locality requirement, and for some systems it’s a legitimate one to write into a job description. It is not a desk requirement. Sitting in the office all day Tuesday does not make you faster on Thursday night’s page. And notice the arithmetic nobody does: if the person is in the office, they still have to get to the office. Thirty minutes, forty-five, an hour, whatever the commute is — the mandate doesn’t eliminate travel time, it just charges it to the employee every single day instead of to the incident.

Kickoff and planning. The most real of the three. The opening phase of a project, the architecture argument, the whiteboard session where the shape of the thing gets decided — those are better in person, and I’d fight for them. They’re also, again, events. A few days. Occasionally a week.

Add those up and you get something like: a couple of onboarding weeks per new hire, a kickoff per project, a few team gatherings a year. That’s a travel budget and an offsite budget. Put a number on it, then put it next to what an office costs — rent, power, water, internet, furniture, cleaning, parking, the facilities person, the whole daily carry, paid every month whether the room is used or not. The travel line is a rounding error against the lease line.

And here’s the part that should bother anyone who actually cares about culture: the travel version is better at the thing the office claims to do. A week working through hard problems together somewhere, with real meals and real evenings, produces more bonding than a year of people sitting near each other with headphones on. If the goal is a team that trusts each other, spend the money on that directly instead of buying a building and hoping proximity generates it as a side effect. Trust is the actual operating system, and you can’t lease it.

What’s left

Strip out the real-estate sunk cost and the middle-management layer whose primary function is confirming that people are at their desks, and what remains of the four-day mandate is a company saying, in the most expensive way available, that it doesn’t know what its people produce.

That was an embarrassing thing to admit in 2019. In 2026 it’s a strategic problem, because the gap between your best engineers and your average ones is now mediated by how well they direct agents — and that difference is completely invisible to attendance. You will promote the wrong people. You already might be.

The fix is not a policy about buildings. It’s writing down what you’re trying to accomplish, being honest about whether the target was ambitious enough, and getting comfortable judging results you didn’t watch happen. That’s harder than a badge reader. It’s also the only thing that works when the typing is done by something that doesn’t commute.

If you’re leading a team where agents are doing a growing share of the building and your existing measurement system is quietly falling apart underneath it, that’s the problem I help engineering leaders and founders fix. Let’s talk.