Trust Is the Operating System
One operating system runs a team whether the teammate is a person or an agent. My blueprint: trust, pods, outcomes over hours, managing agents like interns.
TL;DR: People ask how I manage. One system runs the whole thing, whether the teammate is a person or an agent: trust. I extend it to you, you extend it to the rest of the pod, and everyone extends it to the mission. Hire for the trust you can afford to give. Scale by repeating the small unit, not mutating it. Manage outcomes, not hours. Protect the struggle that grows people. And run your agents through the exact same operating system — as an intern you shouldn’t trust off the bat, on a leash that widens as it earns it, except the agent doesn’t grow from the struggle. You do, by hardening the harness around it.
The membership test
I’ve run teams from three people to twenty, onshore and offshore, and every time someone asks “how do you manage,” they’re expecting a framework — OKRs, a ritual cadence, a tool stack. Those things exist and they matter, but they’re not the operating system. They’re artifacts of it.
The operating system is trust. It runs in three directions at once. I have to trust you with the domain knowledge and the calls I’m not qualified to make myself — that’s the entire reason I hired you instead of doing it myself. You have to trust the rest of the pod the same way, or you’re not a team, you’re a group of people who happen to share a standup. And everyone, including me, has to trust the mission enough to make the tradeoffs that a mission requires instead of optimizing for what looks good in the moment.
Strip away every other question about how to manage and you’re left with one: if I can’t trust you, why are you on the team? Not “why is your code good” — trust isn’t a proxy for skill, and it isn’t earned by credentials. It’s whether I believe you’ll make the call correctly when I’m not in the room, and whether you’ll tell me the truth when you don’t. Everything downstream of that — the pods, the goal-setting, the way I run agents — is just trust, operationalized at a different scale.
Hire for the trust you can extend
You can’t run a trust-based system on people you can’t trust yet, so hiring is where the operating system actually starts, not where it gets tested. I’ve written the long version of this — the case against optimizing for the whiteboard-polished, pedigree-heavy candidate over the person with real trajectory — in The Perfect Hire Is Killing Your Team. The short version for this piece: I hire for drive over polish. Hungry self-starters with an entrepreneurial streak, not people chasing the founder title — you want ownership without the exit plan. AI has made domain knowledge acquirable in a way it never was five years ago; what it hasn’t made acquirable is drive. That’s still the scarce input, and it’s the one thing you’re actually betting on when you extend trust to someone new.
The pyramid of pods
Here’s the part almost everyone skips: how you actually scale a team past the size where you can hold every relationship in your head.
The unit is a pod of three to eight people. Not a rule of thumb — a hard constraint on how many people one person can genuinely trust and be trusted by, bidirectionally, at the density that makes trust real instead of nominal. Past eight, you stop knowing who’s actually stuck and start reading status updates instead of people.
So when the team outgrows one pod, you don’t stretch the pod. You repeat it. A second pod, a third, each one the same size, each one running the same trust relationships internally. The layer above them isn’t one manager stretched across twenty reports — it’s a team of pod leads who cross-talk across the boundaries that would otherwise calcify into silos: front-end talking to back-end talking to mobile, before a ticket forces the conversation. That cross-talk is itself a trust relationship, at a different layer, running the same operating system.
This is the actual answer to “how do you scale from a handful of engineers to fifteen or twenty” — you don’t design a bigger structure, you fractal the small one and put people you trust in charge of the boundary. I lived through most of this the hard way — the hires I got right, the ones I got wrong, the systems I wish I’d built sooner — and wrote it up in From One Engineer to Fifteen. And if you’re still at the size where the pod is the whole team — three to five people — the mechanics of running that without becoming a full-time manager are in How to Manage a 4-Person Engineering Team Without Becoming a Manager. Same operating system, every layer. Trust is the invariant; headcount is just how many times you’ve repeated the unit.
Outcomes over hours, not surveillance dressed as accountability
If trust is real, it shows up first in what you measure. I run goal-based, not hours-based: quarterly OKRs broken into weekly chunks that actually sum to the quarter, so a Friday check-in tells you something true about whether the outcome is on track — not whether someone was active in Slack between nine and five.
When a chunk gets missed, I go in with curiosity, not blame. A two-day task not done in a week means there’s an issue — but “there’s an issue” is a question, not a verdict, and how you ask it is most of the job. It’s not what you say, it’s how you say it. Most misses are a signal about scope, a blocker nobody flagged, or a wrong estimate — not a character problem. Treat it like a character problem the first time and you’ve taught the whole pod to hide the next miss instead of surfacing it, which is strictly worse for you.
What I don’t do is instrument the distrust: no keyloggers, no screenshot tools, no activity trackers pretending to be productivity tools. You hired smart, capable people specifically because you can’t do their job yourself — surveilling them is a tell that you don’t actually believe that, and they will notice. Professionalism is owning the outcome, not performing busyness for a dashboard, and a surveillance culture optimizes for exactly the wrong signal.
This default — trust the pod to recover its own bad call, don’t hover — has a limit. If there’s no consensus and the pod is spinning in circles arguing instead of deciding, I step in. Decision by indecision is a recipe for bad; someone has to break the tie, and sometimes that’s me. Same with a deadline that’s genuinely at risk: the default is hands-off, but hands-off isn’t the same as absent. If a pod is going to miss something that matters, I’ll add bandwidth to save it — including my own hands, which I’ll get to below.
There’s a cost to getting the trust-not-surveillance default wrong, and it isn’t hypothetical — it’s a security posture, not just a management one. If your team is afraid to tell you about the near-miss, you don’t have a security program, you have a countdown. What that looks like when an agent — not a person — is the one making the miss is its own postmortem discipline, which I cover in What an AI Agent Postmortem Should Contain.
And when the person too afraid to flag the miss is a human, not an agent, the failure isn’t technical at all: a scared team is your biggest attack surface, because fear suppresses the exact honest report your security depends on.
Don’t steal the learning
The flip side of trusting people with outcomes is trusting them with the struggle that makes them better at producing outcomes. People grow by working through something hard, not by being handed the answer or having someone quietly clean up behind them. Every time you jump in to save someone from a fixable mistake, you’ve taken the rep away from them — and taught them, without meaning to, that they don’t have to clean their own mess, because you will.
This is getting harder to hold onto in an AI-accelerated shop, because the temptation to just let the model produce the answer and skip the struggle entirely is constant. I’ve written about the specific failure mode where AI closes the knowing-vs-doing gap for juniors before they’ve built the judgment that used to come from doing the grunt work themselves, in We’re About to Stop Making Senior Engineers, and about where Satya Nadella’s “token capital” framing gets the small-team version of this right and wrong, in Nadella Is Right About AI and the Firm. Mostly. The short version: protect the struggle on purpose. It’s the only part of growing an engineer that doesn’t scale, so it’s the part you have to defend deliberately.
The same operating system runs your agents
This is the part of the system nobody else is writing about honestly, and it’s where I want to spend the rest of this.
An agent is the intern you shouldn’t trust off the bat. That’s not a metaphor I’m reaching for — it’s the literal correct posture. It onboards the same way a hire does: you write it a CLAUDE.md the way you’d write onboarding docs for a new engineer, and I’ve made the case that a CLAUDE.md is exactly that document — executable tribal knowledge, not config — in Your CLAUDE.md Is the Onboarding Doc You Never Wrote. Where the analogy holds all the way through: the agent starts on a short, permanent-feeling leash and earns wider scope only as it demonstrates it deserves it, exactly like you’d ladder a new hire from read access to write access to production access over their first quarter, not their first day.
An agent is the intern you shouldn’t trust off the bat.
Where the analogy stops holding is the reason you have to be more careful with the agent, not less: it’s nondeterministic and it fails faster than a person ever could. A junior engineer who’s about to make a bad call gives you tells — hesitation, a question in Slack, a draft PR sitting open for review. An agent doesn’t hesitate. It drives the car into the tree faster than a human ever could, with full confidence the whole way there. The framework I use for calibrating how much autonomy an agent has actually earned — and the signals that it’s drifting off the rails before it hits anything — is in When to Trust an Agent and When to Step In. Amazon found this out at a scale that cost real orders when a mandated coding agent got unsupervised access to infrastructure it wasn’t ready for — the fix wasn’t less AI, it was more humans per deploy, which is the same lesson at enterprise scale: Amazon Let the AI Drive. It Hit a Tree. Guardrails here aren’t a vote of distrust in the tool. They’re physics. You’d put the same guardrails on a new hire’s prod access, just slower ones, because a human’s worst mistake in an hour is smaller than an agent’s worst mistake in a minute.
Here’s the polarity flip that took me longest to actually internalize: with a person, you protect the struggle because that’s how they grow. With an agent, there’s no struggle to protect, because there’s no growth to protect it for — the model on Tuesday isn’t a wiser version of the model that made Monday’s mistake. So the learning doesn’t happen in the worker. It has to happen in the manager, and it shows up as system hardening instead of personal growth: you tighten the CLAUDE.md, you add the guardrail the postmortem revealed you needed, you harden the harness. Every agent mistake is a rep for you, not for it. I’ve written about what actually earns a place in that document after real mistakes, not a speculative template copied once and forgotten, in What I Put in CLAUDE.md After 50 Commits With It — and what it looks like to run this discipline hard enough to actually ship at volume, in 4,154 Commits in Six Months With AI Agents.
That reframe changes the actual question you should be asking before you widen an agent’s scope. It’s not “do I trust this agent” — you don’t, categorically, the same way you don’t extend blind trust to a new hire in week one. The question is: does it matter if the agent nukes the database, if you can restore it in seconds? An agent deleting a production database is a real, documented failure mode, not a hypothetical — I’ve written about the incident and the actual lesson from it, which was to hire more senior engineers, not fewer, in An AI Just Deleted a Production Database in Nine Seconds. Hire More Engineers. The self-repair capability that made that incident recoverable in minutes instead of days is probably just chaos engineering under a different name — you’re not trusting the agent more, you’re driving the cost of its worst plausible mistake toward zero, so the blast radius stops being the thing that scares you.
Do that consistently and something else follows almost automatically: every engineer on your team becomes, in effect, a manager of agent-reports. That’s not a new discipline bolted onto the job — it’s one more layer of the same pyramid of pods I described above, with the same rules: earn trust in small increments, widen the blast radius slowly and cautiously as it compounds, protect yourself from the failure mode you haven’t guardrailed yet instead of the one you already have. The operating system doesn’t change when the teammate is a model instead of a person. Only the growth target does — it moves from the worker to the system around the worker.
Close: no work is beneath the leader
None of this works if trust only flows downward. I’ll file files. I’ll do the menial cleanup task nobody wants, the off-hours bug triage, whatever it takes to hit a deadline the pod is genuinely at risk of missing — because “not my problem” isn’t a sentence a leader gets to use, and because the trust I’m asking the team to extend to me has to be earned the same way I’m asking them to earn mine. When Lavender needed engineering power to hit a deadline, I stepped in and added hands, not just direction. The outcome belongs to the whole team, and that has to be true in both directions or it’s not actually trust, it’s a hierarchy wearing trust’s clothes.
That’s the whole system, top to bottom and person to agent: extend the trust you can afford to extend, in the smallest unit that lets you extend it honestly, measured by outcomes instead of hours, protecting the struggle where it produces growth and hardening the harness where it doesn’t. Everything else — the tools, the rituals, the org chart — is downstream of whether that trust is real.
If you’re building or rebuilding a team under this model — sizing the first pods, figuring out where the agent layer actually earns its autonomy, or just want a second set of eyes on the org you’ve got — that’s a conversation worth having.