AI engineering

We Gated the Code and Left the Prose Wide Open

Teams built real review gates around agent-written code. The PR descriptions, design docs, and incident summaries the same agent writes pass through nothing.

In this post
  1. The gate stops at the diff
  2. A wrong description doesn’t go unread. It steers the review.
  3. The rest of the surface nobody’s gating
  4. What prose actually carries
  5. The qualifiers fall off even when you’re careful
  6. “Just spot the AI writing” isn’t a plan
  7. The line that actually resolves it
  8. What I actually run against this
  9. The Monday version
  10. Where this actually breaks
  11. The gate you already trust, one layer up

TL;DR: Every gate a team builds for agent-written code (review, CI, tests, linters, security scans) sits on the diff. The same agent, in the same run, writes the PR description above it, and nothing checks that at all. Worse than unreviewed: a wrong description sets the frame the reviewer reads the diff in, so it buys a worse review than no description would have. Spotting AI writing by its tells doesn’t fix this: no single tell is proof, and untrained readers are most wrong exactly when they feel most sure. The reason the gates grew on one side of that line and not the other is simpler than neglect: a diff is checkable against a test suite, and a PR description is checkable against nothing. Add one cheap, specific check per artifact class — a description that states what the diff doesn’t do, a design doc someone else has to defend out loud — and the asymmetry closes without adding a tool.

The gate stops at the diff

Watch what happens to one pull request from an agent, end to end. The diff gets a reviewer’s attention. It runs through CI. A test suite has to go green. A linter checks style. On a security-conscious team, a scanner looks for the obvious classes of vulnerability. That’s a real gate, built over years, and even on teams where review is slipping under agent-era volume the machinery around the diff is still there: the CI config, the required checks, the branch protection rule. Whatever else erodes, something still runs against the code. What isn’t graded doesn’t get learned, which is what happened to code quality itself when the training signal never rewarded maintainability, and the review discipline teams have built since is the direct response to that gap.

Above that diff sits a paragraph or two of prose: the PR description. Same agent. Same session. Same run that produced the code the gate just processed. And it passes through nothing. A human reads it once, decides from it whether the diff is worth opening, and moves on. No test suite runs against a PR description. No linter flags a claim that doesn’t hold up. Nobody scans it for the equivalent of a SQL injection. It is the least-verified artifact in the entire pipeline, and it’s the first thing anyone reads.

This is not a new observation in isolation. Rachelle Rathbone, a senior engineer at Atlassian, wrote it up as: “AI Made You Faster. It Didn’t Make You Better.” The point worth building on: the damage isn’t confined to the code. It shows up in the PR descriptions, the Confluence pages, the Slack messages — the places where one person tells another what they did and what to look at. I want to credit that plainly, because it’s correct and it’s the right place to start. What I want to add is narrower, and it’s mine rather than Rathbone’s: the code side has gates and the prose side has none. Note that the piece isn’t letting the code off the hook — it cites GitClear and CodeRabbit to argue that AI-assisted code is measurably worse, and writes that we’re “producing more code, faster” and “also producing worse code, faster.” My claim isn’t that the code is fine. It’s that the code at least gets checked, and the prose doesn’t.

A wrong description doesn’t go unread. It steers the review.

Here’s the part that makes this worse than a missing gate: an unreviewed PR description isn’t neutral. It’s actively worse than no description at all, because of how review actually works.

A reviewer opens a PR that says “refactors the retry logic to handle the timeout edge case.” They read that sentence first, form a hypothesis about what they’re about to see, and then read the diff against that hypothesis. That’s not laziness — it’s how reading code under time pressure works for everyone, including me. You don’t independently re-derive what a diff does from nothing every time; you use the claim as a lens and check whether the code confirms it.

If the claim is right, this is efficient. If the claim is wrong — if the agent’s own understanding of what it just built drifted from what the diff actually does, which is exactly the kind of error I catalogue from my own agent sessions every week — the reviewer isn’t reading the code anymore. They’re confirming a story. The description didn’t fail to add information. It replaced the reviewer’s attention with confirmation bias, on purpose, without anyone deciding that should happen. A wrong description buys you a worse review than an absent one, because an absent description at least leaves the reviewer looking at the diff with no frame to confirm.

That’s the asymmetry: the artifact with zero verification is also the artifact that sets the terms every other, better-verified artifact gets judged by.

There’s a correlational finding that matches this shape, qualifiers intact rather than stripped off the way these numbers usually circulate: a CHI 2025 study out of Microsoft Research and CMU, surveying 319 knowledge workers pre-screened as weekly-or-more GenAI users, found higher self-reported confidence in GenAI’s output was associated with less self-reported critical thinking on that task, while higher confidence in one’s own ability was associated with more (Lee et al., CHI 2025). Not a claim that AI output causes people to think less — correlational, both sides self-report, sample of habitual users rather than a general population. But it lines up with the mechanism above: trust in a claim substitutes for scrutiny of what’s underneath it.

The rest of the surface nobody’s gating

PR descriptions are the clean case because they’re paired directly against a diff that does get checked, which makes the contrast visible. But they’re not the only unguarded artifact coming out of the same sessions. Design docs and RFCs get written by the same agents and reviewed, if at all, by someone skimming for tone rather than checking claims against the system. Incident summaries get drafted from logs an agent read, and the summary is often the only thing anyone reads after the incident is closed — the postmortem becomes institutional memory whether or not it’s accurate. Slack updates and release notes are lower stakes individually, but they’re the same shape: prose generated in the same run as verified work, carrying none of that work’s verification with it.

One artifact in that list has its own answer, which I’ve written up separately in what earns a place in CLAUDE.md after fifty commits. The point for this post is narrower: the unguarded surface is wider than PR descriptions, and it’s wide for the same reason in every case.

What prose actually carries

The reason this lands harder on prose than on the code isn’t incidental. A team doesn’t transfer its understanding of a system through the diff; the diff is the artifact, not the explanation. But I want to be careful here, because I’ve argued the stronger version of this elsewhere and it cuts against the easy claim: following Peter Naur’s 1985 paper, in code was never the hard part, I made the case that you cannot fully write a theory down — documentation is, in Naur’s words, “an auxiliary, secondary product,” and it will never carry the theory. I still think that. So prose isn’t the theory-transfer layer, and I’m not going to promote it to one.

What prose carries is the inputs to the theory: why this approach, what it assumes, what it deliberately doesn’t handle. Constraints are the recoverable part. The theory itself rides in the heads of people who spent real time on the real problem, which is why I ended that post on continuity rather than on documentation. That’s a smaller claim, and it makes the problem worse rather than better — the one recoverable trace of a decision is now routinely generated by something that never held the theory in the first place, and nothing checks it on the way out.

A data point on what it looks like when that transfer fails even for the person who did the writing: in the MIT Media Lab’s “Your Brain on ChatGPT” preprint (Kosmyna et al., still watermarked “under review,” not peer-reviewed), participants wrote an essay with ChatGPT’s help, then were asked, in that first of four sessions, to quote a sentence from what they’d just written. Fifteen of the eighteen participants in the ChatGPT group couldn’t. In the search-engine group and the no-tools group, same session, it was two of eighteen each (Kosmyna et al.) — Session 1, one condition, eighteen people, not a claim about the whole 54-person sample. But it measures the thing underneath the failure a design-doc gate is built to catch: the person credited with the writing never built the theory in the first place. If the author can’t reconstruct a sentence they produced minutes earlier, there was never anything in their head for the prose to be a lossy projection of — and a reviewer reading that prose is downstream of nothing.

The qualifiers fall off even when you’re careful

The clearest example I have is one I’d rather point at than assert, and it’s sitting in the piece that sent me down this road — a careful writer, citing a real study, in an article arguing for more care.

Rathbone cites the same MIT preprint I did, and renders it as: “The study also found that 83% of the AI-assisted group could not accurately recall or quote from essays they had just written.” That’s not fabricated. Fifteen of eighteen is 83.3%, and it’s scoped to the AI-assisted group, which is more care than most coverage of that paper took. They also flag it as a preprint with a small sample and tell you to hold it lightly, which is more than most of its coverage did. What’s gone is narrower and more load-bearing than the preprint caveat: it’s the first of four sessions, and the two control groups sat at two of eighteen on the same question. Without the controls, a reader can’t tell whether 83% is alarming or just what happens when you ask anyone to quote themselves from memory. The controls are what make the finding land, and they’re what fell off.

That’s not a lapse in care. It’s what happens to any sentence that nothing checks — because “does this claim survive checking against the thing it cites?” is a question prose never gets asked, anywhere, by anyone’s process. The argument of this post, demonstrated by its own source — worth more than an example I could have built myself.

“Just spot the AI writing” isn’t a plan

You can learn this. Almost nobody has, and that’s the problem. The instinctive response to “the prose is unreviewed” is “teach people to spot the AI-written parts and read those more carefully.” That doesn’t work, and I want to disagree with Rathbone on this specific point, on the record, because it matters for what actually fixes this.

That post calls em dashes “the #1 dead giveaway,” and as writer-side advice — strip them before you send — that’s fine, and I’d say the same. The claim I want to push on is the broader one running underneath it: “Everyone can tell.” That gets scoped in the piece to the unedited paste, which is the easy case, and the scoping is right. The version that circulates, and the version a reviewer actually acts on, is the unscoped one. Mostly, we can’t.

I run a skill in my own tooling built specifically around AI-writing tells, and its operating rule is blunt: no single tell is proof. Density is the signal, not presence. One em dash is nothing; a paragraph that’s nothing but rhythmic tricolons and hedged transitions is worth a second look. And underneath that rule is something sharper than “people are bad at this.” In a study of 254 Czech native speakers sorting GPT-4o-generated text from human-written text across a range of registers (peer-reviewed, unlike the MIT preprint above), the group that got no feedback made its worst errors precisely when confidence was highest — the signal ran backwards. The same study found that detection is genuinely learnable: participants who got immediate feedback after every trial improved both their accuracy and their calibration.

Read those together and the conclusion isn’t “nobody can tell.” It’s worse for our purposes. Detection is a trainable skill that almost no reviewer has been trained in, and in its untrained state the confidence signal points the wrong way. “Learn to spot the AI writing” asks reviewers to lean on exactly the instinct that misfires hardest — on a text type that study didn’t test. Its stimuli were multi-register prose, not templated technical writing. A PR description is two or three sentences, which doesn’t give you enough text to measure density in — and density, by the rule above, is the only signal there is.

So tell-hunting isn’t a plan. It’s a habit that feels like diligence and isn’t. If the fix has to work regardless of who’s reading and how careful they’re feeling that day, it can’t depend on detection. It has to be a gate — something structural that runs whether or not the reviewer is sharp that morning.

The line that actually resolves it

I drew a line about where AI belongs in a workflow last week, writing up this blog’s own content pipeline, and the line I drew there is the one that explains this problem rather than describing it:

AI is safe wherever its output can be checked against something outside itself. It’s dangerous wherever it has to invent.

That’s from where AI belongs in a one-person content operation, and it applies here without modification. A diff is checkable against something outside itself: a test suite, a linter’s ruleset, a reviewer who can run the code. A PR description is checkable against nothing outside itself — there’s no oracle that returns pass or fail on “this description accurately characterizes this diff,” so nobody built one, because you can’t gate what you can’t check.

The asymmetry isn’t a case of a team getting lazy about documentation. Gates grew wherever a cheap, fast, reliable check existed to hang them on, and nowhere else — the same shape as why maintainability never got rewarded during model training even though everyone agreed it mattered: there was no fast oracle for it either. Nobody skipped reviewing PR descriptions out of neglect. There was nothing to check it against, so nothing to build the gate out of.

What I actually run against this

Since a description genuinely can’t be checked against an external oracle the way a diff can, the practical answer is the closest approximations available, run consistently — not results I’m claiming numbers for, just the mechanism:

The ai-tells grep pass, plus a read pass. The grep catches the mechanical density signal — rhythmic patterns, stacked hedges, tell density worth a second look. It doesn’t decide anything alone; it flags candidates for the read pass, where a human actually checks the claim against the diff or the system it describes.

The editorial-reviewer agent’s fabrication scan. This is the gate I run on every post on this site before publishing, and the principle transfers directly: it has no Write or Edit tools at all, and its whole job is finding any specific claim — an engagement, a number, an incident, an intent — that isn’t traceable to something outside the document. Applied to a PR description: does this claim about what the diff does trace to the diff, or did the agent assert it?

The plan-recon rule that a brief cites only what it has opened. Before a design doc or plan gets acted on, every path, symbol, and claim in it should trace back to a file someone actually read at authoring time, not inferred or remembered from a similar system. Unverified stays tagged unverified instead of stated as fact — the prose-side equivalent of not merging code the tests haven’t run against.

The Monday version

None of this requires new tooling to start. It requires one cheap, specific gate per artifact class, applied consistently, starting this week.

PR descriptions: the author states, in their own words, what the diff does not do. This is the one an agent can’t fake by rereading its own diff more carefully, because answering it requires knowing where the boundary of the change sits — not what’s in the diff, but what a reasonable reader might assume is in the diff and isn’t. That’s a claim the agent that wrote the code has no particular reason to get right, and a human stating it has to actually think about the edges of the change. Rathbone’s version of this test is “if you can’t explain your output without re-reading it, you didn’t write it.” Mine is narrower and harder to fake, because stating a boundary requires knowing where it is.

Design docs: someone who didn’t write it presents it out loud. If they can’t defend a decision when asked, the doc goes back before it becomes the record anyone builds on. This doesn’t require the presenter to have written the code — it requires them to understand it well enough to answer for it, which is a different and cheaper bar than a full independent review.

Incident summaries: every claim in the timeline cites a log line or a commit — the evidentiary half of what an agent postmortem should actually contain. Not “the service degraded around 2pm” — the specific log line that shows it, with a timestamp. If a claim in the summary can’t point at the evidence, it doesn’t go in the summary, because a postmortem nobody can re-derive from evidence is just a story that got agreed on.

Anything else — release notes, Slack updates, status reports: name an owner who will be asked about it out loud, in a standup or a review, at some point. What does the work here is not the writing of the artifact at all, but the fact that someone specific knows they might have to answer for it.

None of these is a tool. Each one is a sentence of policy that converts an unverifiable claim into a checkable one — by making a human commit to it in a way the agent that generated it never had to. That’s the same move the code side of this made years ago, just applied one layer up, to the prose a team actually runs on.

Where this actually breaks

I want to be specific about what these four gates don’t fix, because overclaiming here would repeat the failure this post is about. None of them are as cheap as “add a linter rule”: they cost a human’s attention at the moment the artifact is produced, against a real calendar, on a team already drowning in review load, on the same Faros AI telemetry that shows PR merges with no review at all up 31.3%, correlational, as I said plainly when I wrote it up. And none of them scale the way a test suite does: a suite runs against ten thousand PRs a day at the same marginal cost as one, but “someone defends the doc out loud” doesn’t parallelize. If your agents are producing design docs faster than your team can sit through presentations of them, that’s the real constraint — the same production-versus-attention gap that shows up everywhere agents multiply output without multiplying the humans absorbing it.

The gate you already trust, one layer up

The instinct to gate agent-written code came from a correct read of the risk: agents produce a lot of output, fast, and not all of it is right. That instinct remained correct past the edge of the diff; it simply stopped being applied there, because nobody had built the oracle that would have made the next place to apply it obvious.

The fix isn’t more scrutiny generally, and it isn’t learning to spot the tells — that’s a trainable skill nobody has trained for, and one where feeling certain is the wrong signal to trust. It’s the same move that built the code gates in the first place: find the artifact, find (or manufacture) something outside the artifact that can check it, and make the check unfakeable by whatever produced the claim. A PR description that states what the diff doesn’t do passes that test. A vibe that the writing “reads fine” doesn’t.

If you’re building this out for your own team — where the gates already are, where they should be, and what a cheap version actually looks like for your review load — that’s a conversation I have often as a fractional CTO and vCISO.