Where AI Belongs in a One-Person Content Operation
The agent pipeline that runs this blog, stage by stage: where AI is safe because its output is checkable, and where a human has to stay in the loop.
In this post
TL;DR: This blog runs on a pipeline of Claude Code subagents: one grades an idea against the live search results, one drafts against a brief, one reads the draft for invented facts and uncited claims, one checks the mechanics, and a test suite refuses to pass a post that breaks the site’s rules. It’s fast, and most of it is AI. The line I draw isn’t “AI writes, human edits.” It’s this: AI is safe wherever its output can be checked against something outside itself (a search results page, a source link, a test) and dangerous wherever it has to invent (facts, experience, client stories). So the two load-bearing parts of the pipeline are the fabrication gate and the fact that a human, not an agent, schedules every post. Everything else is plumbing. And the honest footnote: a fast pipeline solves production, not distribution. It will happily make more posts than anyone reads.
The question isn’t whether to use AI. It’s where.
Most arguments about AI and writing are stuck on a binary. One camp says generated content is slop and anyone using it is cheating. The other camp says the tools are good now and the holdouts are sentimental. Both camps are arguing about the wrong unit. “Did AI touch this post” isn’t a useful question. “Which steps did AI do, and what stopped it from being wrong at each one” is.
I run sublimecoding.com alone. There’s no editor, no content team, no second pair of eyes unless I build one. That constraint is why the pipeline exists, and it’s also why I’ve had to be precise about where the agents stop. When one person is the whole operation, a single invented fact doesn’t get caught by a colleague. It ships under my name.
Google’s position is more sensible than either camp. Its guidance on generative AI content doesn’t prohibit the tools. It warns that “using generative AI tools or other similar tools to generate many pages without adding value for users may violate Google’s spam policy on scaled content abuse,” and tells publishers to “focus on accuracy, quality, and relevance, especially when automatically generating the content.” Automation isn’t the problem. Unchecked automation at volume is.
So here’s the actual pipeline, and where I draw the line at each stage.
The pipeline, stage by stage
Every agent below is a markdown file in the repo’s
.claude/agents/ directory. That’s the standard Claude Code subagent
format: frontmatter with a name, a description, a tool allowlist,
and a model, then a system prompt. Each one runs in its own context
window and returns a report. The directory is gitignored, so these
definitions are local tooling rather than part of the published site,
but they’re real files, and what follows describes what they actually
say.
1. Idea grading: serp-gap-validator.
Before anyone writes a word, an agent searches the target query and one
or two close variants, classifies who holds page one (vendor docs,
practitioner blogs, thin or stale content), notes the dates, and greps
the site’s own published corpus for overlap. It returns a GREEN, YELLOW,
or RED verdict. RED means someone already owns the query, the space is
saturated, or the idea would split traffic with a post I’ve already
published. That last check matters more than people think: an idea can
have a wide-open external search page and still be a bad post because
I’d be competing with myself.
2. The brief: me. Title, slug, angle, date, tags, the internal links the post has to carry, and, critically, which real cases the post is allowed to use. The brief is where I decide what’s true and what’s mine to tell. An agent can’t make that call, because it doesn’t know what happened in a room it wasn’t in.
3. Drafting: post-drafter. It reads a
canonical exemplar post and the site’s voice contract, then writes
exactly one file. Its instructions carry hard gates (title length, tag
taxonomy, frontmatter shape) and content rules. The two that matter
most: never invent client engagements, incidents, or numbers, and reuse
only figures already published on the site. Every external stat gets an
inline source link, verified with a web search before it’s asserted. The
drafter doesn’t get to be clever about facts. It gets to be fast about
structure and prose.
4. Editorial review:
editorial-reviewer. This is the fabrication gate.
It’s read-only by instruction, its tool allowlist carries no Write or
Edit tool, and it’s pinned to Opus at high effort, a heavier tier than
the mechanical checkers get, because this is the one read where being
wrong costs the most. It checks, in priority order: any first-person
claim of a specific engagement, incident, metric, or dollar figure that
isn’t either linked to an external source or already present in the
published corpus; any stat or reported event without an inline citation;
voice drift against recent posts; and cannibalization against existing
titles and descriptions. It returns APPROVE or REVISE with numbered
findings, each carrying a severity, the offending quote, and a suggested
edit.
5. Mechanical preflight:
publish-preflight. Runs in parallel with the
editorial read, because the two don’t depend on each other. Also
read-only, on a cheaper model. It checks the rendered SERP title stays
within 70 characters once the ” | Jared Smith” suffix is added, that
tags are in the active taxonomy, that frontmatter is complete, that
every internal /blog/ link resolves to a published slug,
that images referenced actually exist, and that a scheduled date is full
ISO8601 UTC and plausibly in the future.
6. Scheduling: /schedule-post, run by a
human. This skill injects the reviewed draft into the site’s
content store with a future publish timestamp. The post is baked into
the release but hidden everywhere until that UTC moment, then flips
live, and a background rescan submits it to IndexNow within six hours.
The skill’s frontmatter sets
disable-model-invocation: true, which per Claude Code’s skills
documentation means “only you can invoke the skill,” the setting
Anthropic recommends for workflows with side effects “like
/commit, /deploy,” because “you don’t want
Claude deciding to deploy because your code looks ready.” The same skill
makes me assign the post to a topic hub and add one link to it from an
older post, because a post nobody links to is a leaf.
7. Social card build. After injection, a script generates the post’s Open Graph image from the content store. Mechanical, and it has to run after injection because that’s where it reads from.
8. mix test. The site is a Phoenix app,
and its test suite is the last gate. It fails the build if any post’s
rendered title exceeds 70 characters, if a post uses a tag slug that’s
been merged out of the taxonomy, if a post’s social card is missing on
disk, if a post’s FAQ schema answers don’t appear verbatim in the
visible body, or if a post more than 45 days old still has no editorial
inbound link from another post. Then I push to main, which deploys.
Where is AI safe in a content pipeline, and where isn’t it?
AI is safe at any stage where its output can be verified against something that exists outside the model, and unsafe at any stage where the model has to supply the truth itself. Here’s how it maps.
| Stage | What AI does | What a human must do | Why it’s safe (or not) |
|---|---|---|---|
| Idea grading | Searches the query, classifies page one, greps the corpus for overlap | Decide which ideas are worth grading, and which ones I’ve actually lived | Safe: the search results page is observable and the corpus is greppable |
| Brief | Nothing | Set the angle, the allowed real cases, the required links | Human-only: only I know what happened |
| Draft | Structure, prose, citation lookups, frontmatter | Supply the experience the post is built on | Conditional: safe against a brief, dangerous without one |
| Editorial review | Fabrication scan, uncited-claim scan, voice and overlap check | Read the findings and decide what to cut or fix | Safe: every finding points at a quote and a source to check |
| Preflight | Title length, tags, frontmatter, links, images, dates | Nothing, unless it fails | Safe: every check has a right answer |
| Scheduling | Runs the injection, builds the card, proposes the hub and backlink | Type the command, pick the date, approve the hub and backlink | Human-triggered: it publishes under my name |
| Tests | The suite runs; agents can fix what it flags | Push the deploy | Safe: pass or fail |
| Distribution | Can draft channel-native versions | Actually show up in the places readers are | The gap: production scales, attention doesn’t |
Look at which column the safe stages share. Every one of them produces output a second process can check: a SERP someone else can load, a slug that either exists in a JSON file or doesn’t, a character count, a test that goes red. When the agent is wrong, something outside it notices.
The drafting row is the one people get wrong, and it’s why I marked it conditional. A drafter working from a brief that names the real cases, with a rule that it may only reuse published numbers, is doing something checkable: every claim in its output should trace to the brief, a link, or the corpus. A drafter asked to “write a post about lessons from scaling an engineering team” is doing something uncheckable. It will produce a plausible anecdote, a plausible number, and a plausible lesson, and nothing in the pipeline can tell you they didn’t happen, because nothing in the pipeline was in the room.
AI is safe wherever its output can be checked against something outside itself. It’s dangerous wherever it has to invent.
This is the same distinction I make when I help a startup decide where agents belong in its product or its engineering workflow: find the stages where the output has an external referee, automate those hard, and keep a human on the stages where the model is the only source of truth. If you’re drawing that line for your own team and want a second opinion on where it should go, that’s a large part of what I do in fractional engagements.
What the fabrication gate actually catches
The argument for a separate editorial reviewer is easy to make in the abstract. Here’s what it looks like on a real post.
When the draft of my engineer’s ranking of AI side hustles went through the editorial reviewer, it came back with findings no rule in the drafter could have prevented. Two stand out.
The first was a bio line, and it came from me. The brief I handed the drafter described my background from an older, out-of-date version of my bio, written in a chat that didn’t have my site’s context. The drafter reproduced it faithfully, because that’s its job. The reviewer read it against the bio the rest of the site uses and flagged it as blocking. A reader wouldn’t have caught it. A search engine or an AI assistant assembling a picture of who I am from my own pages would get conflicting inputs.
The second was a cost figure. My post on what AI coding agents really cost reports agent consumption at “$1,800 to $3,500 per developer per month.” The draft dropped the unit: “my team’s seats consume $1,800 to $3,500 of API-priced tokens a month” reads as a team total. That changes the number by whatever the team size is, and it would have gone out under my name pointing at my own source post, which says something different. The drafter had the right number and the wrong unit, and the rule “reuse only published figures” doesn’t catch that on its own, because the figure was reused. It took a second reader going back to the source to see the unit had moved.
Both errors share a shape. Neither was an invention from nothing. Both were a real fact bent slightly in the retelling, which is the hardest kind of error for an author to see in their own draft. One came from me, one from the model. That’s why the gate is a separate agent with its own instructions, not a line in the drafter’s prompt telling it to be careful. The drafter already had that line. An author, human or model, is the worst reviewer of its own work, because it reads what it meant rather than what it wrote.
The reviewer is also read-only on purpose. It reports findings with suggested edits and never applies them. If the gate could rewrite the post, I’d have two authors and no reviewer.
Why a human schedules every post
The single mechanical decision I’d defend hardest is that no agent
can start a publish on its own. The /schedule-post skill is
marked so the model can’t invoke it; I have to type it.
The injection script is deterministic and fine, and it could technically still be run by hand from a shell; the guard is the skill setting plus a standing rule. Scheduling is the moment the post becomes a public claim by me, and I want that moment to require my attention. Every gate before it produces a report. Reports are easy to skim. Typing the command is the point where I’ve either read the editorial findings or I’ve consciously decided not to, and either way it’s my decision on the record, not an agent’s inference that the draft “looks ready.”
It also forces the curation work that automation quietly skips. Publishing can be fully automated. Deciding which topic hub a post belongs in, and which older post should link to it, can’t be done well by a script that doesn’t know how the corpus hangs together. The test suite enforces that inbound link eventually, with a grace period, because the failure mode of an automated pipeline is a growing pile of posts that nothing points to.
The mechanics of keeping agents inside those boundaries (which hooks block which edits, what a PostToolUse block can and can’t undo) are in the four Claude Code hooks I run on every project. The choice of which model runs which gate, and the tiering logic that puts review on a heavier model than mechanical checks, lives in one file I keep for model routing. Neither is a content concern, exactly, but a pipeline like this one sits on top of both.
What the checkable stages still need from me
Calling a stage “AI-safe” doesn’t mean it runs unattended. It means the errors it makes are catchable, which is a different thing.
The SERP grader can be wrong about who owns a query. It’s reading a snapshot, and search results move, which is why its instructions say notes older than about two months get re-validated before anyone drafts from them. The preflight can pass a post that’s mechanically perfect and says nothing. The test suite can go green on a post I shouldn’t have written. None of the checkable stages has an opinion about whether the post is worth reading, and none of them should. That judgment is mine.
What makes those stages safe is that when they’re wrong, the wrongness is visible: a verdict with a named page-one result I can open, a failing assertion with a line number. Compare that to the invention stages, where a wrong answer looks exactly like a right one until someone who was there reads it.
The other thing the pipeline needs from me is the raw material. The drafter is good at turning a real case into a readable section. It can’t produce the case. Every post in this pipeline starts with something I’ve actually done, configured, measured, or been wrong about, and the brief is where that enters. If I stopped supplying it, the pipeline would keep running and the posts would keep passing every gate, and they’d be worthless. The gates check that claims are sourced. They can’t check that there was anything worth claiming. Keeping that supply going through the months when nothing measurable happens is its own problem, and I wrote up what I’d do about it separately. What the pipeline doesn’t do is check whether any of it reached anyone. That’s the distribution audit, which grades what happens after publishing rather than the writing.
This is also why I keep the pipeline’s rules and my own working context in plain files rather than in an agent’s memory. The voice contract, the exemplar post, the model routing table, the correction history: they live somewhere I can read, diff, and version. I wrote up that approach in AIOS, the markdown operating system I run across every repo, and the plugins that encode the process disciplines underneath it are in the Claude Code plugin stack I actually run.
If you’re building your own
Start from the line, not the tools. List every step between “I have an idea” and “someone read it.” For each one, ask a single question: when the AI is wrong here, what notices? If the answer is a search page, a source link, a test, or a second agent with a different job and no write access, automate it. If the answer is “nothing, unless I happen to remember,” that step stays human, and no amount of prompt instruction to “be accurate” changes that.
Then put a separate reader on every factual claim, and make publishing something a model can’t do on its own. Those two decisions are small, and they’re the ones that let the rest of the pipeline be fast without the output quietly drifting from the truth.
And put distribution in the definition of done from day one. A fast production line will happily make more well-sourced posts than you’ve built an audience to read, which is the gap I graded my own blog on.
If you’re working out where agents belong in your own company’s workflows, the same framework applies well beyond blog posts. That’s the kind of problem I work through with founders as a fractional CTO and vCISO.