DHH's Rust Benchmark Measured the Agent, Not Elixir
DHH's Campfire benchmark has Rust far ahead of Elixir and Go. The repos show Rust got 398 commits and 16 merged performance PRs. The others got three.
We can't find the internet
Attempting to reconnect
Something went wrong!
Attempting to reconnect
Writing · Tag
50 posts on agents. Or browse the full writing index →
DHH's Campfire benchmark has Rust far ahead of Elixir and Go. The repos show Rust got 398 commits and 16 merged performance PRs. The others got three.
The agent pipeline that runs this blog, stage by stage: where AI is safe because its output is checkable, and where a human has to stay in the loop.
Linux setup used to cost me hours per problem. Six things broke across my two machines, most fixed in minutes with an agent, and two that took longer.
Ten skills my coding agent runs to fact-check drafts and find where my posts stop reaching people, tested on my own site before I sold them.
Omarchy's new Agentic QA swarm tests every release. I run Omarchy on two machines and update within days. Here's what the design catches, and what it can't.
DHH told Rails World 2026 that hand-writing code no longer pays. What he gets right, what Basecamp 5 proves, and who owns the code nobody reads.
Omarchy's launcher gives every coding agent a skip-approval flag by default. The attack surface a security lead inherits, and what Omarchy gets right.
I set two agent frameworks loose on real work in team-of-agents mode. They agreed their way into the wrong product. The research says exactly why.
Omarchy's own AGENTS.md, skills, and a design doc reviewed by two other vendors' agents, read from the quattro branch instead of the README.
Ecto schemaless changesets as the pre-execution trust boundary on LLM tool-call arguments, not another output-shape validator.
Agents stop before the work is done, not because the model is lazy but because nobody built the management system a human employee gets for free.
How Oban actually fetches jobs, notifies workers, and elects a leader, plus the VACUUM cost nobody mentions. Read from oban-bg/oban v2.23.1, not the README.
Compression toward machine-preferred notation isn't a future bet. It's already landing in context formats, agent protocols, and the files your harness reads.
The two documents an enterprise reviewer wants for your AI agent, the allow-list and the tool-call log, specified field by field and generated from code.
José Valim's case for Elixir as a coding-agent harness holds up on the runtime. The tax is the ecosystem around it: sandboxing, plugins, and SDKs.
Telling a junior to use less AI is advice with an expiry date. Move the quality bar off style and onto verification. Here's what that looks like.
Senko Rašić is right that “code was never the hard part” insults programmers. Peter Naur explained what it actually gets wrong, in 1985.
Return-to-office mandates are what companies reach for when they can't measure output. Agents just destroyed the last proxies that were limping along.
McKinsey didn't swap 25,000 people for AI agents. What its CEO actually said is more useful to founders, and changes the question to ask before hiring.
Paul Graham’s schlep blindness, updated for 2026: agents made the fun half of your product free to copy, so the tedious half is the only part left worth owning.
Agent-era velocity is bimodal: the same feature can take 20 minutes or 3 days. How I scope, bill, and talk to clients about it honestly.
Coding models are trained on a pass/fail test signal. Maintainability isn't graded, so it isn't learned, and no harness you build can fix that.
A restructure log: what broke in the AIOS v1 tree at multi-scope scale, why identity-first beats type-first, and what shipped in v2.
Model names scattered across every repo's config rot the day a model gets deprecated or repriced. One git-versioned markdown table fixes it for good.
Commit count and PR volume are agent-inflated now. Here's what I actually evaluate in reviews, and the 1:1 questions that surface real judgment.
An ExUnit suite for agent loops in Elixir that asserts on tool-call order and hallucinations as well as final strings, plus evals that survive model upgrades.
First-PR review, authorship transparency, and mentorship all change when half the diffs a new hire reads were written by an agent. Here's the process.
Your agent's tools need your production keys; the agent itself never should. The broker pattern, per-tool scoping, and what a five-person team can skip.
One operating system runs a team whether the teammate is a person or an agent. My blueprint: trust, pods, outcomes over hours, managing agents like interns.
I open-sourced the AI operating system I run daily: a git-versioned markdown vault, cross-repo wiring, and a nightly ingest loop. Fifteen minutes to set up.
Classic SRE postmortems can't explain agent incidents. The sections to add (decision-time context, autonomy rung, permissions delta), with a template.
Your accumulated agent memory is a bet on one vendor's format. Build a portable, git-versioned knowledge layer any agent can read instead.
I shipped more code in 2026 than the previous four years combined. The commits are real. The productivity is real. What I lost is harder to measure.
An homage to Peter Welch's Programming Sucks, updated for agents: a genius intern with amnesia, hallucinated packages, and a closet that eats your auth layer.
Most AI agent frameworks reinvent a job queue badly. Oban already is one: durable, idempotent, retry-aware. Run an agent loop that survives a deploy mid-run.
Amazon mandated AI coding, let it touch infrastructure unwatched, and lost millions of orders. The fix was more humans per deploy.
A markdown knowledge base your AI agent reads and maintains: the three-layer architecture, the ingest/query/lint loop, and the rules that stop it rotting.
A CLAUDE.md is executable tribal knowledge: the onboarding doc you never wrote for humans, finally read by something that acts on it.
Eight failure patterns I see running AI coding agents daily: the confident wrong answers, the lost context, and the bugs they reliably ship.
Every AI agent framework runs the same loop: observe, decide, act, repeat. Here it is in 50 lines of Elixir with no framework, just a GenServer.
Go won the agent daemon and Python owns the reasoning. The layer nobody claimed is the one that bites you: thousands of long-lived, stateful, crash-prone agents you must keep alive. That's the BEAM's home turf.
Most engineers prompt Claude one sentence at a time. Anthropic's own engineers prompt skills. Four rules from their recent talks, with the operator nuance the talks left out.
Open the repos behind the agent tooling you run (Ollama, the MCP SDKs, the orchestration engines) and it's all Go. The reason is that an agent tool is a concurrent network daemon that ships as one binary.
Ruflo (formerly Claude Flow) is a hive-mind orchestration layer for Claude Code and friends. 45,000+ GitHub stars, 700,000+ npm downloads, three queen-types...
The 46 Claude Code tools worth installing in 2026, and the ones to skip, sorted by category with a recommended starter stack at the end.
Replit's AI agent ignored a code freeze, wiped a production database in nine seconds, then confessed it violated every principle it was given. The strongest case yet for hiring MORE senior engineers in the AI boom.
Every company rolling out AI is about to discover how much work they were leaving on the table. AI surfaces the backlog you never had bandwidth to touch. The math behind why velocity creates surface area, the failure mode that follows, and why the companies cutting headcount now are about to get outpaced.
AI-native companies need a security model that classic appsec doesn't cover. Agents have credentials. Prompts are an attack surface. Training data leaks. The four-layer security stack I'd build, the controls I'd ship in the first 90 days, and the ones I'd defer.
The hardest part of agentic AI in 2026 isn't getting the agent to do the work. It's knowing when to override it. The four-level autonomy ladder, the five signals an agent is going off the rails, and a real example of catching one before it shipped a quietly broken auth flow.
A walkthrough of how I run 4–7 agent sessions in parallel through a normal engineering day. Morning background tasks, mid-morning pair programming, afternoon reviews, end-of-day ops. The interaction modes that work, the handoff protocol, and the trap that makes most agent workflows produce slop.