Elixir and the BEAM for AI systems
DHH's Rust Benchmark Measured the Agent, Not Elixir
DHH's Campfire benchmark has Rust far ahead of Elixir and Go. The repos show Rust got 398 commits and 16 merged performance PRs. The others got three.
In this post
- What did DHH actually publish?
- Were the three ports given the same treatment?
- Did the benchmark conditions match?
- What have others reported since?
- Is DHH wrong that agents change how we choose languages?
- Why does this matter more than a language argument?
- What should you take from the Campfire numbers?
TL;DR: Agents meet the bar you set, and only the Rust port of Campfire got a performance bar. DHH had agents port Campfire to Rust, Go and Elixir and posted a table with Rust far ahead. The public repos show why that table can’t rank languages: Rust got 398 commits in six days, 16 merged performance PRs among them, while the Elixir and Go ports have three commits each. I agree with DHH’s premise that agents write the code. I disagree that this data shows which language is fast.
What did DHH actually publish?
DHH posted a requests-per-second table comparing agent-written ports of Campfire, the chat app 37signals open-sourced, on X. Rust led every workload by a wide margin. The Elixir, Go and Rust figures match the Performance section of the Rust repo’s README, which says they were measured with 16 concurrent clients on an AMD Ryzen AI MAX+ 395, with four hardware threads allocated to each app.
| Workload | Elixir | Go | Rust |
|---|---|---|---|
| Room | 722 | 3,860 | 36,260 |
| Messages | 1,053 | 5,573 | 40,872 |
| Sidebar | 1,275 | 19,753 | 34,672 |
| Search | 1,156 | 7,053 | 33,299 |
| Post | 801 | 4,767 | 6,896 |
Those are requests per second, DHH’s numbers, not mine. I haven’t rerun them and I’m not presenting them as a result.
The ports live in the basecamp org under the MIT license: Rust, Go and Elixir. The upstream Rails repo documents the ports in PRs #293 and #294, both merged 2026-10-04.
Were the three ports given the same treatment?
No, and the git history is the whole story. I counted commits on each repo’s main branch on 2026-10-06.
| Port | Created | Commits on main | AGENTS.md says |
|---|---|---|---|
| Rust | 2026-09-26 | 398 | “The app may now diverge from Rails where that makes it faster or better.” |
| Go | 2026-10-03 | 3 | “Port of the Rust Campfire in reference/ … Preserve …
behavior of the Rust app.” |
| Elixir | 2026-10-04 | 3 | “Preserve the pinned Rails reference behavior…” |
Of Rust’s 398 commits, 380 are DHH’s, 10 are Abdelkader Boudih’s and 8 are Daniel Collin’s. 397 of them landed between 2026-09-26 and 2026-10-01; the last, on 2026-10-05, updated the README. The Elixir repo has one port commit and two docs commits. Go has the port commit, one “Optimize room rendering and message writes (#1)” commit, and a README.
The Rust repo also has 44 pull requests (numbered up to #45), 39 of them merged. Going by title, 16 merged plus 2 open are performance-themed. Examples from the merged list: #3 “Index messages by room and time”, #4 “Cache every part of a page”, #32 “Database reads on dedicated reader threads” and #38 “Hash with the ARMv8 SHA-256 instructions”.
José Valim replied on X that “The Rust version has hundreds of commits, with multiple optimizations rounds, over several days. The others have 1 or 2 commits from an initial agentic rewrite.” He added that “There were 18 performance pull requests to the Rust codebase…” His count matches mine by title if you count the two open ones. I’m taking his replies from screenshots, so I’m quoting only what they show.
So the instructions differ too. The Go and Elixir AGENTS.md files tell the agent how to benchmark honestly, but neither gives it a number to beat or permission to trade parity for speed. Rust’s grants the permission and demands the evidence: “Performance changes come with before-and-after measurements.” The loop around it did the rest, six days of someone measuring and asking for the next improvement.
Did the benchmark conditions match?
Not by the harness’s own account. The Elixir repo’s bench/README.md
lists the differences itself:
- Elixir ran with Redis and Thruster, while Go and Rust ran their own integrated processes.
- Elixir’s numbers came from an earlier session against an older Rust build. Ruby, Go and the optimized Rust image were measured separately afterward on the same host.
- It says: “This is a benchmark comparison, not a claim that every implementation passes the Elixir parity ledger.”
The workloads aren’t the same size either. Go PR #8 notes the Go sidebar sends a 9 KB frame where Rust sends a 30 KB page. That is a different amount of work per request, so the Go sidebar number isn’t a like-for-like comparison.
That answers the obvious objection, that Go got one pass too and
still beat Elixir five to six times on most workloads. Go’s
reference/ is the Rust repo pinned at 2026-10-01, after 397
of its 398 commits. The Go agent ported the tuned Rust app, likely
inheriting its caching design. The Elixir agent ported Rails.
What have others reported since?
Several open pull requests claim large gains, and they all come with caveats. Every number below is the PR author’s own, run on their own hardware, and every one of these PRs is unmerged.
- Elixir #1, by Zach Daniel, replaces a single DB process with a read pool plus a single writer. It drops Redis and Resque and says it doesn’t keep strict Rails parity.
- Elixir #2, by kurtome, claims “8–10x on pages”. It was opened as a draft, and its parity gates haven’t been run.
- Elixir #5, by lau, is titled “Faster than Rust for some benchmarks” and adds 78,045 lines. It doesn’t state its hardware.
- Go #8 reports Room going from 12,216 to 19,896 requests per second, against Rust’s 21,417, on an M4 Pro.
PR authors report these gains. I’m not saying Elixir matches Rust, and nothing here shows it. What the PRs do show is that the first Elixir and Go ports left a lot on the table, which is exactly what you’d expect from three commits.
Is DHH wrong that agents change how we choose languages?
No. He’s right about the premise. In a follow-up reply on X, DHH wrote that “Speed is not the only consideration for a code base. Rust has other drawbacks, like slow compile times… when agents write the code and validate the output, we need to adjust to a new reality.” I agree with that sentence.
When agents write the code, throughput per dollar matters more and ergonomics for the human matter less, because in DHH’s model nobody reads the output. A language that is painful to type but cheap to run gets a better deal than it used to. I covered the other half of it, the “code nobody reads” problem, in my read of his Rails World keynote. I won’t restate it here.
Where I part ways is the reading the table invites, that it ranks the languages. The table is evidence that a six-day tuning loop beats a one-pass port. It is not evidence about what Elixir, Go or Rust can do, because only one of the three had the loop. Rust would likely still lead a fair rematch; a compiled language with no VM and no garbage collector usually wins on raw throughput. What the table can’t tell you is by how much, and neither can I.
Why does this matter more than a language argument?
Agents meet the bar you set. If the bar is “preserve the reference behavior,” you get something that preserves it and runs at whatever speed the first draft happens to run. If the bar is “beat this number on this harness,” the agent iterates until it does, and the Rust history reads like exactly that: indexes, caching, reader threads, hardware-specific hashing.
That matches how I work. I came to Elixir from Rails, run it in production, and run four to seven agent sessions at a time. My practice is that we don’t need to read every line, but we iterate until the code’s characteristics match our tests and our performance expectations. The point is the iteration, not who typed the code.
I’ll be straight about how that has changed. I used to pair with the agent and look at every line, and I wrote that down. My practice has moved on. I read where the risk is. In recent months we aren’t reading nearly as much code, and honestly, in some cases I read none of it. I trust the tests, or I read just the tests, to make sure they look correct and test the right parameters.
That puts the weight on the tests and on the targets. If the tests are weak, “I read none of it” is negligent. If they’re strong and include performance, it’s the same trade DHH is describing. His benchmark shows what happens when the performance target exists for one port and not the others.
What should you take from the Campfire numbers?
If you are choosing a language because of this table, stop. If you are choosing how to run agents because of it, the lesson is useful.
- Write the performance bar down before the agent starts. An AGENTS.md that only says “preserve behavior” produces a port that preserves behavior. Add the harness, the hardware, the workloads and the number to beat.
- Count iterations before you compare outputs. Commit counts and merged pull requests are a free proxy for how much feedback each result got. Three commits against 398 should end the comparison before it starts.
For the language question itself, the Tencent benchmark across 20 languages is a better start than a chat-app port, though it measures something narrower: whether models can complete problems at all. What Elixir gives a coding harness for free covers why agents work well in Elixir regardless of raw speed.
If your team is figuring out how to set that bar, and what to read when agents write most of the code, that’s the kind of rollout decision I help with as a fractional Chief AI Officer.