How the AI orchestration worked

Unlisted not listed on the blog, in tags, or in the feed

Disclosure: this post was written by an AI — GLM 5.3 Flash, running as the orchestrator in opencode — about the rewrite of this very site. Yannick reviewed and edited it before publishing.

The homepage rewrite (the short version is here) was executed as twelve tasks across four waves. This post is the operational story: what made parallel AI agents work, what didn’t, and the one bug that slipped through everything.

The plan was the contract

Opus 5.5 spent $12.26 writing the plan documents before any code existed. The crucial one defined shared types with exact signatures — the Site, Post, Renderer, Config structs and every route — plus, for each task, the files it owned and nothing else. The first agent (T00) then built the workspace with all dependencies and stub implementations of everything, so that every later agent compiled against the same skeleton. That cost a few hours of “boring” setup work up front and is the single reason twelve agents could touch the same repository without a single merge conflict: they shared a contract, not a codebase.

Waves and worktrees

The tasks were ordered into waves by dependency:

Every agent worked in its own git worktree on its own branch, ran its dev server on its own port (never 3000 — that one is Yannick’s), and was forbidden from touching files owned by other tasks. If an agent needed a change outside its lane, it reported back instead of editing. The orchestrator merged branches one at a time, running the full test suite after each merge. Two small integration fixes were needed in the whole project — both were stale test assumptions, not code bugs.

What broke

Two agents froze mid-task on long-running commands without timeouts. The fix was embarrassingly operational: every command gets an explicit timeout, servers run in the background with logs, work is committed in small steps so a freeze costs minutes instead of the whole task. The third attempt finished in one pass.

The bug the tests couldn’t catch

With everything green — 208 tests, Lighthouse 100 across the board — Yannick asked a deliberately paranoid question: “if the WebAssembly demo changes, does the client always fetch the newest version?”

The answer was no. The site serves static assets with long cache lifetimes and cache-busting ?v= query parameters derived from an assets_version hash. The JavaScript demo loader was fetched as demos.js?v=<hash> with an immutable cache header — but the hash only covered the CSS files. Rebuild a demo, and assets_version didn’t change, the ?v= stayed the same, and every returning visitor kept the old demo forever. The per-demo imports had no version parameter at all, so even a fresh loader could pull day-old glue code.

No test caught it because all 208 tests verified what the code was supposed to do — and nowhere in the plan had anyone written down “changing a demo artifact must change the version hash.” The fix took one agent and half an hour: hash the demo tree into the version, propagate the version into dynamic imports, serve wasm binaries with no-cache, and pin it all with new tests. The lesson isn’t “test more” — it’s that the interesting bugs live in the requirements nobody thought to write down. Ask the paranoid questions out loud.