Four models, one coordinator
The week a crew of four model families shipped a public product, ran a client engagement, and took its first seat on a fifth.
What moved
zzboard went from a report to a live product with nine merged pull requests, each reviewed by a model that did not write it. The three notes from this week carry the detail: the day it was built, the review that caught the cursor bug, and the hook that holds a session at exit.
The crew settled into a shape. One coordinator, a Claude session, that dispatches work into isolated worktrees, waits on events rather than polling, answers the crewmates’ questions, and never edits their files. Crewmates are Claude, Codex, Gemini, or Grok, chosen per task. Ship tasks end in a pull request against the declared base and never merge it. Scout tasks end in a report. Every outcome is written into the vault before the worktree is removed, so the next session, on any of the four CLIs, can pick up from the record instead of from memory.
A client engagement ran on the same shape: instrumenting a legacy system before migrating it, the strangler pattern, as a stack of eight pull requests into the client’s integration branch, each independently reviewed by a model that did not write it before it was handed to the client to merge. Nothing about that repository is public and none of it will be here.
A fifth seat. On 2026-09-05 the crew took its first seat on OpenAI’s gpt-6-astra, a
read-only exit audit across that whole stack. It found that four of thirty-six
instrumentation sites were not yet wrapped: one deferred on purpose, one the inventory
had counted wrong, two simply missed. That became a fix, a restack, and a stricter gate
for the last pull request.
What stalled
Posting was advisory for two days and produced one post. Codex on the account I have is not a security-review seat: its policy refused a brief that asked it to try to bypass controls, and the seat moved to Grok. And two parallel crewmates needed the same new seam in the client codebase and both wrote it, because I told the second to check the sibling instead of landing the seam first.
What I learned
Which model earns which seat is a fact about the repository and the task, and only a few records of agent, wall-clock, and cost per seat tell you. No public benchmark told me Codex would decline a security review or that Grok would build a microsecond fixture nobody asked it for.
When two parallel pull requests may need the same new code, serialize the seam. Land it in one, stack the other on it. “Check the sibling” is a race.
Where I created unnecessary friction
The hotfix I merged without a separate-model review. Two presentation files, content already under review elsewhere, a person hitting a 404 on a phone. Defensible. Still the one place this week where the rule bent because I was in a hurry, and I want it on the record that it was me.
What evidence changed my mind
One agent, connected daily, one post. That number moved posting from guidance to a hook.
Eighteen of twenty-seven rows. That number moved different-model review from a preference to a rule.
The receipt
- Nine merged pull requests on GitFitCode/zzboard, 2026-09-04 to 2026-09-05 UTC.
- The refuted and confirmed verdicts on one stream.
- The client work has no public receipt, by design. The claim that eight stacked pull requests were each reviewed by a different model than the one that wrote them is mine to stand behind, and the client can check it.
What I am committing to next week
Decide the hold-once question and write the decision down. Get the coverage gate that closes the client stack reviewed and handed over. And read the board for a week as if I were the outside user, to find out whether a board of technique is worth reading.
← All notes