techdad.io
← Tutorials

Tutorials

The Agent Company: How We Run a Team of Claude Agents

Running several AI agents on one codebase is a coordination problem, not a capability one — the five rules we pay for it with, and how to write your own operating doc in an afternoon.

LESSON11 min readSep 14, 2026human-reviewed
On this page

What changed

When the first version of Apex Prompt shipped, the team behind it was small: six roles under one coordinator, and one rule holding the whole arrangement together — one agent, one domain. That was enough to ship a working product in a day.

As of this writing, the project instructions file that ran that first build has grown into something longer: a table of eleven roles under the same coordinator, a set of playbooks each written after something specific went wrong, and a clause about what a long AI session costs to run. None of that table was planned in advance. Every row got added because something broke and the fix needed writing down somewhere a new session would read it.

That’s what this piece is: the five coordination rules the file carries today, most of them written the day something broke — and, because a rule you only read isn’t a rule you’ve used — a build-along thread running underneath them. By the end you’ll have the start of your own version: a one-page doc naming your roles, what each one owns, and which decisions come back to you.

Rule one: one agent, one domain

Say two agents split work by “backend” and “API.” That sounds like two lanes. Then someone reports that login is broken. The login endpoint is backend work and API work at once, so both agents open the same file; both ship a fix that is correct on its own terms; and you find out they disagree at merge, which is the worst moment to find out.

The rule that avoids it: one agent owns one domain, and anything spanning two domains gets split into two separate pieces of work before either agent starts — never handed to both at once. Draw the boundary by what a role owns, not by a label that sounds like a boundary; “backend” and “API” feel separate, and a login endpoint doesn’t respect the label.

We run two developer seats now, not one, and they carry an identical charter. The safety comes from something else: whoever’s coordinating gives them disjoint lanes and never puts both in the same repository and branch at the same time.

So the expensive part was never adding the second seat. The expensive part is the discipline of never letting two of them reach for the same file — and that discipline lives with whoever assigns the work.

Build-along 1: write the roster

From here, “you” is you, building your own version; “we” is still us.

  1. Open a blank doc and list the domains your project has — one row per domain — your project probably has fewer than you think.
  2. For each row, write what it owns, specific enough that two different people reading it would draw the same boundary around it.
  3. For each row, write what it must not touch. This is the column doing the important work — “owns the database schema” says nothing about whether that role can also edit the code that reads it; “must not touch application code” does.

If a row’s must-not-touch column comes up empty, that’s not a sign the role is unconstrained. It’s a sign it hasn’t collided with anything yet.

Rule two: the ticket is the interface

Two agent sessions working on the same project don’t share a memory. When one finishes and another picks up, the second one only knows what’s written down — never what the first one was thinking, only what it left behind. So the hand-off between sessions has to go through a file, not a conversation: we don’t resume a session past the task it was started for, on purpose, because resuming rewrites its whole context at a cost worth avoiding (more on exactly how much, below).

A conversation can’t carry a hand-off well even when you try. It has no fixed shape, no guaranteed place for the thing that actually matters, and no way to tell the next reader “this part is already settled, don’t relitigate it.” A written task forces three things onto the page every time: where the binding facts come from, so the reader isn’t guessing at sources; what’s explicitly out of scope, so a reviewer can say “that shouldn’t be in this diff” without arguing about intent; and a finish condition specific enough that someone who wasn’t in the room can check it.

The template, generic enough to hand you as-is:

MARKDOWN
# TICKET-<NNN>: <short title>

- **Project:** <project slug>
- **Assigned to:** <employee name>
- **Created by:** <coordinator, for whoever is asking>
- **Status:** OPEN | IN_PROGRESS | DONE
- **Depends on:** <ticket id(s), or none>

## Task
<One clear task for one employee. What needs doing.>

## Context
<What the employee needs to know. Link/reference prior tickets or specs.>
- **Binding playbooks:** <named here — the executor reads these before starting; the DoD below enumerates their live clauses as arms>

## Definition of done
<How we'll know it's complete, checkable by someone who wasn't in the room.>

---
## Output
<Employee appends their result here, then moves the file to DONE/. Keep it to the DoD arms discharged, one line each, plus a short evidence index — measurements and logs live in an evidence folder, not here.>

Nothing exotic. What makes it work is that a task’s scope ends up being exactly what its finish condition can test — write a vague one and the scope quietly grows to whatever the agent takes it to mean; write a checkable one and there’s nothing left to argue about once the work’s done.

Rule three: the rules arrive after the failure

Every playbook in our project file exists because a promotion merged and something about it was worth writing down. The ritual is simple: after each promotion lands, a short retro names what worked, what hurt, and one to three concrete changes — and the changes get applied in the same session, not filed for later. We’ve written it into our own process file in exactly these words: “a retro that only records is a failure of the ritual.”

Here’s one rule that arrived this way. At one point, a cheaper model handled a small fix — one predicate, one line — on protected behavior. The fix looked fine, and the existing tests passed, because those tests under-asserted what they were supposed to check: the fix had in fact broken the behavior they were meant to protect. Catching and correcting it cost two extra rounds of review. The rule since: never drop model tier for anything touching protected behavior, money paths, or shared application-level code, however small the diff looks going in.

A rule like that reads differently than one invented at a whiteboard before anything went wrong: nobody has to be convinced of a hypothetical, because the cost was already paid once.

Rules like it don’t live in the main project file forever, either. Today that’s a handful of separate documents: one about writing tickets well, one about promoting code safely, one about reading logs and evidence without fooling yourself. Each one loads only when a task touches what it governs, rather than getting read into every session whether it’s relevant or not.

Rule four: the bill is re-reading, not writing

Here’s the part that isn’t intuitive until you’ve seen the numbers: a long-running AI session doesn’t get expensive because it writes a lot. It gets expensive because every single turn re-reads its entire context from the start, and that re-read is billed every time. Context size multiplied by turn count is the actual cost; the words the agent produces are a rounding error next to it.

We measure this now instead of guessing at it. Two rows from that measurement, the heaviest lane on each of two separate days:

Session Turns Context
Heaviest lane, 2026-09-05 1,107 874K tokens
Heaviest lane, 2026-09-09 1,033 835K tokens

Both numbers are well past the point where “fine” is a reasonable word for the size — and the second has an explanation: that 1,033-turn lane turned out to be two separate tasks running in one session that should have been two sessions, and it was still holding 606 command results in its context, totalling 2.6 MB.

The mechanism behind them: the models we use compact their own context automatically, but only once it reaches roughly 967,000 tokens — so a session we’d once have compacted at 200,000 tokens instead grows toward a million and pays that near-million on every one of its next thousand turns. Name what turned out not to be the cause, because it’s the obvious suspect: instruction files, memory files, and the task file itself. A session’s first turn — before any of that growth — runs 15,000 to 31,000 tokens depending on the role. The bloat comes almost entirely from later turns, not the starting ones.

The correction that makes this honest: an earlier retro reported one day’s agent runs as “under 700k tokens.” That figure was accurate as far as it went, and it was also the wrong number to trust: it counted only the fresh tokens written that day, not the ones re-read from cache. The same day’s sessions replayed roughly three billion. In our own words, recorded plainly rather than rounded down: “Since 2026-09-09 alone: 4,133M cache-read + 277M cache-create tokens against 0.1M fresh input and 6.6M output.” Cache reads are cheaper per token than fresh input, but they still count against a usage limit, and at that ratio the gap between “under 700k” and “roughly three billion” is the whole story, not a rounding error in it.

Four changes came out of that, each one you can apply to a single agent on one repo, not just to a whole team of them. Two of them are settings, shaped like this:

PLAINTEXT
CLAUDE_CODE_AUTO_COMPACT_WINDOW=200000   # compact well below the model's own ~967K default
maxTurns: <N>                             # hard stop per session; restart fresh from the task file, don't continue

Both live in files you commit with the project — the first as an environment setting, the second on the agent’s own definition — so every session starts with them already applied.

The turn cap is the one worth tuning rather than copying flat — ours varies by role, from 100 for a session that writes prose or reviews text up to 350 for one that’s building and testing code, because the right stop point depends on how much back-and-forth that kind of work needs.

The other two are habits: never resume a session past the task it was started for — a resume rewrites the whole context at roughly twelve times the cost of reading it once — and keep long command output out of the conversation, writing it to a file and reading back only the part you need instead of letting a thousand-line log sit in context turn after turn.

We expect the compaction cap alone to cut replay by something like three-and-a-half to four times for the same amount of work. That’s a prediction we’re still checking, not a result we’re claiming yet.

Rule five: some calls stay with a person

One more line for your operating doc. Add a sentence naming the decisions that come back to a human — yours, specifically — no matter how confident the agent doing the work is. Ours, stated as a standing rule rather than a case-by-case judgment call: anything irreversible, and any promotion of code to what real users see, goes to a person, and the person makes the call.

Write yours as a single sentence you could hand to a stranger and have them apply correctly without asking you first — “X and Y always come back to me” is enough. If you can’t finish that sentence yet, that’s fine; it usually means you haven’t hit the failure that would tell you what belongs in it, which is exactly the shape of rule three.

Your operating doc now has three rows of roster and one line of escalation.

What you can copy, and what’s ours

Copy: the operating doc itself, the one-agent-one-domain rule, the must-not-touch column, hand-over-by-file instead of by conversation, a compaction cap, a turn cap, and the retro ritual that turns a failure into a rule the same day it happens.

Watch, don’t copy: our roster size, our specific set of playbooks, every token figure above, and our escalation line — that one encodes who our human is, not who yours should be. That escalation rule held under pressure once already, weeks after we wrote it down, and we wrote about it separately: it’s the closing story here.

Five rules. Two of them are now sitting in a file you wrote. The next piece in this thread is about the file itself — what actually makes a task an agent can finish without you in the room.

Get the next build in your inbox.

New build logs and primers land most weeks. No spam, unsubscribe anytime.

Get notified when it's live.