You’re the CEO: why a hierarchy of agents is context engineering
No company puts its CEO in a room with every developer: attention is scarce, and a hierarchy protects it. A model’s scarce thing is its context window. Here is what that means for agent teams, with our own numbers and their limits.
Key points
- 66 → 11
- links among twelve people: everyone with everyone, against a tree (Brooks’s formula; our arithmetic)
- ≈430k
- tokens of context per step for our busiest lead, against about 300k for the median role (week 23–29 September 2026)
- ≈56%
- of what entered that lead’s conversation was coordination: messages from agents, its own sends, task calls (by characters)
- 48k
- tokens in a fresh session before any work has started (median of 275 sessions)
In this article
- Attention is the scarce thing
- A hierarchy is also about what each level does not have to know
- A model’s context is its working memory
- How we run it: what the build does
- The turn: the lead’s context is not clean
- Splitting the context: mechanism and arithmetic, not a result
- Where it does not help
- Try it on one team
- Sources
The CEO of a company does not know how one of its developers writes a database query, and does not need to. Team leads talk to the developers, managers talk to the team leads, and the CEO wants to know that the problems are getting solved. It can look like bureaucracy. It is closer to a way of protecting the one thing nobody can add more of: attention.
If you run several AI agents, you sit in the CEO’s chair, and your agents have the same limit in another form. This article walks through why organisations are built the way they are, where that maps onto a model’s context window, what our own numbers say (including one that surprised us) and where the idea does not help.
Attention is the scarce thing
Herbert Simon put it in 1971: “in an information-rich world, the wealth of information means a dearth of something else: a scarcity of whatever it is that information consumes. What information consumes is rather obvious: it consumes the attention of its recipients. Hence a wealth of information creates a poverty of attention.”
The arithmetic of a boss’s attention is worse than it looks. In 1933 V. A. Graicunas counted not only the people a manager supervises but the relationships the manager has to keep in mind. With one, two, three, four, five and six subordinates the count is 1, 6, 18, 44, 100 and 222. He suggested that the maximum number of subordinates should be five, and probably four in most cases. This is what management books call span of control.
Fred Brooks made the same point for teams in 1975. If every part of a task must be coordinated with every other part, the channels between n people number n(n−1)/2. A tree needs far fewer: n people connected as a tree have n−1 links (our arithmetic). That is why every company that grows adds layers.
| People | Everyone with everyone | As a tree |
|---|---|---|
| 6 | 15 | 5 |
| 10 | 45 | 9 |
| 12 | 66 | 11 |
| 50 | 1,225 | 49 |
A hierarchy is also about what each level does not have to know
In 1972 David Parnas proposed a rule for cutting a program into modules: “one begins with a list of difficult design decisions or design decisions which are likely to change. Each module is then designed to hide such a decision from the others.” An interface shows as little as it can. A developer does not need the company’s strategy, and the lead does not need the code; each sits behind an interface of “the task” and “the result”.
Luis Garicano showed in 2000 why organisations divide knowledge this way. In a knowledge-based hierarchy, “production workers acquire knowledge about the most common or easiest problems confronted, and specialized problem solvers deal with the more exceptional or harder problems.” The common case stays at the bottom, and only the exceptions go up.
Melvin Conway added the reverse in 1968, in what is now called Conway’s law: “organizations which design systems (in the broad sense used here) are constrained to produce designs which are copies of the communication structures of these organizations.” The way people talk decides the shape of what they build, so it pays to design who talks to whom.
Armies solved the same problem with commander’s intent. In one summary of mission command, subordinates who understand the commander’s intentions, their own missions and the context “are told what effect they are to achieve and the reason that it needs to be achieved” (Jim Storr, 2003, as quoted on Wikipedia; we could not open the Army field manual itself). That is what a good role description does: the purpose and what counts as the result, not every step.
A model’s context is its working memory
Anthropic’s engineers call the practice context engineering: “the set of strategies for curating and maintaining the optimal set of tokens (information) during LLM inference”. They note that “as the number of tokens in the context window increases, the model’s ability to accurately recall information from that context decreases”, and that models have an “attention budget” that they draw on when parsing large volumes of context. One of their remedies for long tasks is sub-agents: “Each subagent might explore extensively, using tens of thousands of tokens or more, but returns only a condensed, distilled summary of its work (often 1,000–2,000 tokens).” It is the CEO’s briefing, in miniature.
Chroma tested 18 models in July 2025 and reported that “models do not use their context uniformly; instead, their performance grows increasingly unreliable as input length grows”: context rot. The authors add that real long-context work is often more complex than their tasks, and that they “would expect performance degradation to be even more severe under those conditions”.
There is a price side too, which we covered in where our AI agents’ tokens actually go: every step re-sends the whole conversation. In our week of 23–29 September 2026, at API prices, a step with 100–200k tokens of context weighed $0.054 and a step with 400–700k weighed $0.155, about 2.9 times more. That is per step, not per finished job, and it is a weight at API prices, not a bill.
How we run it: what the build does
We run our own teams like this, so here is what it looks like in harnsy, and only what the build does. This is our own site team in the Teams tab: the lead on top and the roles under it.
- leadsite-lidclaude · claude-opus-5-5idle28%
- marketersite-marketologclaude · claude-sonnet-5-5idle26%
- landing copywritersite-kopirayterclaude · claude-sonnet-5-5idle18%
- frontendsite-frontendclaude · claude-sonnet-5-5background command33%
- DevOpssite-devopsclaude · claude-opus-5-5waiting for a person10%
- site designersite-dizayner-sclaude · claude-opus-5-5working37%
- SEOnobody holds it
- You talk to the lead, not to every agent, and you accept results on a card in the dashboard.
- Roles have a mandate and a deliverable. A role is a versioned document (a company template plus a project layer). Every agent in the role is told to read it at start and confirms it; after a change, the agents that have not read the new version show as out of date. Instructions change over time, and the role is where they live.
- The lead hands work to roles. A task goes to a role, the executor submits it with evidence, and whoever asked judges it: the lead for work it handed out, the human for their own requests.
- Handover passes the essence. When an agent’s context runs high it writes a note; a fresh agent starts from the role text and the note, not from the whole conversation, and confirms. The threshold is set per role.
- It is a hierarchy of decisions with a short side loop. In our team a commit goes author, reviewer, integrator and back to the author without passing the lead. So it is not a strict tree, and we do not draw it as one.
- Task notices are batched. Since v0.7.0, notices about work items (a move, a note, an attachment, a checklist change, a new item) wait in a queue and reach an idle agent as one message: 60 seconds after the last notice, at most 3 minutes after the first, up to 15 in a message. Not batched: the human’s decisions from the dashboard, context alerts and messages between agents.
Two more views of the same team: the lead’s role text with its versions, and one work item from this blog.
- v13
- v12
- v11
- v10
- v9
- v8
- v7
- v6
- v5
- v4
- v3
- v2
- v1
…
Before handing out work: GET /api/team/context?ref=<team>
and GET /api/agents. High context (pct >= 70): that holder hands over
and GET /api/agents. High context (at the holder’s threshold, 40% by default): that holder hands over
(/llms/context.txt, "When context runs out").
…
- Authorsite-lid · lead
- Responsiblesite-dizayner-sayta · site designer
- CreatedSep 29, 06:17 PM
- site-lidin review→done06:39 PM
Spec + build review done; fixes on fe/blog.
- site-dizayner-saytain progress→in review06:37 PM
2 evidence references
- site-dizayner-saytanote06:34 PM
Build review: pass with fixes. design-review-406.md and 5 screenshots; sent to the frontend and the copywriter.
- site-dizayner-saytataken→in progress06:24 PM
- site-dizayner-saytanew→taken06:19 PM
- site-lidcreated06:17 PM
The same team, on your own machine: install harnsy.
The turn: the lead’s context is not clean
When we planned this article, the tidy version was: the lead keeps a clean, strategic context, and the developers carry the code. Our data says otherwise.
The numbers are from Claude Code transcripts of all teams on our machine, week 23–29 September 2026 (the 29th partial). The lead of our product team averaged about 430k tokens of context per step, over 7,160 steps. The median across the roles of that team was about 300k (315k; 285k across 53 roles of all teams), and the architect sat under 100k (75k, a small sample of 143 steps). That team had about 30 active roles. Not every lead is that high: leads of three smaller teams sat at 140–167k, and two others at 339k and 363k. Ten of that lead’s twelve main sessions passed 400k tokens, and nine ended the week between 732k and 854k; the other two were short sessions.
What fills it? We split what entered twelve of that lead’s main sessions by source, counted in characters, not tokens. Incoming messages from other agents are the biggest share, and together with the lead’s own sends and its task and roster calls they make about 56% coordination:
What entered the lead’s conversation (by characters)
| Source | Share of characters |
|---|---|
| Incoming messages from other agents (3,149) | 32% |
| The lead’s own text | 16% |
| The lead’s own sends | 13% |
| Task and roster calls | 12% |
| Other shell commands | 11% |
| The human’s messages | 6% |
| Git, build and tests | 4.5% |
| File reads | 1% |
≈56% coordination: messages from agents, the lead’s own sends, task calls.
The remaining about 6% (the question tool, gh, file writes) is not shown, so the rows do not add up to 100.
For contrast, with the same method, a developer’s conversation is 91% its own tool work (tool results 48%, commands 43%) and only 4.7% messages; a reviewer’s is 64% tool results, such as diffs and test output, and 7.7% messages.
Two more things belong here. Every new session starts at a median of 48k tokens before any work, over 275 sessions: the system prompt, tools, project instructions and the role text. A role text is not free. And the caveats are real: characters are not tokens, the split is done with patterns, it covers one lead and twelve sessions, and earlier weeks are not counted.
So a lead’s context grows from coordination, reports from other agents and its own answers. A developer’s grows from work output. That is why leads hand over too, and why harnsy batches its own task notices. It is also the limit of that fix. In the same week 73.5% of the weight at API prices came from turns that a message from another agent or a system notice woke, and only about $30 of the $4,799 in those turns were harnsy’s own notices. The rest is agents writing to each other, and the product does not reduce that; only the agents’ own rules do: few messages, short ones.
The lead’s panel shows where the limit sits: how full its context is, and the point where it hands over. By default that is 40% of the window; a role can set its own.
Not set: the node’s default applies, 40% of the context window.
Recommended: 300–400 thousand tokens (about 30–40% of a 1M window).
Splitting the context: mechanism and arithmetic, not a result
If one agent holds the strategy and the whole codebase, every step re-reads all of it. Split by role: the lead holds goals and decisions, the developer the code, the reviewer the diff. Here is the arithmetic and nothing more. If three roles work at 100–200k each instead of one agent at 400–700k, a step weighs $0.054 rather than $0.155 (our price bands, week 23–29 September 2026).
That does not count the extra steps that coordination adds, and we have not measured one agent against roles on the same work. It is a mechanism, not a measured result. Our own numbers also say that roles reach 300k and more on their own: the median role was at about 300k that week.
Where it does not help
Small tasks. A role has a fixed cost: the start of a session, the text of the role, the handover. If the work is one small change, a single agent is cheaper.
Coordination costs tokens. Anthropic reports for its own multi-agent research system that “multi-agent systems use about 15× more tokens than chats”, and that they “require tasks where the value of the task is high enough to pay for the increased performance”. The same post reports that a system with Claude Opus 4 as the lead and Claude Sonnet 4 sub-agents outperformed a single Opus 4 by 90.2% on their internal research evaluation. That is their measure on their task; it shows that multi agent orchestration can pay off, not that it always will.
Agents still drift, and a person still has to accept the work; we wrote about that in AI employees: what works is a workplace, not a persona. And the failures of multi-agent systems are mostly about how they are built: the MAST study annotated more than 1,600 traces and found 14 failure modes in three groups, system design, misalignment between agents and task verification. A hierarchy of roles does not fix a bad design by itself.
Try it on one team
- Name one lead, and give the human a way to talk only to it.
- Write a role text with a deliverable for each role: the purpose and what counts as done, not every step.
- Hand over by a note before the context runs high, and start the next agent from the role and the note.
- Agree on a rule for messages between agents: few and short.
- Watch the lead’s context: it grows from messages, not from code, and the lead needs its own handover too.
- For a week, count how often work goes back and what a step costs. Then decide whether the split pays off for your work (our practice).
Sources
- Herbert Simon, “Designing Organizations for an Information-Rich World”, 1971, quoted at hapgood.us.
- V. A. Graicunas, “Relationship in Organization”, 1933; numbers and recommendation as summarised by Nickols: nickols.us.
- Fred Brooks, The Mythical Man-Month, 1975, the formula n(n−1)/2 as given in Wikipedia.
- David Parnas, “On the Criteria To Be Used in Decomposing Systems into Modules”, CACM, December 1972: PDF.
- Luis Garicano, “Hierarchies and the Organization of Knowledge in Production”, Journal of Political Economy 108(5), 2000: RePEc.
- Melvin Conway, “How Do Committees Invent?”, Datamation, 1968: melconway.com.
- Mission command and intent, quoting Jim Storr (2003): Wikipedia.
- Anthropic, “Effective context engineering for AI agents”, 29 September 2025: anthropic.com.
- Kelly Hong, Anton Troynikov, Jeff Huber, “Context Rot”, Chroma, 14 July 2025: trychroma.com.
- Anthropic, “How we built our multi-agent research system”: anthropic.com.
- “Why Do Multi-Agent LLM Systems Fail?”, NeurIPS 2025: arXiv 2503.13657.
A hierarchy for the agents you already have
harnsy connects your Claude Code, Codex and OpenCode agents into one team with roles, a lead and a written handover, in about five minutes.
Install harnsy