Where our AI agents’ tokens actually go: a week of 65,000 requests
If you run several Claude Code, Codex or OpenCode agents at once, you probably reached the weekly limit sooner than you expected. We wanted to know why, so we counted. This article is not about money: we run on subscriptions, and the question is how the size of the context weighs on your usage limits. (Anthropic’s docs say a long session keeps using the limit for the whole conversation; how much a token weighs, they do not say.) This is what one week of transcripts from one person’s agents says about what Claude Code costs and why usage limits run out: 65,487 requests, one machine, no theory.
Key numbers
- 65,487
- requests and 20.8 billion input tokens in the week
- $6,527
- is the weight of the week’s tokens at API prices, about $932 a day
- 65%
- of the weight is re-reading the conversation from cache ($4,245); model output is 12% ($804)
- 99%
- of input is served from cache
- 3.4×
- heavier a step at 400–700k tokens of context than under 100k
In this article
- How we counted
- Claude Code usage limits start with one fact: every step re-sends the conversation
- Context length sets the weight of a step: the Claude Code context window in numbers
- What wakes an agent up
- Idle time, the cache and the Claude Code usage limits
- Tool output and screenshots stay in the context
- The model matters less than it seems
- What we are changing at home (a hypothesis)
- Count it yourself
- What to take from this
- Where harnsy fits
This article is for anyone who uses Claude Code for real work. Most of it holds for one agent as much as for ten, and the parts about several agents say so. Near the end we describe what we are changing in our own setup and where harnsy, the tool we build, fits in. That part is short and marked.
How we counted
We read the transcripts Claude Code keeps for every session: one line per request, each with its token counts. We priced every request at the published API rate for its model and for each kind of token (input, cache write, cache read, output). Then we grouped the requests three ways: by line of the weight, by the size of the context at that step, and by what started the turn. A turn is one wake-up of an agent, made of one or more requests. Each request is priced at its own model’s standard rate: Opus 5.5 $6,147 (63,096 requests), Opus 5 $278 (23–24 September), Fable 5.1 $82, Sonnet 5 $19, Haiku 4.5 $1, with no batch discount and no regional multiplier. 464 of the 65,487 requests went to another provider’s model and carry no dollars in the total. For context, two months of the same transcripts hold 143,000 requests with a 98.8% cache hit rate.
Claude Code usage limits start with one fact: every step re-sends the conversation
An agent step looks small: read a file, run a test, write a line. But every step sends the model the whole conversation again, from the first message to the last tool result. The model does not remember your session the way you do. It reads it again, every time.
Prompt caching softens this. The server keeps the start of the conversation for a while, and reading it back costs a tenth or less of the normal input price. On a subscription the cache lives for one hour (five minutes for subagents). That is why 99% of our input tokens were served from cache. It is also why the weight is still large: a light read (at API prices) of 318,000 tokens, repeated tens of thousands of times, adds up.
This is how the week’s tokens weigh, at API prices ($6,527 in all):
Where the week’s weight went
| Line | At API prices | Share |
|---|---|---|
| Cache read | $4,245 | 65.0% |
| Cache write, 1 hour | $1,234 | 18.9% |
| Model output | $804 | 12.3% |
| Cache write, 5 minutes | $244 | 3.7% |
| Uncached input | $1 | 0.0% |
Two things stand out. Model output, the part we think of as the work, is 12% of the weight. Re-reading the conversation is 65%. And the average step carried 318,000 tokens of context, which is a lot to read a hundred times before lunch.
Why is the cache read the biggest line if it is so light per token at API prices? Volume. On Opus 5.5 a cache read costs $0.20 per million tokens against $4 for fresh input. But the week read 20.8 billion input tokens, and 99% of them came from cache. A tiny price, multiplied by an enormous count, is still the largest number in the count.
Context length sets the weight of a step: the Claude Code context window in numbers
Because every step re-reads the context, a step weighs more the longer the conversation is. We grouped all 65,487 requests by the size of the context at that step:
Weight by the size of the context at the step
| Context at the step | At API prices | Requests |
|---|---|---|
| up to 100k | $486 | 10,835 |
| 100–200k | $780 | 14,347 |
| 200–400k | $1,669 | 17,966 |
| 400–700k | $3,042 | 19,599 |
| over 700k | $550 | 2,740 |
55% of the week’s weight is steps above 400k.
Weight of one step by context size
| Context at the step | Per step | Against under 100k |
|---|---|---|
| up to 100k | $0.045 | |
| 100–200k | $0.054 | |
| 200–400k | $0.093 | |
| 400–700k | $0.155 | 3.4× |
| over 700k | $0.201 |
A step at 400–700k tokens weighs 3.4 times a step under 100k. The 400–700k band alone is $3,042, 46.6% of the week, and steps above 400k together make up 55% of the total. With a context window of a million tokens there is a lot of room to grow into.
Sessions behave the same way. Of 245 main sessions, 90 reached 400,000 tokens or more, and those 90 produced 82% of the weight. And contexts rarely shrink: only 27 times in the week did an agent’s context halve, through compaction or /clear. Most simply kept growing.
The lesson is not “use a smaller window”. The window is a ceiling, not a target. What matters is how long a conversation runs before someone cuts it. Start a fresh session when the topic changes, compact when a task is done rather than when the window is nearly full, and treat a conversation that has grown past a few hundred thousand tokens as heavy at every step, because it is.
What wakes an agent up
In a team of agents, most steps do not start with you. Something wakes an agent: a message from another agent, your prompt, a subagent it launched. We split the week’s turns by what started them:
What starts a turn
| Started by | At API prices | Share |
|---|---|---|
| A message from another agent or a system notice | $4,799 | 73.5% |
| A human | $1,141 | 17.5% |
| Subagents | $426 | 6.5% |
73.5% of the weight: turns started by another agent.
Almost three quarters of the weight sits in turns that another agent started, and the woken agent was carrying 415,000 tokens of context. Short reply turns, one to three requests such as “ok, got it”, made up 5,486 turns and $1,146, 18% of the weight.
There is a floor here. Waking an agent that carries 415,000 tokens weighs at least $0.08 on Opus 5.5 at API prices, even if its whole answer is one word. That is nothing once. It is a lot five thousand times. This describes multi-agent work; a single session has far fewer wake-ups.
What can be done? Ask agents to answer only when there is something to report: an acknowledgement sent to an agent with a large context is a full-weight step. Send several updates as one message instead of many. And give quick questions a lighter home: a subagent, whose context averaged 100,000 tokens, can take a small question without waking the agent that carries 415,000.
Idle time, the cache and the Claude Code usage limits
The cache expires. In the week, an agent sat idle for more than an hour 166 times. Each time its cache expired, and the next step had to write about 396,000 tokens again: $475 in total, 7.3% of the weight, about $2.90 each.
The one-hour cache is worth its extra weight at API prices. Writing to it cost about $454 more than a five-minute cache would have (measured tokens, at the price difference). We estimate that re-writing the context after the 1,498 pauses of five minutes to an hour, had the cache lasted five minutes, would have cost about $2,960, roughly 6.5 times more. That a one-hour write weighs 1.6 times a five-minute write in your limit is our assumption; at API prices it does. So the practical rule is short: before a long idle, hand the work over or run /compact, so that what waits is a short conversation, not a 400,000-token one.
Nothing here says that idling is bad. Agents wait for builds, for reviews, for you. The weight comes from waiting with a large context and then resuming. A short context is light to write again even after the cache has expired.
How does this weight map to the weekly limit? Anthropic does not publish the formula, so this is only an estimate: we saw about $30 of API-priced work move the weekly limit by 1% (two accounts, one a Max account and one whose plan we did not record; one day). Other use of the same login also moves the percent, so the true figure is at least this.
Tool output and screenshots stay in the context
Tools are light to run and heavy to keep. We counted 55,277 tool results, about 107 million characters in total. 12.5% of the results are longer than 5,000 characters and make up 61% of the volume. The agents also took in 2,349 images, mostly screenshots.
Writing that output is almost free. The weight comes later: it stays in the conversation and is re-read on every following step, at whatever the context weighs by then. A few habits keep it out:
- Pipe long command output through
tailorhead, or ask for one summary line. - Run
git show --statbefore a full diff. - Read a file with
offsetandlimitinstead of all of it. - Send big logs to a subagent and take back a summary.
- Take a screenshot only when text will not do, and keep it small.
Screenshots deserve their own mention. An agent that checks its own interface work may take many and keep them all, and each stays in the conversation the way text does. If a check needs one screenshot, take one, at a size that answers the question.
The model matters less than it seems
It is natural to reach for a lighter model first. In our numbers the model matters, but less than context length.
We took 64,999 requests that ran on Opus 5 between 14 August and 24 September (about six weeks, not the week above) and priced the same tokens at Opus 5.5’s published rates. $12,321 became $6,166. Cache reads are 2.5 times cheaper on Opus 5.5, and every other line is 1.25 times cheaper. Only the price list changed; the tokens did not. Anthropic says Opus 5.5 at medium effort uses about half the tokens on multi-step code, so the real saving may be larger.
Now Sonnet 5.5 against Opus 5.5 at equal steps. Input and output are half the price ($2 and $10 per million tokens, against $4 and $20). Yet the total drops only about 16% on average, and about 10% for the lead role. The reason is that cache reads are 65% of our week’s weight (40–81% by role), and both models price them the same, $0.20 per million tokens. The 16% comes from about 3,000 requests of the earlier period, all on Opus 5.5, assuming the same number of steps: arithmetic on old steps, not a measurement. A pilot of per-role models started on 29 September, with some roles on Sonnet 5.5; there are no results yet. One documented fact may matter more than the 16%: Opus and Sonnet have separate limits (Anthropic: each applies only to requests to that model family; the sizes are not published), so roles on Sonnet do not draw on the Opus limit.
| Model | Input | Cache write, 5 min | Cache write, 1 h | Cache read | Output |
|---|---|---|---|---|---|
| Opus 5.5 | $4 | $5 | $8 | $0.20 | $20 |
| Opus 5 | $5 | $6.25 | $10 | $0.50 | $25 |
| Sonnet 5.5 | $2 | $2.50 | $4 | $0.20 | $10 |
| Fable 5.1 | $10 | $12.50 | $20 | $0.25 | $50 |
| Haiku 4.5 | $1 | $1.25 | $2 | $0.10 | $5 |
Opus 5 against Opus 5.5, same tokens
| Priced at | At API prices |
|---|---|
| Opus 5 prices | $12,321 |
| Opus 5.5 prices | $6,166 |
Cache reads 2.5 times cheaper, every other line 1.25 times.
So where does that leave the choice of model? A lighter model helps, but its saving is capped by the cache-read line, which barely moves. Price lists change (a planned Sonnet 5 price rise was cancelled in September 2026). The safer habit is to look at your own transcripts after every change, of price or of model, rather than trust the list.
What we are changing at home (a hypothesis)
If context length is the multiplier, the lever is to keep it short. Our own change is to hand a role over to a fresh agent at 300–400,000 tokens (30–40% of a 1M window), not to let it run up to 700,000 (70%).
This is a hypothesis, not a result. Modelled on our week, handing over cuts about 30–35% of the weight (35% at 300k, 30% at 400k), against about 13% at 700k. The model counts handovers across all main sessions of the week and assumes the successor starts at 60k. We assume the real saving is about half of the modelled one. It also has a price: more handovers, roughly 190 a week at 300k instead of about 45 at 700k, so the handover note has to be good. We have not validated this in practice. A follow-up measurement is planned.
Why not hand over even earlier? A handover is not free. The fresh agent starts from a written note and has to read again what it needs, and a handover that loses a decision costs more than the tokens it saved. That is a reason to be careful with this lever, and a reason to measure it before calling it a win.
Count it yourself
You can check your own numbers. Claude Code writes every request to ~/.claude/projects/*/*.jsonl. Each assistant message carries a message.usage object with input_tokens, cache_creation_input_tokens, cache_read_input_tokens and output_tokens. Two traps: one request can appear on several lines of a transcript (streaming), so deduplicate by message id, or the sum is inflated; and cache_creation_input_tokens mixes five-minute and one-hour writes, whose split is in cache_creation.ephemeral_5m_input_tokens and ephemeral_1h_input_tokens. Subagents write to */subagents/*.jsonl. Add the tokens up per session, multiply each kind by its published price, and sort the sessions by their largest context. The long ones will be at the top.
What to take from this
If you keep only five lines from this article:
- Look at context length first. It multiplies the weight of every step.
- Cut before you idle: hand over or run
/compactbefore a long wait. - Keep tool output short, and send big logs to a subagent.
- Keep messages between agents few and short.
- Compare models on your own transcripts, not on the price list.
Where harnsy fits
harnsy is a local app that connects Claude Code, Codex and OpenCode agents into one team with roles. It does not shrink your tokens by itself, but a few things in it follow from the numbers above:
- It tells the lead when an agent’s context runs high. The threshold is set per role; in v0.7.0 the default is 40% of the window, and the dashboard suggests 300–400 thousand tokens.
- Handover runs on a written note: the agent writes it, the lead seats a fresh agent, the newcomer confirms. The work outlives the context window. It is not infinite memory.
- Since v0.7.0 each role has its own harness, model and reasoning effort.
- The lead is told when an agent sits idle.
How it installs, and what it changes on your machine, is on the install page.
Keep your agents’ memory short
harnsy connects your Claude Code, Codex and OpenCode agents into one team with roles, and hands work over before a context grows too long.
Install harnsy