Claude Sonnet vs Opus for coding agents: choose a model and effort per role
“Sonnet or Opus?” is asked as if one answer fits a whole team. In a team of agents the roles do different work, so the answer differs by role, and it is more about fit and risk than about the price. Here is what the vendor says, what the prices really do, and what we have and have not measured.
Key points
- $0.20
- per million tokens to read from the cache, the same on Sonnet 5.5 and on Opus 5.5; every other price line on Sonnet is half
- 10–34%
- what a role would save by moving to Sonnet 5.5 at the same number of steps (16% on average); arithmetic, not a measurement
- 0 of 29,771
- requests in our product project between 10 and 29 September, before the pilot, ran on Sonnet 5.5 or on high effort, so we have no result yet
- −2%
- the estimated change of a week’s weight in our product project ($3,033 at API prices) with the role table as applied: the saving and the extra cost largely cancel
In this article
- What the vendor says
- The price, and the one line that does not move
- Why the saving is smaller than it looks
- Effort changes the number of steps, not only the length of an answer
- Role by role: our pilot table
- Two things to keep off the table
- Other harnesses are the same decision
- What is measured, and what is still a pilot
- Where harnsy fits
- Sources
What the vendor says
Anthropic released Sonnet 5.5 on 28 September 2026 and describes it as “a faster, lower-cost complement to Claude Opus 5.5”. Its prompting guide for the model adds that “for the hardest long-horizon work, an Opus model is the better choice”. Claude Code’s cost page puts the split in one line: “Sonnet handles most coding tasks well and costs less than Opus. Reserve Opus for complex architectural decisions or multi-step reasoning.” For teams of agents it says: “Use Sonnet for teammates.”
The announcement’s benchmarks point the same way, with a catch: they are not medium-effort results.
| Benchmark | Sonnet 5.5 | Opus 5.5 | Effort in the announcement |
|---|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 66.4% | xhigh or max |
| FrontierCode 1.1 | 46.2% | 54.4% | max |
| CursorBench 4.0 | 55.5% | 57.8% | not stated |
| GDPval-AA | 1844 | 1846 | not stated |
Read plainly: on everyday work the two are close, Sonnet is ahead on Terminal-Bench, and on the hardest code Opus leads by eight points. That is a reason to match the model to the job, and nothing yet about your own repository.
The price, and the one line that does not move
Anthropic’s published API prices for the two models, per million tokens:
| Line | Sonnet 5.5 | Opus 5.5 |
|---|---|---|
| Input | $2 | $4 |
| Cache write, 5 min | $2.50 | $5 |
| Cache write, 1 h | $4 | $8 |
| Cache read | $0.20 | $0.20 |
| Output | $10 | $20 |
Cache read costs the same on both models.
Everything on Sonnet costs half, except the cache read. Anthropic prices a cache read on Opus 5.5 at 0.05 of its input price and on Sonnet 5.5 at 0.1, and both land on $0.20. An agent step re-reads the whole conversation, so that one equal line decides more than the rest. Where our agents’ tokens go has the week we counted; here we only need what follows from it.
Why the saving is smaller than it looks
In our week of 23–29 September 2026 (our product project, all Opus 5.5, $3,033 at API prices), cache reads were the largest part of a role’s weight: 81% for the lead, 76% for the landing copywriter, 68% for the reviewer, and 32% for the architect, whose work is short. We priced the same steps on Sonnet 5.5. The saving runs from 10% for the lead to 34% for the architect, and 16% across the week ($495 of the $3,033, weighted by cost, not a plain mean of the roles), at the same number of steps.
| Role | Cache read, share of weight | Sonnet, same steps |
|---|---|---|
| Lead | 81% | −10% |
| Landing copywriter | 76% | −12% |
| Reviewer | 68% | −16% |
| Frontend | 65% | −17% |
| Developer | 64% | −18% |
| Docs | 58% | −21% |
| Architect | 32% | −34% |
The longer a role’s context, the less Sonnet saves. The lead and the copywriter carry about 430k tokens per step and save 10–12%; the architect carries 92k and saves 34%, on $13 a week.
Two cautions. If Sonnet needs about 20% more steps for the same job (our assumption, not a measurement), the saving of the long-context roles disappears. And the seven bars are the roles with a full week of work; short-lived roles cost too little to matter.
Effort changes the number of steps, not only the length of an answer
Effort is a setting on the model that Anthropic describes as affecting “all tokens” in a response: text, tool calls and thinking. It is “a behavioral signal, not a strict token budget”. For an agent the tool calls matter most. Lower effort means fewer and terser tool calls; higher effort means more of them, a plan explained before acting, and longer summaries. Every extra tool call is another step, and every step re-reads the whole conversation.
- low
The most efficient: simpler tasks that need speed and the lowest cost, such as subagents.
- medium
Balanced: agentic tasks that need a balance of speed, cost and performance.
Opus 5.5 defaultSonnet 5.5 in Claude Code - high
Complex reasoning, difficult coding problems, agentic tasks.
Sonnet 5.5 in the API - xhigh
Long-running agentic and coding tasks, over 30 minutes, with token budgets in the millions.
- max
No constraint on token spending; for the deepest reasoning.
For agentic coding on Sonnet 5.5 Anthropic suggests starting at medium for well-specified tasks and moving to high for harder or longer ones. It also notes that at low the model can skip verifying a change, and that at low and medium on long tasks it is more likely to stop and check in before finishing.
In our week, the model’s own answer, thinking included, was only 8–29% of a role’s weight, and thinking was 12–29% of that. So effort would reach our bill mainly through extra steps, not through longer answers. That is a reading of the vendor’s description and of our split; we have no measured run at high.
Role by role: our pilot table
Most of the reason to pick per role is fit and risk: where a mistake is expensive, use the stronger setting; where the work is well-scoped, use the faster one. Cost comes second. This is the table in harnsy’s own company role templates since 29 September 2026 (every project inherits it; none overrides it), with one exception marked below.
Pilot table, since 29 Sep 2026
| Role | Model | Effort | Why |
|---|---|---|---|
| Lead | Opus 5.5 | medium | Decisions and acceptance; a mistake spreads to the team. Cache reads are 81% of its weight, so Sonnet would save only 10%. |
| Developer | Opus 5.5 | medium | Code with invariants. Anthropic’s benchmark gap is widest on hard code. |
| Reviewer | Opus 5.5 | high | The last filter before the code goes live, worth the stronger setting. Anthropic lists high for difficult coding and agentic tasks. |
| Integration | Sonnet 5.5 | medium | A step-by-step procedure: merge, build, check, deploy. Sonnet leads on Terminal-Bench. |
| Frontend | Sonnet 5.5 | medium | Well-scoped UI work against clear rules. |
| Docs, landing copywriter, marketer | Sonnet 5.5 | medium | Documents and text: well-scoped, routine work, rarely sent back (docs 7%, copywriter 0 of 98). |
| Editions expert | Sonnet 5.5 | high | Rules that end up in a license; high compensates for the lighter model. |
| Designer, UX | Opus 5.5 | medium | Taste and judgement, at a small cost. |
| Lawyer | Opus 5.5 | high | Legal text, where an error costs money. |
| Roles that touch keys and cryptography | Opus 5.5 | high | Recommended, not applied yet: no template row. Their work goes back the most. |
We chose the rows partly by how often a reviewer sent a role’s work back. These are the reviewer’s reject messages by the author’s role since 24 September 2026, with the sample sizes:
- Roles that touch keys and cryptography: 25% (13 of 52), a small sample.
- Lawyer: 25% (1 of 4), a very small sample.
- Developer: 12% (32 of 271).
- Integration: 9% (8 of 86).
- Docs: 7% (4 of 54).
- Designer: 7% (1 of 15), a very small sample.
- Frontend: 2% (2 of 87).
- Landing copywriter: 0 of 98.
A send-back rate says how often a check found something, not how good the author was. Read it as a hint for where to spend a stronger setting, and no more.
Two things to keep off the table
Do not switch the model inside a live session. Each model has its own prompt cache, so after a switch “the next request reads the entire conversation history with no cache hits”, in the words of the Claude Code documentation; Claude Code asks you to confirm while the cache is still warm. Choose the model when an agent starts, or at a handover. The same page notes that opusplan, which uses Opus in plan mode and Sonnet for execution, makes every plan-mode toggle a model switch, and each one starts a fresh cache.
Effort is kinder. On Opus 5.5 and Sonnet 5.5, changing effort in Claude Code keeps the cache (with an API key or a Claude subscription). In the API, a per-message effort change, which is in beta, also keeps it, while a new top-level effort value between requests does not, so hold it constant inside a conversation that relies on the cache.
One more thing from Anthropic’s guide for Sonnet 5.5 belongs on the list of risks, not on the list of findings. It says the model is trained to resist injected instructions and “sometimes it treats a genuine user message as a possible injection”, especially when a message from the user lands right after a tool result in the middle of a turn. In a team where messages arrive while an agent is working, that is worth watching. We list it as a risk: we have not seen it on our own work.
Other harnesses are the same decision
“Pick per role” does not stop at one vendor. harnsy’s launch settings start with the harness (Claude Code, Codex or OpenCode) and only then the model. Our own numbers on Codex are thin and from a different period: over all the months of its session files, an average step carried 113–131k tokens of context, against 318k for Claude across all our projects in the week above, and 96.5% of its input came from the cache, against 99.0%. In September Codex carried about a twentieth of Claude’s input tokens. That says something about how we use it, not about how good it is. We have no quality comparison, so we make none.
What is measured, and what is still a pilot
- Measured
From 10 to 29 September, before the pilot, our product project sent 29,771 requests, all on Opus 5.5 at medium, none on Sonnet 5.5 or on high. The cache-read shares, the weight per role, the reviewer’s send-backs by role, and the price arithmetic.
- Estimated
The role table as applied moves a week by about −2%: the Sonnet roles save about $146 and the high roles add about $81, net about −$65, on a $3,033 week in our product project at API prices. The table as first proposed, with more roles on high and Sonnet, came to −1.5%. High is assumed to cost 20% more (range 10–40%).
- Not measured yet
How Sonnet 5.5 behaves on our work. What high does to the number of steps. How much of the weekly limit Sonnet uses against Opus: Anthropic publishes no ratio, and Claude Code has separate limit messages for the Opus and Sonnet families.
The table went into the company role templates on 29 September 2026, so it is new and there are no results. A second measured week starts on or after 6 October. It compares the pilot week with the baseline week of 23–29 September: spend per role, how much of the Sonnet limit against the Opus limit is used, send-back rate per role, human rework, the integration role’s failed deploys and git mistakes, and speed. It ends with a decision per role: keep, revert or change the effort.
Where harnsy fits
In harnsy each role has a Launch section with three fields in this order: harness, model, effort. Harness is a choice, and an empty one inherits. Model is free text with suggestions; prefer a full model id, because an alias moves to the next version by itself. Effort appears only when the harness has levels: Claude has low, medium, high, xhigh and max, Codex has its own list, and OpenCode has none. Changing the harness clears the model and the effort.
This project’s own launch: new agents of the role start with it; open and restore keep a seat’s harness and take the model and effort only on it.
There are two levels: the company template, which every project inherits, and the project’s own role, which overrides it. A value set on a single start beats both. A change applies to new sessions only (a start, a handover, a restore): saving restarts nobody. Opening or restoring an agent keeps the harness it already runs and takes the role’s model and effort only if that harness is the role’s.
The picker does not make any choice cheaper by itself; it makes a choice you have made stick to a role, and it makes the next handover start on the setting you picked. Our own table is only a starting point for your first week, and our second week will say what is worth keeping. Install harnsy to set launch per role.
Sources
- Anthropic, “Introducing Claude Sonnet 5.5”, 28 September 2026: anthropic.com.
- Claude API pricing: platform.claude.com.
- Effort: platform.claude.com.
- Prompting Claude Sonnet 5.5: platform.claude.com.
- Claude Code, model configuration: code.claude.com.
- Claude Code, manage costs: code.claude.com.
- Claude Code, how it uses prompt caching: code.claude.com.
Set the model and effort for each role
harnsy connects your Claude Code, Codex and OpenCode agents into one team with roles, and each role starts with the harness, model and effort you chose.
Install harnsy