Claude Opus 5 API is $5.00 per million input tokens, $25.00 per million output tokens and $0.50 for cached input — Anthropic’s published list rates for its flagship reasoning model, available at face value rather than through a marked-up reseller. At that card it costs exactly half of Claude Fable 5, the more expensive sibling in the same family. But the rate card is the least useful number in this article: Opus 5 bills its thinking as output and thinks a lot, so the price you actually pay is a function of how long it deliberates. The same rate card stays current on Claude Opus 5 as Anthropic reprices; this is the straight version of what the pricing does to a real invoice.
Here is the situation most readers are in. You have a workload — a coding agent, a pipeline of document-understanding jobs, a long-horizon research task — and you are deciding whether the model at the top of the independent leaderboards is worth its cost. You landed here because the per-token price is easy to find and hard to reconcile with the invoices people actually report. That gap is not a pricing bug. It is the difference between the model’s smallest unit of measurement and its real unit of work. The token is what you are charged for; the task is what you are buying.
The rate card, in full
• Input — $5.00 per million tokens
• Output — $25.00 per million tokens
• Cached input — $0.50 per million tokens (an 80% reduction from the input price)
These are vendor-reported figures, published by Anthropic and current as of August 2026. Two things are worth noticing before we move on.
First, the context window is one million tokens. That matters for the cache: a stable repository or document corpus that fits in a single 1M-token context is exactly the workload where cached input at $0.50 does most of the work.
Second, $25 per million output tokens is not unusual for a frontier reasoning model — but Opus 5 is not a model you call for one output token. Everything it answers begins with thousands of tokens of deliberation, and in Anthropic’s API those reasoning tokens are billed as output. That single mechanic is responsible for almost every “why is my bill so high” story about this model, and it is the reason the token price alone is a trap.

What a reasoning model does to your bill
The cleanest read on this comes from Artificial Analysis, the independent evaluator that runs a fixed battery of tasks against every model it tracks. Its latest measurements, checked August 22, 2026, put Claude Opus 5 (Adaptive Reasoning, Max Effort) at the top of the Intelligence Index at 63.05, first of 185 models.
The same run reveals the cost structure. The full Intelligence Index evaluation on Claude Opus 5 cost $3,836.05, consuming 100 million output tokens — against a median of 72 million for the models in that tier. Turn that into a per-task figure and you get $2.34 per completed task. That is the number that should drive your forecasting, because it is the price of the actual thing you are buying: a finished answer.
To see why the per-task number is so much more informative than the per-token one, compare against GPT-5.6 Sol, the other flagship on the board. Artificial Analysis measured Sol at $1.23 per task on the same index — about half of Opus 5’s cost — while its token prices are in the same premium range. The difference is almost entirely deliberation: Opus 5 burns more output tokens per unit of work, and at $25 per million, tokens spent thinking are the most expensive tokens on the invoice.
There is a physical signature of this too, from OrcaRouter’s own telemetry: over a rolling seven-day window, Claude Opus 5 showed a p50 time-to-first-token of 7.34 seconds, against 1.33 seconds for the volume workhorse GPT-5.6 Luna. Opus 5 thinks longer before it says anything at all. It is an output-quality model, not a latency model, and that trade is built into its cost.
Budget in tasks, not tokens
If you are forecasting spend for the quarter, price-per-million is the wrong denominator. Reasoning models make token-based forecasts unreliable because token consumption per task is a moving target: it depends on the effort setting, on how much context the model re-reads, and on the task itself.
The honest way to plan is to take a representative sample of your real workload, run it at your intended effort setting, and read off the per-task cost. The picture that comes back is consistent: this is a premium reasoning model whose cost per completed task — $2.34 on the independent index — is roughly twice that of its main premium competitor.
That does not mean it is overpriced. It means the question is never “is $25 per million output too much?” It is “does a 63.05 Intelligence Index, on my tasks, justify paying about twice the per-task price of the nearest alternative?”

The effort dial is a cost lever
Opus 5’s “Adaptive Reasoning” exposes an effort setting, and Artificial Analysis runs the model at four of them on the same index: max, xhigh, high, medium. The score ladder is worth reading as a price list, not a spec sheet:
• Max effort — Intelligence Index 63.05
• Xhigh — 62.52
• High — 61.48
• Medium — 58.64
Every step down costs you a little index performance, and every step down also spends fewer output tokens on deliberation — which, at $25 per million, is where the savings live. Medium effort still scores 58.64, a figure that would have been a flagship result twelve months ago, at a fraction of the output-token burn of the max configuration.
The practical rule: tune effort to the task, not to the launch-day default. Agentic and long-horizon work is where the independent scores moved most at high effort settings, so spend there. Simple extraction, classification and formatting tasks do not need a model that argues with itself for three pages of output before returning a boolean. On those, the medium setting is usually the cheapest correct answer.
How $5 / $25 sits in the market
Cross-model price comparisons are where most articles go wrong, so the rule here is to compare tiers you can actually buy today, at post-cut prices.
| Model (config) | Input / 1M | Output / 1M | Cached input / 1M |
| Claude Opus 5 | $5.00 | $25.00 | $0.50 |
| Claude Fable 5 | $10.00 | $50.00 | — |
| GPT-5.6 Sol (post-cut) | $5.00 | $30.00 | — |
| GPT-5.6 Luna (post-cut) | $0.20 | $1.20 | — |
All prices above are vendor-reported list prices. The GPT-5.6 figures reflect OpenAI’s cut, so the comparisons hold as of August 2026.
Three takeaways from that table. First, Opus 5 at $5/$25 is positioned, in Artificial Analysis’ own words, “near-frontier at half the price of Claude Fable 5” — it is the cheaper flagship by a factor of two on the rate card, and it leads Fable 5 on the independent Intelligence Index (63.05 vs 62.07). Second, at $5/$25 it is not a volume model: Luna at $0.20/$1.20 is where the throughput lives, and OrcaRouter’s telemetry confirms it — Luna carried about 21,271.6 million tokens in a rolling seven-day window against Opus 5’s 491.5 million. Third, this is why the per-task lens matters so much: on the rate card Opus 5 and Sol are near-identical at the input, yet on the independent index their per-task costs differ by roughly two to one.
Paying the list price, not a markup
Finally, a word on distribution. Opus 5 is available through the vendor’s own API and several third-party platforms, and the reseller you choose can change your effective price even when the model cannot. The one hard rule worth carrying around: a router’s markup is your cost, and it is the one part of the pricing you can actually control.
OrcaRouter carries Claude Opus 5 in its catalog and passes the Anthropic list price through at 0% markup — $5 in, $25 out, $0.50 cached, no spread on top. The practical consequence of that model is that a vendor price change reaches you the day it happens rather than the quarter after, and the per-task math in this article is the per-task math on your invoice. Whatever platform you use, ask it the one question that matters: are you paying the list price, or the list price plus a tax on a model that is already expensive per task?
The takeaway
Claude Opus 5’s rate card is simple and genuinely competitive — $5 in, $25 out, $0.50 cached, half the price of Fable 5 for the top independent index score. The pricing trap is not the card, it is the unit of measurement: this model spends output tokens thinking, so budget per completed task, not per token. Tune effort down for routine work, use the cache where your context is stable, price a real sample of your workload, and make sure the platform between you and the model is passing the list price through rather than marking it up. Do that, and the most expensive reasoning model on the board becomes a predictable line item instead of a surprise.
Sourcing note: all token prices are vendor-reported list rates from Anthropic and OpenAI (GPT-5.6 figures are post-cut), current as of August 2026. Intelligence Index scores, per-task cost ($2.34), full-run cost ($3,836.05 on 100M output tokens, median 72M) and the effort ladder are independent measurements from Artificial Analysis’ live model page, checked August 22, 2026. Time-to-first-token and token-volume figures are OrcaRouter’s own telemetry over a seven-day window, checked August 22, 2026. Vendor pricing changes without notice.
