Anthropic Claude Pricing Guide 2026: Every Model, Real Workload Math

Executive summary
Anthropic's Claude lineup in 2026 is a clean three-tier ladder: Haiku 4.5 for speed and volume, Sonnet 4.6 for balanced production work, Opus 4.6 for the hardest reasoning and agentic tasks. Pricing is straightforward and uniform across tiers (output is always 5× input), which makes budgeting easier than for some competitors. The biggest cost lever is not tier selection but prompt caching, which Anthropic offers at one of the steepest discounts in the industry — up to 90% off cached input tokens.
This guide breaks down published 2026 rates, shows real-workload math, and explains the four discounts that change the effective price. Run your actual numbers in the LLM cost calculator.
2026 standard API pricing (per 1M tokens)
| Model | Input | Output | Context | Output/Input ratio | | ----------------- | ----- | ------ | ------- | ------------------ | | Claude Haiku 4.5 | $1.00 | $5.00 | 200k | 5× | | Claude Sonnet 4.6 | $3.00 | $15.00 | 200k | 5× | | Claude Opus 4.6 | $5.00 | $25.00 | 200k | 5× |
Earlier models still served (Claude Sonnet 4.5, Opus 4.1, Opus 4) retain their original pricing until officially deprecated. Check the Anthropic console for current availability.
Four discounts that change the effective price
Prompt caching is Anthropic's largest savings lever. The first time you send a prompt with a cached prefix, you pay 1.25× the input rate to write the cache. Every subsequent read of that cached prefix is 0.1× the input rate — a 92% reduction. The cache lives for five minutes by default, with a one-hour option for an additional 2× write cost. For workloads with stable system prompts and tool definitions, caching commonly cuts effective input cost by 60–85%.
Batch API offers 50% off both input and output rates for requests that complete asynchronously within 24 hours. Anthropic's batch endpoint is straightforward and the cost reduction is real. Use it for nightly summarization, embeddings, eval runs, and any user-tolerant async work.
Volume commitments unlock enterprise discounts. Anthropic does not publish these rates but they kick in around $50k–$100k/month committed spend. Negotiate at renewal; the leverage grows with your spend.
Cross-region pricing parity. Unlike some vendors, Anthropic's regional endpoints (US, EU, AWS Bedrock, GCP Vertex) keep substantially similar pricing. There is no cost-hopping opportunity, but also no surprise regional markup.
Real-workload math
Customer support, 2M conversations/month, 2,500 input / 600 output tokens.
Standard Claude Haiku 4.5: $1 × 5,000 + $5 × 1,200 = $11,000/month. With prompt caching (80% of input cached): effective input = $1 × 1,000 + $0.10 × 4,000 = $1,400. Output unchanged at $6,000. Total $7,400/month — a 33% reduction with one engineering change.
Long-form analysis, 100k requests/month, 30,000 input / 4,000 output tokens.
Standard Claude Sonnet 4.6: $3 × 3,000 + $15 × 400 = $15,000/month. With batch API (50% off): $7,500/month.
Agentic coding workload, 10k sessions/month, 50,000 input / 12,000 output tokens.
Claude Opus 4.6 standard: $5 × 500 + $25 × 120 = $5,500/month. With caching across stable tool definitions (60% input cached): effective input = $5 × 200 + $0.50 × 300 = $1,150. Output unchanged at $3,000. Total $4,150/month.
The pattern is consistent: caching plus right-tier selection cuts production Claude spend by 30–70% from the standard rate-card math.
When to pick each Claude tier
Haiku 4.5 for high-volume routine work: classification, routing, short summarization, support deflection. It is the default tier for cost-sensitive production workloads. Where it falls short is multi-step reasoning and long-context synthesis.
Sonnet 4.6 is the workhorse. Drafting, extraction, document Q&A, code review, agent steps. For most production AI features at most companies, Sonnet 4.6 is the right default — the quality-per-dollar curve hits its sweet spot here.
Opus 4.6 for the hardest reasoning, multi-step agents, complex coding, and tasks where output quality directly drives revenue. Reserve for the use cases that justify the 5× premium over Sonnet.
How Claude compares on price
Against OpenAI, Claude tiers are roughly 1–2× more expensive at the equivalent level (Sonnet vs GPT-5; Haiku vs GPT-5 mini). Against Google Gemini 3.1 Pro Preview, Claude Sonnet is roughly 2.5× more expensive on input. Against Qwen-Max, Claude Opus is roughly 3× more expensive. See our GPT vs Claude comparison for the head-to-head workload math.
The case for paying the Claude premium rests on three real strengths: instruction-following on long prompts, refusal behavior in regulated contexts, and writing quality on long-form synthesis. If those matter for your use case, the price difference is usually worth it. If they do not, a cheaper vendor wins on TCO.
Hidden costs to budget for
Vision input. Claude charges for image tokens at the standard input rate; a typical image runs 1,500–2,000 tokens. Heavy multimodal workloads should be modeled explicitly.
Tool use overhead. Tool definitions are sent on every request. With caching this is cheap; without caching it is a meaningful line item.
Eval and red-team runs. Running 5,000-example evals across Haiku, Sonnet, and Opus during a model upgrade is a real cost — budget for it as part of any migration.
Failed and retried requests. Anthropic does not charge for true API errors, but does charge for outputs that fail your downstream validators and get retried. Counted across a month this can be 5–15% of spend.
A practical adoption playbook
Start with Sonnet 4.6 as your default. Move high-volume routine workloads to Haiku 4.5 after a 100-example eval shows quality holds. Reserve Opus 4.6 for the specific use cases where the quality difference is measurable in business outcomes. Enable prompt caching from day one. Move tolerant workloads to batch. Track cost per feature, not just total spend.
For the broader cost-reduction toolkit, see how to reduce context-length costs and run your numbers in the calculator.
Bottom line
Claude's pricing is honest, ladder-shaped, and reward-aligned: pay more for harder work, save through caching and batch. The teams that lose money on Claude are the ones running Opus by default. The teams that win are the ones with disciplined tier-routing and aggressive caching. The savings sit in the operating discipline, not the rate card.
Related tools for lowering LLM spend
- Use the LLM cost calculator to estimate per-request spend before you ship.
- Run the prompt optimizer to trim wasted tokens while keeping your intent intact.
- Package the task with the context engineer when you need a cleaner brief and fewer retries.
- If you want hands-on support, see Reduce My Cost for tailored guidance.