Quick answer
In 2026, LLM API pricing ranges from $0.05 per million input tokens (cheapest credible model: Qwen3-8B, GPT-5 nano) to $30 per million input tokens (most expensive: GPT-5.4 pro). Output tokens are typically 4-10× the input rate. The single biggest cost lever is picking the right tier for the task, not the cheapest vendor.
Use the full table below as a 2026 reference. Plug your actual prompt into the LLM cost calculator for per-message cost.
How to read this table
All prices are USD per 1 million tokens, charged separately for input (what you send) and output (what the model generates) at standard non-cached, non-batch rates. Context column is the maximum tokens per request. Caching, batch, and volume discounts can reduce these rates by 30-90% — see the LLM cost optimization playbook.
Full pricing table — sorted cheapest to most expensive (by input price)
| # | Provider | Model | Input ($/1M) | Output ($/1M) | Out/In ratio | Context | Tier | Best use | | --- | ---------- | ---------------------- | ------------ | ------------- | ------------ | ------- | ------------ | ---------------------------------- | | 1 | OpenAI | GPT-5 nano | $0.05 | $0.40 | 8× | 400k | Nano | High-volume routing/classification | | 2 | Alibaba | Qwen3-8B (Together) | $0.05 | $0.10 | 2× | 128k | Nano | Cheapest credible model | | 3 | OpenAI | GPT-4.1 nano | $0.05 | $0.20 | 4× | 1M | Nano | Legacy long-context cheap | | 4 | Alibaba | Qwen-Turbo | $0.05 | $0.20 | 4× | 128k | Nano | High throughput multilingual | | 5 | Google | Gemini 2.5 Flash Lite | $0.10 | $0.40 | 4× | 200k | Nano | Multimodal at low cost | | 6 | Google | Gemini 2.0 Flash Lite | $0.08 | $0.30 | 4× | 1M | Nano | Long context cheap | | 7 | Cloudflare | mistral-small-3.1-24b | $0.35 | $0.56 | 1.6× | 128k | Mini | Vision + text balanced | | 8 | Alibaba | Qwen3-32B (open) | $0.15 | $0.60 | 4× | 128k | Mini | Open-weights mid-tier | | 9 | OpenAI | GPT-5 mini | $0.25 | $2.00 | 8× | 400k | Mini | Default mini-tier choice | | 10 | OpenAI | GPT-4.1 mini | $0.40 | $1.60 | 4× | 1M | Mini | Long context balance | | 11 | Google | Gemini 3 Flash Preview | $0.50 | $3.00 | 6× | 200k | Mini | Latest Gemini Flash | | 12 | Alibaba | Qwen-Plus | $0.40 | $1.20 | 3× | 128k | Mini | Multilingual mid-tier | | 13 | DeepSeek | DeepSeek V3 (AWS) | $0.27 | $1.10 | 4× | 128k | Mini | Coding-strong cheap | | 14 | xAI | Grok-4-fast | $0.20 | $0.50 | 2.5× | 256k | Mini | xAI cheap fast option | | 15 | xAI | Grok-code-fast-1 | $0.20 | $1.50 | 7.5× | 256k | Mini | Code-specialized | | 16 | Anthropic | Claude Haiku 4.5 | $1.00 | $5.00 | 5× | 200k | Mid | Anthropic cheapest | | 17 | Google | Gemini 2.5 Pro | $1.25 | $10.00 | 8× | 200k | Mid | Strong reasoning | | 18 | OpenAI | GPT-5 | $1.25 | $10.00 | 8× | 400k | Mid | OpenAI flagship standard | | 19 | Alibaba | Qwen-Max | $1.60 | $6.40 | 4× | 128k | Mid | Cheapest Alibaba flagship | | 20 | Google | Gemini 3.1 Pro Preview | $2.00 | $12.00 | 6× | 200k | Mid | Latest Gemini Pro | | 21 | OpenAI | GPT-4.1 | $2.00 | $8.00 | 4× | 1M | Mid | Long-context flagship | | 22 | OpenAI | GPT-5.4 | $2.50 | $15.00 | 6× | 1.05M | Mid-Flagship | OpenAI long-context flagship | | 23 | Anthropic | Claude Sonnet 4.6 | $3.00 | $15.00 | 5× | 200k | Mid | Anthropic workhorse | | 24 | Anthropic | Claude Sonnet 4.5 | $3.00 | $15.00 | 5× | 200k | Mid | Previous Sonnet tier | | 25 | xAI | Grok-4 | $3.00 | $15.00 | 5× | 256k | Mid | xAI flagship | | 26 | Anthropic | Claude Opus 4.6 | $5.00 | $25.00 | 5× | 200k | Flagship | Anthropic flagship | | 27 | OpenAI | o4-mini | $1.10 | $4.40 | 4× | 200k | Reasoning | Cheap reasoning model | | 28 | OpenAI | o3 | $10.00 | $40.00 | 4× | 200k | Reasoning | Hard reasoning legacy | | 29 | Anthropic | Claude Opus 4.1 | $15.00 | $75.00 | 5× | 200k | Flagship | Premium reasoning | | 30 | OpenAI | GPT-5.4 pro | $30.00 | $180.00 | 6× | 1.05M | Flagship+ | Most expensive credible model |
Pricing per first-party publication or major third-party host as of May 2026. Always verify against the vendor's pricing page before committing to a workload.
What this table actually means for cost
The cheap-to-flagship ratio is 600×. GPT-5 nano at $0.05/1M input vs GPT-5.4 pro at $30/1M input. The right tier for the right task matters more than the right vendor.
Output is 4-10× input on every tier. Plan for this when estimating spend. Force structured output and set hard token limits to control it.
Sub-$1/1M input is enough for most production work. Classification, routing, extraction, short summarization, and most chatbot turns work fine on this tier.
$1-$3/1M input is the workhorse range. Drafting, document Q&A, multi-step tool use, code review.
$3+/1M input should be reserved for hard work. Multi-step reasoning, complex agents, long-context synthesis with quality requirements.
Effective cost after discounts
Sticker prices are the ceiling, not the bill. Three discounts change the real number.
Prompt caching — OpenAI 50% off cached input, Anthropic up to 90%, Google 75%. For workloads with a stable system prompt, expect 30-60% off effective input cost.
Batch API — OpenAI, Anthropic, Google all offer 50% off both input and output for async workloads completing within 24 hours.
Volume commits — at $50k+/month, enterprise rate cards become available with all three major vendors.
A workload running on Claude Sonnet 4.6 at $3/$15 with prompt caching enabled and 60% of input cached costs an effective $1.30/$15 — closer to the mid-tier than the flagship rate.
Cost math examples at 2026 rates
A 2,000 input / 400 output token chatbot call
- GPT-5 nano: $0.0002
- Qwen3-8B: $0.00011 (cheapest)
- GPT-5 mini: $0.0013
- Claude Haiku 4.5: $0.004
- GPT-5: $0.0065
- Claude Sonnet 4.6: $0.012
- Claude Opus 4.6: $0.02
- GPT-5.4 pro: $0.132 (most expensive)
The cheapest to most expensive ratio is 1,200×. For a single product feature handling 1M calls/month, that is the difference between $110/month and $132,000/month.
A 30,000 input / 4,000 output token document analysis
- GPT-5 mini: $0.0155
- GPT-5: $0.0775
- Claude Sonnet 4.6: $0.15
- GPT-5.4: $0.135
- Claude Opus 4.6: $0.25
For 50,000 reports/month, GPT-5 mini delivers $775/month while Claude Opus 4.6 delivers $12,500/month at this volume.
Run your real prompts in the LLM cost calculator for accurate per-workload math.
Cheapest model by use case (2026)
| Use case | Recommended cheapest credible | Why | | ----------------------------- | ---------------------------------------------------------------- | ------------------------------------- | | Classification / routing | GPT-5 nano or Qwen3-8B | Sub-$0.10/1M, accuracy holds at scale | | Short summarization | GPT-5 nano | Fast, cheap, structured output | | Translation | Qwen-Turbo | Strong multilingual at $0.05 input | | Customer support chatbot | GPT-5 mini | Best quality-per-dollar with caching | | Code completion | Grok-code-fast-1 or GPT-5 mini | Code-specialized cheap tier | | Long-context Q&A | GPT-4.1 mini (1M context) | Best long-context cheap option | | Reasoning under cost pressure | o4-mini | Cheapest credible reasoning model | | Image generation pairing | (see image generation pricing) | Different price band entirely |
Most expensive ≠ best for your task
The $30/$180 GPT-5.4 pro tier is the right choice in roughly 1-3% of production AI calls. The market has aggressively segmented in 2026 — the cheap end is much more capable than it was 18 months ago, and the flagship end is much more expensive. Default to the lowest tier that passes your eval set, and escalate only for tasks that fail.
FAQ
What is the cheapest LLM in 2026?
Qwen3-8B hosted via Together AI at $0.05 input / $0.10 output per million tokens, and GPT-5 nano at $0.05 input / $0.40 output, are tied as the cheapest credible production models. For volume routine work, either is roughly 100× cheaper than a flagship model.
Why is output always more expensive than input?
Output generation requires the model to run autoregressive inference one token at a time, while input is processed in parallel. The compute cost ratio is genuinely 4-10× and shows up in vendor pricing accordingly.
How does prompt caching change the price?
Anthropic offers up to 90% off cached input tokens, OpenAI 50%, Google 75%. For workloads with a stable system prompt over 1,024 tokens, expect effective input cost to drop 30-60% after caching is enabled.
What is the cost difference between flagship and mini tiers?
Typically 5-10× within a vendor. GPT-5 mini ($0.25/$2) versus GPT-5 ($1.25/$10) is exactly 5×. Claude Haiku 4.5 ($1/$5) versus Claude Sonnet 4.6 ($3/$15) is 3×. Tier selection moves the bill more than vendor selection.
Are batch API prices included in this table?
No. Batch rates are 50% off both input and output across OpenAI, Anthropic, and Google. To get the batch rate, divide the table values by 2.
Which vendor is cheapest overall?
It depends on tier. At the nano tier, OpenAI and Qwen tie. At mini, OpenAI is cheapest. At mid/flagship, prices converge within 30% across OpenAI, Anthropic, and Google. There is no single "cheapest vendor" — there is a cheapest tier-vendor combination per workload.
How often does this pricing change?
Major vendors typically reprice 2-4 times per year. New tiers are added quarterly. Bookmark this page and run your numbers against the LLM cost calculator before any major capacity planning.
How do I cut my LLM bill if I'm overspending?
Follow the LLM cost optimization playbook. The top three actions: enable prompt caching, route routine workloads to a mini-tier model, and audit your system prompts for bloat. Together these are usually a 50-70% reduction.