Estimate request cost from real prompts and compare current token pricing across OpenAI, Anthropic, Google, Meta, xAI, and DeepSeek.
Calculator
LLM Cost Calculator
Enter a real prompt, choose a reply length, and instantly see cost per request against published per-1M-token pricing.
Our recommended LLM(s)
Best fit for you
Google Gemini
$0.000032 per message
Gemini 2.5 Flash Lite
Low-cost Gemini option for high-volume, cost-sensitive workloads.
Also worth a look
OpenAI
$0.000032 per message
GPT-5 nano
Lowest-cost OpenAI model for routing, classification, and ultra-high-volume requests.
OpenAI
$0.000032 per message
GPT-4.1 nano
Ultra-low-cost OpenAI model for simple, high-throughput tasks.
OpenAI
$0.0001 per message
GPT-4.1 mini
Cheaper long-context OpenAI model for balanced production workflows.
OpenAI
$0.0002 per message
GPT-5 mini
Current mini model for routine production tasks and high-volume workflows.
OpenAI
$0.0004 per message
o4-mini
Fast reasoning model for cheaper chain-of-thought and structured tasks.
Google Gemini
$0.0004 per message
Gemini 2.5 Flash
Fast Google model for balanced cost, latency, and multimodal capability.
OpenAI
$0.0004 per message
GPT-5.4 mini
Latest mini model in the GPT-5.4 family for coding and agentic workflows.
Anthropic
$0.0004 per message
Claude Haiku 4.5
Anthropic’s fastest low-cost model for high-volume, latency-sensitive work.
Google Gemini
$0.0004 per message
Gemini 3 Flash Preview
Google’s preview Flash model for fast multimodal reasoning and agentic workflows.
OpenAI
$0.0006 per message
GPT-4.1
Long-context OpenAI model for general-purpose production work.
OpenAI
$0.0006 per message
o3
OpenAI reasoning model for harder analytical and planning workloads.
OpenAI
$0.0008 per message
GPT-5
OpenAI's current flagship general-purpose model for complex work.
Google Gemini
$0.0010 per message
Gemini 3.1 Pro Preview
Google's latest Gemini Pro preview for complex multimodal reasoning workloads.
Anthropic
$0.0012 per message
Claude Sonnet 4.6
Anthropic's current balanced production model for everyday writing, analysis, and automation.
Anthropic
$0.0012 per message
Claude Sonnet 4.5
Previous Sonnet tier that remains useful as a stable comparison point.
OpenAI
$0.0012 per message
GPT-5.4
OpenAI's latest flagship general-purpose model for complex work.
Google Gemini
$0.0014 per message
Gemini 2.5 Pro
Current Gemini Pro model for high-capability multimodal reasoning and coding.
Anthropic
$0.0020 per message
Claude Opus 4.6
Anthropic's current flagship model for the hardest reasoning, coding, and agentic tasks.
OpenAI
$0.0024 per message
GPT-5.5
OpenAI’s newest frontier model for the most complex professional work.
Anthropic
$0.0060 per message
Claude Opus 4.1
Anthropic's current flagship model for complex reasoning and analysis.
What is an LLM calculator?
An LLM calculator keeps the math simple. Drop in your prompt, choose how long you expect the model to reply, and the tool shows the total price. You can switch models to see how the cost changes, then copy the output for your next update or LLM cost reduction review.
Use it before you promise a service level, ask for more budget, or send a savings note to leadership. Pair it with our latest cost guides for context your stakeholders can read quickly.
How to estimate LLM costs
- 1. Gather real prompts. Collect a mix of short and long prompts from the workflows you want to support.
- 2. Count tokens. Paste those prompts into the calculator or use our batching tips to plan efficient runs.
- 3. Compare providers. Switch between GPT-5, Claude Sonnet 4.6, Gemini 3 Flash Preview, and the latest Llama and DeepSeek models to see where spend and quality meet your goals.
- 4. Share the numbers. Export the results or copy-paste the summary into your finance or product review.
Tokens vs. requests explained
- Tokens are small chunks of text. Providers bill you for the tokens you send in and the tokens they send back.
- Requests are the API calls. Each request usually has a minimum cost plus the token charges.
- Why it matters: Fewer tokens per request means lower spend and faster responses, but batching work into fewer requests can also cut overhead fees.
Need help cut your AI/LLM cost? Our experts are standing by, Contact us here .
FAQ
What is an LLM calculator?
An LLM calculator estimates the per-request cost of running a prompt against a specific model by multiplying input and output tokens by the published rate.
How do I estimate LLM costs?
Gather real prompts, count tokens, compare across providers, and use a calculator to convert per-1M-token rates into per-message cost.
What is the difference between tokens and requests?
Tokens are small chunks of text billed per use. Requests are API calls; each request includes input and output tokens charged separately.