LLM API Pricing Calculator — Compare AI Model Costs
Enter your monthly input and output tokens, pick a workload, and instantly compare monthly and annual API costs across 10 providers including OpenAI, Anthropic, Google, DeepSeek, Mistral and xAI — 20 models ranked cheapest first.
Conversational assistants & support bots
Amazon Nova Lite (Bedrock)
300K context · $0.06/M in · $0.24/M out
$0.18
per month
$2.16/yr
Estimated cost for 1M input + 500K output tokens per month. In/Out columns show USD per million tokens.
| # | Model | Monthly cost |
|---|---|---|
| 1 | Nova Lite (Bedrock) Amazon | $0.18 |
| 2 | Mistral Small Mistral | $0.25 |
| 3 | Gemini 2.0 Flash | $0.30 |
| 4 | GPT-4o-mini OpenAI | $0.45 |
| 5 | Gemini 2.5 Flash | $0.45 |
| 6 | Llama 4 Maverick (via Groq) Meta | $0.50 |
| 7 | DeepSeek-V3 DeepSeek | $0.82 |
| 8 | Llama 3.3 70B Groq | $0.985 |
| 9 | DeepSeek-R1 DeepSeek | $1.65 |
| 10 | Nova Pro (Bedrock) Amazon | $2.40 |
| 11 | Claude Haiku 3.5 Anthropic | $2.80 |
| 12 | o4-mini OpenAI | $3.30 |
| 13 | Mistral Large Mistral | $5.00 |
| 14 | Gemini 2.5 Pro | $6.25 |
| 15 | GPT-4o OpenAI | $7.50 |
| 16 | Command R+ Cohere | $7.50 |
| 17 | Claude Sonnet 4 Anthropic | $10.50 |
| 18 | Grok 3 xAI | $10.50 |
| 19 | o3 OpenAI | $30.00 |
| 20 | Claude Opus 4 Anthropic | $52.50 |
How the calculator works
Monthly cost = (input tokens ÷ 1M × input price) + (output tokens ÷ 1M × output price). Annual cost is the monthly figure × 12.
Enter your token usage
Set your estimated monthly input and output tokens with the sliders, or type exact figures. Defaults start at 1M input and 500K output tokens.
Pick a workload
Choose Chat, Code, Vision or Reasoning. The best-value pick is filtered to models that actually support that capability.
Compare every model
See monthly and annual cost for every model ranked cheapest-first, with the best-value pick highlighted for your workload.
LLM API pricing per million tokens
Standard on-demand rates in USD per 1,000,000 tokens, grouped by provider. Last updated 2026-07-22.
| Model | Input $/M | Output $/M |
|---|---|---|
Nova Lite (Bedrock) Amazon | $0.06 | $0.24 |
Nova Pro (Bedrock) Amazon | $0.80 | $3.20 |
Claude Haiku 3.5 Anthropic | $0.80 | $4.00 |
Claude Sonnet 4 Anthropic | $3.00 | $15.00 |
Claude Opus 4 Anthropic | $15.00 | $75.00 |
Command R+ Cohere | $2.50 | $10.00 |
DeepSeek-V3 DeepSeek | $0.27 | $1.10 |
DeepSeek-R1 DeepSeek | $0.55 | $2.19 |
Gemini 2.0 Flash | $0.10 | $0.40 |
Gemini 2.5 Flash | $0.15 | $0.60 |
Gemini 2.5 Pro | $1.25 | $10.00 |
Llama 3.3 70B Groq | $0.59 | $0.79 |
Llama 4 Maverick (via Groq) Meta | $0.20 | $0.60 |
Mistral Small Mistral | $0.10 | $0.30 |
Mistral Large Mistral | $2.00 | $6.00 |
GPT-4o-mini OpenAI | $0.15 | $0.60 |
o4-mini OpenAI | $1.10 | $4.40 |
GPT-4o OpenAI | $2.50 | $10.00 |
o3 OpenAI | $10.00 | $40.00 |
Grok 3 xAI | $3.00 | $15.00 |
Frequently asked questions
How is the monthly LLM API cost calculated?
Monthly cost = (input tokens / 1,000,000) × input price + (output tokens / 1,000,000) × output price. The calculator multiplies your estimated monthly input and output tokens by each model’s per-million-token rate, then annualises the result by multiplying by 12.
What is the cheapest LLM API?
It depends on your workload and token mix. Budget models such as Gemini 2.0 Flash, Mistral Small and Amazon Nova Lite are among the lowest per-million-token rates, while reasoning-heavy models cost far more. Enter your real input/output volumes above to see the cheapest option for your specific usage.
Why do input and output tokens have different prices?
Output tokens are generated one at a time during decoding, which uses more compute than reading the input prompt. As a result, output pricing is typically 2–5× higher than input pricing for the same model. A workload that generates a lot of output is far more sensitive to the output rate.
How does the best-value pick work?
For the workload you select (Chat, Code, Vision or Reasoning), the calculator filters to models that support that capability and highlights the one with the lowest monthly cost. If no model declares that capability, it falls back to the cheapest model overall.
Are these prices accurate?
All figures are estimates compiled from publicly available provider pricing and may lag behind official rate cards. Use them for planning and comparison, then confirm on the provider’s pricing page before committing. Prices were last updated on the date shown in the comparison table.
Does the calculator include prompt caching or batch discounts?
No. Estimates use standard on-demand pay-as-you-go rates. Providers such as OpenAI, Anthropic and Google offer prompt caching, batch APIs and committed-use discounts that can cut costs substantially, so real spend may be lower than the figures shown here.