AI Cost Estimator
Workload Configuration
Reduces input cost ~90% for cached portion
Monthly Cost Comparison
Detailed Cost Breakdown
| Model ↕ | Provider ↕ | Input Cost/Day ↕ | Output Cost/Day ↕ | Total/Day ↕ | Monthly ↑ |
|---|---|---|---|---|---|
| DeepSeek V4 FlashCHEAPEST | DeepSeek | $0.070 | $0.140 | $0.210 | $6.30 |
| GPT-5-nano | OpenAI | $0.025 | $0.200 | $0.225 | $6.75 |
| Llama 4 Scout (Groq) | Meta/Groq | $0.055 | $0.170 | $0.225 | $6.75 |
| GPT-4.1-nano | OpenAI | $0.050 | $0.200 | $0.250 | $7.50 |
| Gemini 2.5 Flash-Lite | $0.050 | $0.200 | $0.250 | $7.50 | |
| Mistral Small 4 | Mistral | $0.075 | $0.300 | $0.375 | $11.25 |
| DeepSeek V4 Pro | DeepSeek | $0.217 | $0.435 | $0.652 | $19.57 |
| GPT-5.4 Nano | OpenAI | $0.100 | $0.625 | $0.725 | $21.75 |
| Gemini 3.1 Flash-Lite | $0.125 | $0.750 | $0.875 | $26.25 | |
| GPT-4.1-mini | OpenAI | $0.200 | $0.800 | $1.00 | $30.00 |
| Mistral Large 3 | Mistral | $0.250 | $0.750 | $1.00 | $30.00 |
| GPT-5-mini | OpenAI | $0.125 | $1.00 | $1.13 | $33.75 |
| Mistral Medium 3.1 | Mistral | $0.200 | $1.00 | $1.20 | $36.00 |
| Gemini 2.5 Flash | $0.150 | $1.25 | $1.40 | $42.00 | |
| Grok Build 0.1 | xAI | $0.500 | $1.00 | $1.50 | $45.00 |
| Gemini 3 Flash | $0.250 | $1.50 | $1.75 | $52.50 | |
| Grok 4.3 | xAI | $0.625 | $1.25 | $1.88 | $56.25 |
| Grok 4.20 | xAI | $0.625 | $1.25 | $1.88 | $56.25 |
| GPT-5.4 Mini | OpenAI | $0.375 | $2.25 | $2.63 | $78.75 |
| o4-mini | OpenAI | $0.550 | $2.20 | $2.75 | $82.50 |
| o3-mini | OpenAI | $0.550 | $2.20 | $2.75 | $82.50 |
| Claude Haiku 4.5 | Anthropic | $0.500 | $2.50 | $3.00 | $90.00 |
| Mistral Medium 3.5 | Mistral | $0.750 | $3.75 | $4.50 | $135.00 |
| o3 | OpenAI | $1.00 | $4.00 | $5.00 | $150.00 |
| GPT-4.1 | OpenAI | $1.00 | $4.00 | $5.00 | $150.00 |
| Gemini 3.5 Flash | $0.750 | $4.50 | $5.25 | $157.50 | |
| GPT-5.1 | OpenAI | $0.625 | $5.00 | $5.63 | $168.75 |
| GPT-5 | OpenAI | $0.625 | $5.00 | $5.63 | $168.75 |
| Gemini 2.5 Pro | $0.625 | $5.00 | $5.63 | $168.75 | |
| Gemini 3.1 Pro | $1.00 | $6.00 | $7.00 | $210.00 | |
| GPT-5.3-Codex | OpenAI | $0.875 | $7.00 | $7.88 | $236.25 |
| GPT-5.2 | OpenAI | $0.875 | $7.00 | $7.88 | $236.25 |
| GPT-5.4 | OpenAI | $1.25 | $7.50 | $8.75 | $262.50 |
| Claude Sonnet 4.6 | Anthropic | $1.50 | $7.50 | $9.00 | $270.00 |
| Claude Sonnet 4.5 | Anthropic | $1.50 | $7.50 | $9.00 | $270.00 |
| Claude Sonnet 4 | Anthropic | $1.50 | $7.50 | $9.00 | $270.00 |
| Claude Opus 4.8 | Anthropic | $2.50 | $12.50 | $15.00 | $450.00 |
| Claude Opus 4.7 | Anthropic | $2.50 | $12.50 | $15.00 | $450.00 |
| Claude Opus 4.6 | Anthropic | $2.50 | $12.50 | $15.00 | $450.00 |
| Claude Opus 4.5 | Anthropic | $2.50 | $12.50 | $15.00 | $450.00 |
| GPT-5.5 | OpenAI | $2.50 | $15.00 | $17.50 | $525.00 |
| Claude Fable 5 | Anthropic | $5.00 | $25.00 | $30.00 | $900.00 |
| Claude Opus 4.1 | Anthropic | $7.50 | $37.50 | $45.00 | $1,350.00 |
| o3-pro | OpenAI | $10.00 | $40.00 | $50.00 | $1,500.00 |
| GPT-5.5 Pro | OpenAI | $15.00 | $90.00 | $105.00 | $3,150.00 |
| GPT-5.4 ProPRICIEST | OpenAI | $15.00 | $90.00 | $105.00 | $3,150.00 |
Cost vs. Capability Insights
Compared to the cheapest option (DeepSeek V4 Flash):
Workload Summary
Pricing data last updated: July 19, 2026. Prices per 1M tokens. Estimates are approximate and may vary with actual usage patterns.
What This Tool Does
LLM API pricing in 2026 spans from $0.14 input / $0.28 output per million tokens (DeepSeek V4 Flash) up to $30 input / $180 output (GPT-5.5 Pro), across 32 current models from 7 providers. Flagship rates: GPT-5.5 $5 input / $30 output, Claude Opus 4.8 $5 input / $25 output, Gemini 3.1 Pro $2 input / $12 output. Every price in the table below is verified against the vendor's official pricing page. This estimator turns those per-token rates into monthly workload costs.
Use How to Use for execution steps and FAQ for constraints, policies, and edge cases.
Last updated:
This tool is provided as-is for convenience. Output should be verified before use in any production or critical context.
Agent Invocation
Best Path For Builders
Browser workflow
Runs instantly in the browser with private local processing and copy/export-ready output.
Browser Workflow
This tool is optimized for instant in-browser execution with local data handling. Run it here and copy/export the output directly.
/ai-cost-estimator/
For automation planning, fetch the canonical contract at /api/tool/ai-cost-estimator.json.
How to Use AI Cost Estimator
- 1
Select your AI models
Choose models you're using: GPT, Claude, Gemini, Llama, etc. The tool displays current pricing per 1M input and output tokens. Add multiple models if you're comparing or using a mix.
- 2
Estimate token counts
Input your expected monthly usage: number of requests, average input tokens (rule of thumb: 1 token ≈ 4 characters), and average output tokens. Or paste a sample prompt to auto-calculate token count.
- 3
Factor in batching and caching
Some models offer cheaper batch processing or prompt caching. Account for these if applicable. Prompt caching (e.g., Claude) reduces per-token costs for repeated inputs.
- 4
Calculate total and per-request costs
The tool shows monthly cost, per-request cost, and cost per feature/endpoint. Compare pricing across models to choose the best fit for your workload and budget.
Frequently Asked Questions
What is AI Cost Estimator?
How do I use AI Cost Estimator?
Is AI Cost Estimator free?
Does AI Cost Estimator store or send my data?
How accurate are the cost estimates?
How much do LLM APIs cost per million tokens across providers in 2026?
Current list prices per 1 million tokens for 32 models across OpenAI, Anthropic, Google, xAI, DeepSeek, Mistral, Meta/Groq. Input and output tokens are billed separately, and output is typically several times more expensive — for chat-style workloads with long answers, output cost usually dominates the bill.
| Model | Provider | Input / 1M | Output / 1M | Tier |
|---|---|---|---|---|
| GPT-5.5 | OpenAI | $5 | $30 | premium |
| GPT-5.5 Pro | OpenAI | $30 | $180 | premium |
| GPT-5.4 | OpenAI | $2.50 | $15 | premium |
| GPT-5.4 Pro | OpenAI | $30 | $180 | premium |
| GPT-5.4 Mini | OpenAI | $0.75 | $4.50 | mid |
| GPT-5.4 Nano | OpenAI | $0.20 | $1.25 | budget |
| GPT-5.3-Codex | OpenAI | $1.75 | $14 | premium |
| Claude Fable 5 | Anthropic | $10 | $50 | premium |
| Claude Opus 4.8 | Anthropic | $5 | $25 | premium |
| Claude Opus 4.7 | Anthropic | $5 | $25 | premium |
| Claude Opus 4.6 | Anthropic | $5 | $25 | premium |
| Claude Opus 4.5 | Anthropic | $5 | $25 | premium |
| Claude Opus 4.1 | Anthropic | $15 | $75 | premium |
| Claude Sonnet 4.6 | Anthropic | $3 | $15 | mid |
| Claude Sonnet 4.5 | Anthropic | $3 | $15 | mid |
| Claude Haiku 4.5 | Anthropic | $1 | $5 | budget |
| Gemini 3.1 Pro | $2 | $12 | premium | |
| Gemini 3.5 Flash | $1.50 | $9 | premium | |
| Gemini 3 Flash | $0.50 | $3 | mid | |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | budget | |
| Gemini 2.5 Pro | $1.25 | $10 | premium | |
| Gemini 2.5 Flash | $0.30 | $2.50 | mid | |
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | budget | |
| Grok 4.3 | xAI | $1.25 | $2.50 | premium |
| Grok 4.20 | xAI | $1.25 | $2.50 | premium |
| Grok Build 0.1 | xAI | $1 | $2 | mid |
| DeepSeek V4 Flash | DeepSeek | $0.14 | $0.28 | budget |
| DeepSeek V4 Pro | DeepSeek | $0.43 | $0.87 | mid |
| Mistral Medium 3.5 | Mistral | $1.50 | $7.50 | premium |
| Mistral Large 3 | Mistral | $0.50 | $1.50 | mid |
| Mistral Small 4 | Mistral | $0.15 | $0.60 | budget |
| Llama 4 Scout (Groq) | Meta/Groq | $0.11 | $0.34 | budget |
Sources: official pricing pages of OpenAI, Anthropic, Google AI, xAI, DeepSeek, Mistral, and Groq. Lineup verified 2026-06-21; data refreshed July 19, 2026. Prices in USD; caching, batch, and long-context tiers can change the effective rate.
Which provider is cheapest per million tokens?
Among current frontier-lab models, DeepSeek V4 Flash is the cheapest in this dataset at $0.14 input / $0.28 output per million tokens. Open-weight models served via Groq undercut most closed models for high-volume pipelines. At the premium end, GPT-5.5 Pro tops the table at $30 input / $180 output. The gap between the cheapest and the most expensive model is well over two orders of magnitude, which is why model routing — sending easy requests to budget tiers — is usually the single largest cost lever.
How do I estimate a real monthly bill from these rates?
- Estimate average input and output tokens per request. As a rule of thumb, 1,000 English words is roughly 750 tokens on GPT-family tokenizers.
- Multiply by requests per day and by 30, then apply the per-million rates from the table separately for input and output.
- Account for prompt caching where supported — repeated system prompts and context can bill at a fraction of the input rate.
- Use the estimator above to run these numbers per model and compare providers side by side; the calculation runs client-side.