Model your monthly LLM API bill across 23 models and six providers using your own request volume and token assumptions, with input and output priced separately because they differ by several times.
Estimate your monthly AI API costs across all major providers. Set your expected usage and compare pricing for OpenAI, Anthropic, Google, DeepSeek, Mistral, and xAI models side by side.
Usage Presets
| Model | Provider | Daily | Monthly | Yearly |
|---|---|---|---|---|
| GPT-4.1 nanocheapest | OpenAI | $0.50 | $15.00 | $182.50 |
| Gemini 2.0 Flash | $0.50 | $15.00 | $182.50 | |
| GPT-4o mini | OpenAI | $0.75 | $22.50 | $273.75 |
| Gemini 2.5 Flash | $0.75 | $22.50 | $273.75 | |
| Grok 3 mini | xAI | $0.80 | $24.00 | $292.00 |
| DeepSeek V3 | DeepSeek | $1.37 | $41.10 | $500.05 |
| GPT-4.1 mini | OpenAI | $2.00 | $60.00 | $730.00 |
| DeepSeek R1 | DeepSeek | $2.74 | $82.20 | $1,000 |
| Claude Haiku 4.5 | Anthropic | $4.80 | $144.00 | $1,752 |
| Claude 3.5 Haiku | Anthropic | $4.80 | $144.00 | $1,752 |
| o4-mini | OpenAI | $5.50 | $165.00 | $2,008 |
| Mistral Large | Mistral | $8.00 | $240.00 | $2,920 |
| GPT-4.1 | OpenAI | $10.00 | $300.00 | $3,650 |
| o3 | OpenAI | $10.00 | $300.00 | $3,650 |
| Gemini 2.5 Pro | $11.25 | $337.50 | $4,106 | |
| GPT-4o | OpenAI | $12.50 | $375.00 | $4,563 |
| GPT-5.3 Codex | OpenAI | $15.75 | $472.50 | $5,749 |
| Claude Sonnet 4.6 | Anthropic | $18.00 | $540.00 | $6,570 |
| Claude 4 Sonnet | Anthropic | $18.00 | $540.00 | $6,570 |
| Claude 3.5 Sonnet | Anthropic | $18.00 | $540.00 | $6,570 |
| Grok 3 | xAI | $18.00 | $540.00 | $6,570 |
| Claude Opus 4.7 | Anthropic | $90.00 | $2,700 | $32,850 |
| Claude 4 Opus | Anthropic | $90.00 | $2,700 | $32,850 |
Prices as of early 2026 and may change. Open-weight models (Llama 4, etc.) are free to self-host but incur compute costs not shown here. Batch API pricing and cached prompt discounts not included.
The mistake that produces shocking bills is estimating from input length alone. Output tokens typically cost three to five times more than input, so a workload generating long responses costs far more than its prompt size suggests. Model both sides separately or the estimate will be wrong by a multiple, not a margin.
The second-order effect is retries. A pipeline that retries failed calls twice has three times the token spend on the failure path, and agentic loops that re-read context on every turn compound it further. If you are running agents rather than single calls, the Agent Loop Cost Predictor models that shape specifically.