A side-by-side comparison of large language models on the three axes that actually drive a decision: context window size, input and output pricing, and what each one can do.
Compare 25 major LLM models side by side: context windows, pricing, capabilities, and release dates. Filter by provider, tier, or capability. Prices are per 1M tokens.
Provider
Tier
Capability
| Model ↕ | Provider | Context ↕ | Max Out | Input $/1M ↑ | Output $/1M ↕ | Capabilities | Released ↕ |
|---|---|---|---|---|---|---|---|
| Llama 4 Scout | Meta | 512K | 16K | Free | Free | visionopen | 2025-04 |
| Llama 4 Maverick | Meta | 1.0M | 16K | Free | Free | visionopen | 2025-04 |
| GPT-4.1 nano | OpenAI | 1.0M | 33K | $0.10 | $0.40 | vision | 2025-04 |
| Gemini 2.0 Flash | 1.0M | 8K | $0.10 | $0.40 | vision | 2025-02 | |
| GPT-4o mini | OpenAI | 128K | 16K | $0.15 | $0.60 | vision | 2024-07 |
| Gemini 2.5 Flash | 1.0M | 66K | $0.15 | $0.60 | reasonvision | 2025-04 | |
| DeepSeek V3 | DeepSeek | 128K | 8K | $0.27 | $1.10 | open | 2024-12 |
| Grok 3 mini | xAI | 131K | 16K | $0.30 | $0.50 | reason | 2025-03 |
| GPT-4.1 mini | OpenAI | 1.0M | 33K | $0.40 | $1.60 | vision | 2025-04 |
| DeepSeek R1 | DeepSeek | 128K | 8K | $0.55 | $2.19 | reasonopen | 2025-01 |
| Claude Haiku 4.5 | Anthropic | 200K | 32K | $0.80 | $4.00 | reasonvision | 2025-10 |
| Claude 3.5 Haiku | Anthropic | 200K | 8K | $0.80 | $4.00 | vision | 2024-10 |
| o4-mini | OpenAI | 200K | 100K | $1.10 | $4.40 | reasonvision | 2025-04 |
| Gemini 2.5 Pro | 1.0M | 66K | $1.25 | $10.00 | reasonvision | 2025-03 | |
| GPT-5.3 Codex | OpenAI | 400K | 128K | $1.75 | $14.00 | reasonvision | 2026-02 |
| GPT-4.1 | OpenAI | 1.0M | 33K | $2.00 | $8.00 | vision | 2025-04 |
| o3 | OpenAI | 200K | 100K | $2.00 | $8.00 | reasonvision | 2025-04 |
| Mistral Large | Mistral | 128K | 8K | $2.00 | $6.00 | vision | 2024-11 |
| GPT-4o | OpenAI | 128K | 16K | $2.50 | $10.00 | vision | 2024-05 |
| Claude Sonnet 4.6 | Anthropic | 1.0M | 128K | $3.00 | $15.00 | reasonvision | 2026-02 |
| Claude 4 Sonnet | Anthropic | 200K | 64K | $3.00 | $15.00 | reasonvision | 2025-05 |
| Claude 3.5 Sonnet | Anthropic | 200K | 8K | $3.00 | $15.00 | vision | 2024-10 |
| Grok 3 | xAI | 131K | 16K | $3.00 | $15.00 | reasonvision | 2025-02 |
| Claude Opus 4.7 | Anthropic | 1.0M | 128K | $15.00 | $75.00 | reasonvision | 2026-04 |
| Claude 4 Opus | Anthropic | 200K | 32K | $15.00 | $75.00 | reasonvision | 2025-06 |
Showing 25 of 25 models. Prices as of early 2026 and may change. Open-source models are free to self-host; hosted API pricing varies.
Model selection is usually framed as which is best and is nearly always actually which is cheapest that clears the bar. Most production workloads are classification, extraction, summarization, or routing, and the frontier model is overkill for all four. The interesting comparison is between the smallest model that passes your evals and the one you defaulted to.
Context window deserves scepticism too. Advertised windows are the maximum the model accepts, not the range over which it reliably attends. Recall degrades toward the middle of very long contexts, so a large window is not a substitute for retrieval.