Every model, one rate card.
Context windows, blended rates and what each model is actually good at. New releases land within days of general availability and never change your price.
| Model | Provider | Tier | Context | Per 1M tokens | Best at |
|---|---|---|---|---|---|
| GPT-5.2 | OpenAI | Frontier | 400K | $9.00 | General reasoning and tool use |
| GPT-5 mini | OpenAI | Fast | 400K | $0.60 | High-volume everyday work |
| o4-pro | OpenAI | Reasoning | 200K | $24.00 | Long-horizon problem solving |
| Claude Opus 4.5 | Anthropic | Frontier | 500K | $11.25 | Long documents and careful analysis |
| Claude Sonnet 4.5 | Anthropic | Balanced | 500K | $3.75 | Coding and agentic work |
| Claude Haiku 4.5 | Anthropic | Fast | 200K | $0.90 | Classification and extraction |
| Gemini 3 Pro | Frontier | 2M | $8.40 | Huge context, video and audio | |
| Gemini 3 Flash | Fast | 1M | $0.45 | Cheapest capable multimodal | |
| Llama 4 Maverick | Meta | Balanced | 1M | $0.75 | Open weights you can self-host later |
| Llama 4 Scout | Meta | Fast | 10M | $0.35 | Enormous context at low cost |
| Mistral Large 3 | Mistral | Balanced | 256K | $2.40 | European hosting, multilingual |
| Codestral | Mistral | Code | 256K | $0.90 | Fill-in-the-middle completion |
| Grok 4 | xAI | Frontier | 256K | $6.00 | Real-time search grounding |
| DeepSeek V4 | DeepSeek | Balanced | 256K | $0.84 | Strong maths at low cost |
| DeepSeek R2 | DeepSeek | Reasoning | 128K | $2.20 | Visible chain of thought |
| Sonar Pro | Perplexity | Search | 200K | $4.50 | Live web research with citations |
| Sonar | Perplexity | Search | 128K | $1.00 | Fast grounded answers |
| Sonar Reasoning | Perplexity | Reasoning | 128K | $3.00 | Multi-step research with sources |
Rates are blended per million tokens, inclusive of routing, and apply identically on every paid plan. Frontier and reasoning tiers are excluded from the free plan.
What the tiers mean.
Fast
Sub-second first token, cents per thousand messages. Classification, extraction, rewriting, routing decisions, anything you run in bulk.
Balanced
The working default. Coding, drafting, analysis with a real context window. Where most Auto-Route traffic lands.
Frontier
Best available judgement. Long documents, ambiguous briefs, work you would otherwise hand to a senior person.
Reasoning
Deliberate multi-step problem solving. Slower and dearer per call, decisively better on maths, proofs and planning.
Code
Fill-in-the-middle completion tuned for editors and agents rather than conversation.
Let Auto-Route choose
Or skip the taxonomy entirely. Auto-Route reads each message and picks the cheapest tier that clears the bar.
Try them against each other.
Side-by-side runs one prompt through four models at once, with cost and latency under every column.
Open side-by-side
