Get early access Open the demo

Every model, one rate card.

Context windows, blended rates and what each model is actually good at. New releases land within days of general availability and never change your price.

ModelProviderTierContextPer 1M tokensBest at
GPT-5.2 OpenAI Frontier 400K $9.00 General reasoning and tool use
GPT-5 mini OpenAI Fast 400K $0.60 High-volume everyday work
o4-pro OpenAI Reasoning 200K $24.00 Long-horizon problem solving
Claude Opus 4.5 Anthropic Frontier 500K $11.25 Long documents and careful analysis
Claude Sonnet 4.5 Anthropic Balanced 500K $3.75 Coding and agentic work
Claude Haiku 4.5 Anthropic Fast 200K $0.90 Classification and extraction
Gemini 3 Pro Google Frontier 2M $8.40 Huge context, video and audio
Gemini 3 Flash Google Fast 1M $0.45 Cheapest capable multimodal
Llama 4 Maverick Meta Balanced 1M $0.75 Open weights you can self-host later
Llama 4 Scout Meta Fast 10M $0.35 Enormous context at low cost
Mistral Large 3 Mistral Balanced 256K $2.40 European hosting, multilingual
Codestral Mistral Code 256K $0.90 Fill-in-the-middle completion
Grok 4 xAI Frontier 256K $6.00 Real-time search grounding
DeepSeek V4 DeepSeek Balanced 256K $0.84 Strong maths at low cost
DeepSeek R2 DeepSeek Reasoning 128K $2.20 Visible chain of thought
Sonar Pro Perplexity Search 200K $4.50 Live web research with citations
Sonar Perplexity Search 128K $1.00 Fast grounded answers
Sonar Reasoning Perplexity Reasoning 128K $3.00 Multi-step research with sources

Rates are blended per million tokens, inclusive of routing, and apply identically on every paid plan. Frontier and reasoning tiers are excluded from the free plan.

Tiers

What the tiers mean.

Fast

Sub-second first token, cents per thousand messages. Classification, extraction, rewriting, routing decisions, anything you run in bulk.

Balanced

The working default. Coding, drafting, analysis with a real context window. Where most Auto-Route traffic lands.

Frontier

Best available judgement. Long documents, ambiguous briefs, work you would otherwise hand to a senior person.

Reasoning

Deliberate multi-step problem solving. Slower and dearer per call, decisively better on maths, proofs and planning.

Code

Fill-in-the-middle completion tuned for editors and agents rather than conversation.

Let Auto-Route choose

Or skip the taxonomy entirely. Auto-Route reads each message and picks the cheapest tier that clears the bar.

Try them against each other.

Side-by-side runs one prompt through four models at once, with cost and latency under every column.

Open side-by-side