Support bot · 1M user msgs / mo
~800 tokens in + 250 out per message
- GPT-4o mini~$270
- GPT-5.6 Luna~$460
- Claude Haiku 4.5~$2,050
- DeepSeek V4 Flash (peak)~$682
Radar
Snapshot 2026-09 - USD per 1M tokens unless noted. Built for budgeting and architecture choices, not a live price feed.
19 tiers
| Model | Input / 1M | Output / 1M | Context | Rate limits |
|---|---|---|---|---|
GPT-6 AstraOpenAIfrontier Current OpenAI flagship (gpt-6-astra). Short-context standard rates shown. Cyber-sensitive capabilities are gated. Official pricing → | $10.00 | $50.00 | 1M | Rolling out; Fast mode is 2x Standard price. Long-context (>272K input) doubles input / 1.5x output. |
GPT-5.6 SolOpenAIfrontier Promotional $4/$20 at least through 2026-11-21. Launch card was $5/$30 - treat that as the post-promo fallback until OpenAI says otherwise. Official pricing → | $4.00 | $20.00 | 1M (SKU-dependent) | Tiered RPM/TPM by usage tier; long-context rates are higher |
GPT-5.6 TerraOpenAIfrontier Balanced GPT-5.6 workhorse after the July 30 cut. Prefer over Sol unless you need the top 5.6 SKU. Official pricing → | $2.00 | $12.00 | 1M (SKU-dependent) | Same tier system as Sol; usually the everyday GPT-5.6 default |
GPT-5.6 LunaOpenAIcheap Current GPT-5.6 budget lane after the July 30 cut. Still more expensive than GPT-4o mini on raw tokens. Official pricing → | $0.20 | $1.20 | 1M (SKU-dependent) | Higher practical throughput than Sol/Terra for most tiers |
GPT-4o miniOpenAIcheap Still a strong OpenAI budget default for classification, routing, and high-QPS assistants. Official pricing → | $0.15 | $0.60 | 128K | Higher practical throughput than frontier SKUs for most tiers |
OpenAI o4-miniOpenAIreasoning Compact o-series reasoning. Succeeded in the lineup by GPT-5 mini / Luna for some workloads - verify current default. Official pricing → | $1.10 | $4.40 | 200K | Reasoning tokens inflate bills; stricter than chat SKUs |
Claude Opus 5Anthropicfrontier Current Opus-class production default. Fable 5 remains the higher $10/$50 premium tier. Official pricing → | $5.00 | $25.00 | 1M | Org-level rate limits; prompt caching can cut repeat input cost |
Claude Sonnet 5Anthropicfrontier Introductory $2/$10 is now the standard rate. The planned bump to $3/$15 did not happen. Official pricing → | $2.00 | $10.00 | 1M | Org-level rate limits; prompt caching available |
Claude Haiku 4.5Anthropiccheap Use as a fast router/extractor in front of Sonnet/Opus. Official pricing → | $1.00 | $5.00 | 200K | Usually the highest-throughput Claude tier |
Gemini 3.5 FlashGooglecheap Current Flash default. Faster than 2.5 Flash but no longer the cheap Google SKU - Flash-Lite is cheaper. Official pricing → | $1.50 | $9.00 | 1M | Free tier + paid; output price includes thinking tokens |
Gemini 3.1 ProGooglefrontier Rates shown for prompts <=200K tokens. Gemini 3.5 Pro is still 'coming soon' with no public price. Official pricing → | $2.00 | $12.00 | 1M | Preview ID gemini-3.1-pro-preview; long prompts (>200K) use $4 / $18 |
DeepSeek V4 FlashDeepSeekcheap Peak cache-miss rates shown (effective 2026-08-16). Off-peak is $0.22 / $0.66. Cache hits are much cheaper. Official pricing → | $0.44 | $1.32 | 1M | Peak hours 01:00-04:00 and 06:00-10:00 UTC; off-peak is half |
DeepSeek V4 ProDeepSeekreasoning Peak cache-miss rates shown. Off-peak is $0.66 / $1.98. Thinking mode is default - count output carefully. Official pricing → | $1.32 | $3.96 | 1M | Same peak window as Flash; long CoT chews TPM |
Mistral Small 4Mistralcheap Current high-volume Mistral budget lane for chat/classify. Official pricing → | $0.15 | $0.60 | 128K | Platform tiers; EU-friendly option |
Mistral Large 3Mistralfrontier Current Mistral flagship on La Plateforme (also open-weight). Official pricing → | $0.50 | $1.50 | 256K | Lower than US hyperscalers for many tiers |
text-embedding-3-largeOpenAIembed Output N/A. Dimension truncation can reduce storage cost, not API input cost. Official pricing → | $0.13 | n/a | 8K input | Usually high TPM; batch embeddings when possible |
text-embedding-3-smallOpenAIembed Default cheap RAG embedder for many prototypes. Official pricing → | $0.02 | n/a | 8K input | Very high practical throughput |
Llama 70B-class (self-host)Meta weights / your GPUselfhost Rough all-in: often $0.3-$2 / 1M tokens depending on utilization, quantization, and idle GPUs. Idle hardware dominates cost. Official pricing → | Hardware-bound | Hardware-bound | 128K (model-dependent) | Limited by your GPUs / queue - you own the quota |
DeepSeek V4-class (self-host)DeepSeek weights / your GPUselfhost API is usually cheaper until you have sustained high utilization or data-residency constraints. Peak API rates make self-host more interesting than in July. Official pricing → | Hardware-bound | Hardware-bound | 1M (model-dependent) | Long reasoning traces = fewer concurrent users per GPU |
~800 tokens in + 250 out per message
~3K tokens in (context) + 400 out
Input-only embedding job