Support bot · 1M user msgs / mo
~800 tokens in + 250 out per message
- GPT-4o mini~$270
- Claude Haiku~$1,640
- Gemini Flash~$180
- DeepSeek Chat~$490
Radar
Snapshot 2026-07 - USD per 1M tokens unless noted. Built for budgeting and architecture choices, not a live price feed.
15 tiers
| Model | Input / 1M | Output / 1M | Context | Rate limits |
|---|---|---|---|---|
GPT-4oOpenAIfrontier Default multimodal workhorse. Watch image token billing separately. Official pricing → | $2.50 | $10.00 | 128K | Tiered RPM/TPM by usage tier; new orgs start low |
GPT-4o miniOpenAIcheap Best OpenAI default for classification, routing, and high-QPS assistants. Official pricing → | $0.15 | $0.60 | 128K | Higher practical throughput than frontier SKUs for most tiers |
OpenAI o1OpenAIreasoning Output can include hidden reasoning tokens - estimate 2–5× naive chat cost. Official pricing → | $15.00 | $60.00 | 200K (SKU-dependent) | Stricter than chat models; reasoning tokens inflate bills |
Claude Sonnet 4Anthropicfrontier Strong for coding agents. Enable prompt caching on stable system prompts. Official pricing → | $3.00 | $15.00 | 200K+ | Org-level rate limits; prompt caching can cut repeat input cost |
Claude HaikuAnthropiccheap Use as a fast router/extractor in front of Sonnet. Official pricing → | $0.80 | $4.00 | 200K | Usually the highest-throughput Claude tier |
Gemini 2.0 FlashGooglecheap Aggressive price/perf for long docs. Confirm AI Studio vs Vertex pricing. Official pricing → | $0.10 | $0.40 | 1M (check SKU) | Free tier + paid; long-context has separate quotas |
Gemini 1.5/2.x ProGooglefrontier Competitive when you actually use huge context; otherwise Flash often wins. Official pricing → | $1.25 | $5.00 | 1M–2M (SKU-dependent) | Long-context requests can hit stricter concurrent limits |
DeepSeek Chat / V3-classDeepSeekcheap Price pressure king. Cache / off-peak discounts sometimes apply - check current page. Official pricing → | $0.27 | $1.10 | 64K–128K | API limits vary; watch region latency from SEA |
DeepSeek-R1DeepSeekreasoning API is cheap vs o1; self-host shifts cost to GPUs. Count output tokens carefully. Official pricing → | $0.55 | $2.19 | 64K–128K | Long CoT outputs chew TPM fast |
Mistral SmallMistralcheap Good budget European vendor lane for chat/classify. Official pricing → | $0.10 | $0.30 | 32K+ | Platform tiers; EU-friendly option |
Mistral LargeMistralfrontier Useful when you want a non-US/CN frontier API with solid coding. Official pricing → | $2.00 | $6.00 | 128K+ | Lower than US hyperscalers for some tiers |
text-embedding-3-largeOpenAIembed Output N/A. Dimension truncation can reduce storage cost, not API input cost. Official pricing → | $0.13 | n/a | 8K input | Usually high TPM; batch embeddings when possible |
text-embedding-3-smallOpenAIembed Default cheap RAG embedder for many prototypes. Official pricing → | $0.02 | n/a | 8K input | Very high practical throughput |
Llama 70B-class (self-host)Meta weights / your GPUselfhost Rough all-in: often $0.3–$2 / 1M tokens depending on utilization, quantization, and idle GPUs. Idle hardware dominates cost. Official pricing → | Hardware-bound | Hardware-bound | 128K (model-dependent) | Limited by your GPUs / queue - you own the quota |
DeepSeek-R1 (self-host)DeepSeek weights / your GPUselfhost API is usually cheaper until you have sustained high utilization or data-residency constraints. Official pricing → | Hardware-bound | Hardware-bound | 64K–128K | Long reasoning traces = fewer concurrent users per GPU |
~800 tokens in + 250 out per message
~3K tokens in (context) + 400 out
Input-only embedding job