← All Radar guides

Radar

Pricing

Snapshot 2026-09 - USD per 1M tokens unless noted. Built for budgeting and architecture choices, not a live price feed.

19 tiers

ModelInput / 1MOutput / 1MContextRate limits
GPT-6 AstraOpenAIfrontier

Current OpenAI flagship (gpt-6-astra). Short-context standard rates shown. Cyber-sensitive capabilities are gated.

Official pricing →
$10.00$50.001MRolling out; Fast mode is 2x Standard price. Long-context (>272K input) doubles input / 1.5x output.
GPT-5.6 SolOpenAIfrontier

Promotional $4/$20 at least through 2026-11-21. Launch card was $5/$30 - treat that as the post-promo fallback until OpenAI says otherwise.

Official pricing →
$4.00$20.001M (SKU-dependent)Tiered RPM/TPM by usage tier; long-context rates are higher
GPT-5.6 TerraOpenAIfrontier

Balanced GPT-5.6 workhorse after the July 30 cut. Prefer over Sol unless you need the top 5.6 SKU.

Official pricing →
$2.00$12.001M (SKU-dependent)Same tier system as Sol; usually the everyday GPT-5.6 default
GPT-5.6 LunaOpenAIcheap

Current GPT-5.6 budget lane after the July 30 cut. Still more expensive than GPT-4o mini on raw tokens.

Official pricing →
$0.20$1.201M (SKU-dependent)Higher practical throughput than Sol/Terra for most tiers
GPT-4o miniOpenAIcheap

Still a strong OpenAI budget default for classification, routing, and high-QPS assistants.

Official pricing →
$0.15$0.60128KHigher practical throughput than frontier SKUs for most tiers
OpenAI o4-miniOpenAIreasoning

Compact o-series reasoning. Succeeded in the lineup by GPT-5 mini / Luna for some workloads - verify current default.

Official pricing →
$1.10$4.40200KReasoning tokens inflate bills; stricter than chat SKUs
Claude Opus 5Anthropicfrontier

Current Opus-class production default. Fable 5 remains the higher $10/$50 premium tier.

Official pricing →
$5.00$25.001MOrg-level rate limits; prompt caching can cut repeat input cost
Claude Sonnet 5Anthropicfrontier

Introductory $2/$10 is now the standard rate. The planned bump to $3/$15 did not happen.

Official pricing →
$2.00$10.001MOrg-level rate limits; prompt caching available
Claude Haiku 4.5Anthropiccheap

Use as a fast router/extractor in front of Sonnet/Opus.

Official pricing →
$1.00$5.00200KUsually the highest-throughput Claude tier
Gemini 3.5 FlashGooglecheap

Current Flash default. Faster than 2.5 Flash but no longer the cheap Google SKU - Flash-Lite is cheaper.

Official pricing →
$1.50$9.001MFree tier + paid; output price includes thinking tokens
Gemini 3.1 ProGooglefrontier

Rates shown for prompts <=200K tokens. Gemini 3.5 Pro is still 'coming soon' with no public price.

Official pricing →
$2.00$12.001MPreview ID gemini-3.1-pro-preview; long prompts (>200K) use $4 / $18
DeepSeek V4 FlashDeepSeekcheap

Peak cache-miss rates shown (effective 2026-08-16). Off-peak is $0.22 / $0.66. Cache hits are much cheaper.

Official pricing →
$0.44$1.321MPeak hours 01:00-04:00 and 06:00-10:00 UTC; off-peak is half
DeepSeek V4 ProDeepSeekreasoning

Peak cache-miss rates shown. Off-peak is $0.66 / $1.98. Thinking mode is default - count output carefully.

Official pricing →
$1.32$3.961MSame peak window as Flash; long CoT chews TPM
Mistral Small 4Mistralcheap

Current high-volume Mistral budget lane for chat/classify.

Official pricing →
$0.15$0.60128KPlatform tiers; EU-friendly option
Mistral Large 3Mistralfrontier

Current Mistral flagship on La Plateforme (also open-weight).

Official pricing →
$0.50$1.50256KLower than US hyperscalers for many tiers
text-embedding-3-largeOpenAIembed

Output N/A. Dimension truncation can reduce storage cost, not API input cost.

Official pricing →
$0.13n/a8K inputUsually high TPM; batch embeddings when possible
text-embedding-3-smallOpenAIembed

Default cheap RAG embedder for many prototypes.

Official pricing →
$0.02n/a8K inputVery high practical throughput
Llama 70B-class (self-host)Meta weights / your GPUselfhost

Rough all-in: often $0.3-$2 / 1M tokens depending on utilization, quantization, and idle GPUs. Idle hardware dominates cost.

Official pricing →
Hardware-boundHardware-bound128K (model-dependent)Limited by your GPUs / queue - you own the quota
DeepSeek V4-class (self-host)DeepSeek weights / your GPUselfhost

API is usually cheaper until you have sustained high utilization or data-residency constraints. Peak API rates make self-host more interesting than in July.

Official pricing →
Hardware-boundHardware-bound1M (model-dependent)Long reasoning traces = fewer concurrent users per GPU

Rough monthly scenarios

Support bot · 1M user msgs / mo

~800 tokens in + 250 out per message

  • GPT-4o mini~$270
  • GPT-5.6 Luna~$460
  • Claude Haiku 4.5~$2,050
  • DeepSeek V4 Flash (peak)~$682

RAG answers · 200K queries / mo

~3K tokens in (context) + 400 out

  • GPT-5.6 Terra~$2,160
  • Claude Sonnet 5~$2,000
  • Gemini 3.5 Flash~$1,620
  • DeepSeek V4 Flash (peak)~$370

Embed a 50M-token corpus (once)

Input-only embedding job

  • emb-3-small~$1.00
  • emb-3-large~$6.50
  • Self-host BGEGPU hours, not API $

Rate-limit & cost tips