TokenCost.io
Live Pricing: Claude Sonnet 5, GPT-5.6 Sol/Luna, DeepSeek V4 & Gemini 3.7 (August 2026)

AI API & Cloud Compute Cost Simulator

Accurately forecast your monthly LLM API bills, simulate prompt caching savings, and compare database & serverless egress across all frontier providers.

15+ Frontier Models
Up to 90% Cache Savings
Sub-second Client Simulation
Full Stack Backend TCO
Quick Workload Presets

Workload Parameters

250M tokens/mo
req
1k100k1M5M10M+
tok
1002,000 (RAG)8,000 (Agent)32,000 (Doc QA)64k+
tok
50 (Short JSON)500 (Chat response)2,000 (Code)4,000 (Essay)8k+
50% Cache Hit
10% Light reuse50% Standard RAG75% Agent Loop90% High caching

Monthly Economics Summary

Total Tokens250M200M In / 50M Out
Cheapest ModelDeepSeek V4 Flash$7.80 /mo
Top ReasoningClaude Fable 5$3,600.00 /mo
Prompt Caching Active (50% Cache Hit)Saves up to 90% on input tokens for Claude, DeepSeek & Gemini
DeepSeek V4 FlashDeepSeekLowest Cost

Disruptive sub-cent pricing at $0.03/1M input and 1.31M token context window.

Context: 1.3MSpeed: ~180 tpsMMLU: 89.4%
Input: $0.03/1MOutput: $0.09/1M
Input: $3.30Output: $4.50Save $2.70
$7.80/mo
$0.078 per 1k reqs
Get Keys
Llama 4 ScoutMetaSpeed Champion

Ultra-fast lightweight Llama 4 model running with 1.31M context window at 290 tokens/sec.

Context: 1.3MSpeed: ~290 tpsMMLU: 87.5%
Input: $0.11/1MOutput: $0.34/1M
Input: $22.00Output: $17.00
$39.00/mo
$0.390 per 1k reqs
Get Keys
DeepSeek V3.2DeepSeek

Upgraded V3 architecture delivering reliable function calling and 163k context.

Context: 163.8kSpeed: ~80 tpsMMLU: 90.2%
Input: $0.26/1MOutput: $0.38/1M
Input: $28.60Output: $19.00Save $23.40
$47.60/mo
$0.476 per 1k reqs
Get Keys
Qwen 3.8 FlashQwenSpeed Champion

Ultra-fast 1M context model running at 170 TPS at only $0.15/1M input.

Context: 1MSpeed: ~170 tpsMMLU: 88%
Input: $0.15/1MOutput: $0.47/1M
Input: $30.00Output: $23.50
$53.50/mo
$0.535 per 1k reqs
Get Keys
Mistral Small 4MistralLowest Cost

Fast, lightweight model from Mistral with 262k context at only $0.15/1M tokens.

Context: 262.1kSpeed: ~150 tpsMMLU: 84.8%
Input: $0.15/1MOutput: $0.60/1M
Input: $30.00Output: $30.00
$60.00/mo
$0.600 per 1k reqs
Get Keys
Llama 4 MaverickMetaOpen Weights King

Meta’s frontier Llama 4 flagship model with 1,048,576 token context and complete weights openness.

Context: 1MSpeed: ~220 tpsMMLU: 91.2%
Input: $0.20/1MOutput: $0.80/1M
Input: $40.00Output: $40.00
$80.00/mo
$0.800 per 1k reqs
Get Keys
GPT-5.6 LunaOpenAIBest Value

OpenAI’s fast workhorse model with 1.05M context at an ultra-low $0.20/1M input price.

Context: 1.1MSpeed: ~165 tpsMMLU: 88.6%
Input: $0.20/1MOutput: $1.20/1M
Input: $25.00Output: $60.00Save $15.00
$85.00/mo
$0.850 per 1k reqs
Get Keys
Gemini 3.5 Flash LiteGoogleSpeed Champion

Sub-second real-time streaming model running at 185 tokens/second with 1M context window.

Context: 1MSpeed: ~185 tpsMMLU: 86.8%
Input: $0.30/1MOutput: $2.50/1M
Input: $37.50Output: $125.00Save $22.50
$162.50/mo
$1.63 per 1k reqs
Get Keys
DeepSeek R1 (Latest)DeepSeek

Updated R1 open reasoning model with expanded 163k context and competitive math performance.

Context: 163.8kSpeed: ~50 tpsMMLU: 91.8%
Input: $0.50/1MOutput: $2.15/1M
Input: $62.50Output: $107.50Save $37.50
$170.00/mo
$1.70 per 1k reqs
Get Keys
DeepSeek V4 ProDeepSeekBest Value

Frontier 1M-context flagship model from DeepSeek delivering frontier parity at $0.66/1M.

Context: 1MSpeed: ~75 tpsMMLU: 94.1%
Input: $0.66/1MOutput: $1.98/1M
Input: $72.60Output: $99.00Save $59.40
$171.60/mo
$1.72 per 1k reqs
Get Keys
Llama 3.3 70B InstructMeta

Widely deployed open model with 131k context and flat symmetric token pricing.

Context: 131.1kSpeed: ~180 tpsMMLU: 86.4%
Input: $0.71/1MOutput: $0.71/1M
Input: $142.00Output: $35.50
$177.50/mo
$1.78 per 1k reqs
Get Keys
Devstral 2 (Mistral)MistralTop Coding Engine

Specialized 2nd generation developer code model with 262k context and superior syntax precision.

Context: 262.1kSpeed: ~110 tpsMMLU: 86%
Input: $0.44/1MOutput: $2.20/1M
Input: $88.00Output: $110.00
$198.00/mo
$1.98 per 1k reqs
Get Keys
ByteDance Seed-2.0-CodeByteDance

ByteDance’s high-throughput code synthesis engine with 262k context window.

Context: 262.1kSpeed: ~120 tpsMMLU: 87.2%
Input: $0.50/1MOutput: $3.00/1M
Input: $100.00Output: $150.00
$250.00/mo
$2.50 per 1k reqs
Get Keys
Gemini 3.7 FlashGoogleSpeed Champion

Google’s state-of-the-art 3.7 generation Flash model with 1,048,576 token context and native multimodal reasoning.

Context: 1MSpeed: ~150 tpsMMLU: 92.1%
Input: $0.75/1MOutput: $3.75/1M
Input: $93.75Output: $187.50Save $56.25
$281.25/mo
$2.81 per 1k reqs
Get Keys
xAI Grok Build 0.1xAI

Specialized code-focused Grok model offering aggressive $1.00 input / $2.00 output pricing.

Context: 256kSpeed: ~130 tpsMMLU: 87%
Input: $1.00/1MOutput: $2.00/1M
Input: $200.00Output: $100.00
$300.00/mo
$3.00 per 1k reqs
Get Keys
Kimi K2.7 CodeMoonshotTop Coding Engine

Moonshot’s specialized coding intelligence optimized for deep reasoning and long-context codebase understanding.

Context: 262.1kSpeed: ~90 tpsMMLU: 88.5%
Input: $0.66/1MOutput: $3.40/1M
Input: $132.00Output: $170.00
$302.00/mo
$3.02 per 1k reqs
Get Keys
GPT-5.4 MiniOpenAI

Compact 5.4 model offering 400k context and fast execution.

Context: 400kSpeed: ~140 tpsMMLU: 86.4%
Input: $0.75/1MOutput: $4.50/1M
Input: $93.75Output: $225.00Save $56.25
$318.75/mo
$3.19 per 1k reqs
Get Keys
Mistral Medium 3.5Mistral

Mistral’s flagship model with 262k context, European hosting, and state-of-the-art multilingual fluency.

Context: 262.1kSpeed: ~90 tpsMMLU: 89.6%
Input: $1.50/1MOutput: $7.50/1M
Input: $300.00Output: $375.00
$675.00/mo
$6.75 per 1k reqs
Get Keys
xAI Grok 4.6xAIBest Value

xAI’s flagship model with 500,000 context window, competitive $2.00/$6.00 pricing, and real-time knowledge.

Context: 500kSpeed: ~110 tpsMMLU: 92.8%
Input: $2.00/1MOutput: $6.00/1M
Input: $400.00Output: $300.00
$700.00/mo
$7.00 per 1k reqs
Get Keys
Qwen 3.8 MaxQwenBest Value

Alibaba’s frontier 3.8 Max model with 1M context and superior multilingual instruction compliance.

Context: 1MSpeed: ~80 tpsMMLU: 93%
Input: $2.00/1MOutput: $6.00/1M
Input: $400.00Output: $300.00
$700.00/mo
$7.00 per 1k reqs
Get Keys
Claude Sonnet 5AnthropicMost Popular

Anthropic’s flagship 5th-generation model with 1,000,000 token context, 90% prompt caching discount, and top SWE-bench benchmarks.

Context: 1MSpeed: ~90 tpsMMLU: 93.8%
Input: $2.00/1MOutput: $10.00/1M
Input: $220.00Output: $500.00Save $180.00
$720.00/mo
$7.20 per 1k reqs
Get Keys
GPT-5.6 SolOpenAIMost Popular

OpenAI’s premier 5.6-generation flagship model with 1.05M context and unified multimodal capabilities.

Context: 1.1MSpeed: ~100 tpsMMLU: 93.4%
Input: $2.00/1MOutput: $10.00/1M
Input: $250.00Output: $500.00Save $150.00
$750.00/mo
$7.50 per 1k reqs
Get Keys
GPT-5.6 TerraOpenAIBest Reasoning

Advanced reasoning model from OpenAI combining RL thinking tokens with 1.05M token context.

Context: 1.1MSpeed: ~70 tpsMMLU: 95%
Input: $2.00/1MOutput: $12.00/1M
Input: $250.00Output: $600.00Save $150.00
$850.00/mo
$8.50 per 1k reqs
Get Keys
GPT-5.3 CodexOpenAITop Coding Engine

OpenAI’s specialized 5.3 generation developer model fine-tuned on billions of code repositories.

Context: 400kSpeed: ~95 tpsMMLU: 92%
Input: $1.75/1MOutput: $14.00/1M
Input: $218.75Output: $700.00Save $131.25
$918.75/mo
$9.19 per 1k reqs
Get Keys
Cohere Command ACohere

Cohere’s enterprise flagship with verifiable citations and built-in grounding for production RAG.

Context: 256kSpeed: ~70 tpsMMLU: 88%
Input: $2.50/1MOutput: $10.00/1M
Input: $500.00Output: $500.00
$1,000.00/mo
$10.00 per 1k reqs
Get Keys
Claude Sonnet 4.6Anthropic

Established 4.6 generation Sonnet with 1M context and 90% prompt caching discounts.

Context: 1MSpeed: ~85 tpsMMLU: 91.5%
Input: $3.00/1MOutput: $15.00/1M
Input: $330.00Output: $750.00Save $270.00
$1,080.00/mo
$10.80 per 1k reqs
Get Keys
Claude Opus 5AnthropicFrontier Flagship

Anthropic’s highest capability model with 1M context, exceptional nuance, and superior domain comprehension.

Context: 1MSpeed: ~60 tpsMMLU: 95.2%
Input: $5.00/1MOutput: $25.00/1M
Input: $550.00Output: $1,250.00Save $450.00
$1,800.00/mo
$18.00 per 1k reqs
Get Keys
Claude Fable 5AnthropicBest Reasoning

Specialized deep reasoning model from Anthropic with 1M context and extensive chain-of-thought verification.

Context: 1MSpeed: ~50 tpsMMLU: 96.5%
Input: $10.00/1MOutput: $50.00/1M
Input: $1,100.00Output: $2,500.00Save $900.00
$3,600.00/mo
$36.00 per 1k reqs
Get Keys
Head-to-Head Showdowns

Popular Model Pricing Comparisons

View all 15+ models
Direct Comparison

Claude Sonnet 5 vs GPT-5.6 Sol

Comprehensive side-by-side cost simulator, 1M context token economics, prompt caching discounts, and SWE-bench comparison.

Detailed Cost & Latency Table Compare →
Direct Comparison

GPT-5.3 Codex vs Claude Sonnet 5

Evaluate developer token pricing, repository-level refactoring, and SWE-bench accuracy.

Detailed Cost & Latency Table Compare →
Direct Comparison

Kimi K2.7 Code vs Devstral 2

Moonshot AI’s specialized developer intelligence vs Mistral’s Devstral 2.

Detailed Cost & Latency Table Compare →
Direct Comparison

xAI Grok 4.6 vs GPT-5.6 Sol

Real-time search-grounded intelligence against OpenAI’s unified multimodal flagship.

Detailed Cost & Latency Table Compare →
Direct Comparison

Qwen 3.8 Max vs Claude Sonnet 5

Alibaba’s top multilingual powerhouse against Anthropic’s coding flagship.

Detailed Cost & Latency Table Compare →
Direct Comparison

DeepSeek V4 Pro vs GPT-5.6 Terra

Evaluate token savings, mathematical benchmark parity, and 1M context chain-of-thought pricing.

Detailed Cost & Latency Table Compare →
Direct Comparison

GPT-5.6 Luna vs Gemini 3.7 Flash

Battle of the workhorse models: Pricing, 1M context limits, and token calculation.

Detailed Cost & Latency Table Compare →
Direct Comparison

DeepSeek V4 Flash vs GPT-5.6 Luna

Can DeepSeek V4 Flash ($0.03/1M) beat OpenAI’s fastest intelligence ($0.20/1M)?

Detailed Cost & Latency Table Compare →
Direct Comparison

Llama 4 Maverick vs GPT-5.6 Luna

Meta’s frontier open model against OpenAI’s lightweight powerhouse.

Detailed Cost & Latency Table Compare →
Direct Comparison

Claude Opus 5 vs Claude Fable 5

Anthropic’s 1M context synthesis master against its specialized reasoning engine.

Detailed Cost & Latency Table Compare →
Direct Comparison

Gemini 3.7 Flash vs Claude Sonnet 5

Google’s 150 TPS multimodal leader against Anthropic’s coding champion.

Detailed Cost & Latency Table Compare →
Direct Comparison

Devstral 2 vs Claude Sonnet 5

Mistral’s dedicated developer model against Anthropic’s flagship coder.

Detailed Cost & Latency Table Compare →
Production Architectures

Workload-Specific Cost Estimators

Estimate token usage formulas, system prompt overheads, and vector database query costs for your specific AI product archetype.

Architectural Optimization

How to Cut Your AI API Bills by 60% to 90%

Modern LLM billing has evolved far beyond raw token counts. High-performing engineering teams leverage three primary levers to maximize ROI:

  • Prompt Prefix Caching: Structure system instructions and tool schemas at the very start of prompts. Anthropic and DeepSeek reduce cached token costs by 90% ($0.20/1M vs $2.00/1M on Claude Sonnet 5).
  • Model Routing & Tiering: Route simple classification and data extraction queries to ultra-fast sub-dollar models (DeepSeek V4 Flash, GPT-5.6 Luna, Gemini 3.5 Flash Lite), reserving frontier reasoning models (Claude Fable 5, GPT-5.6 Terra) strictly for complex logic.
  • Output Token Restraint: Because output tokens are priced 3x-5x higher than input tokens, enforce strict JSON schemas and concise max_tokens constraints.

Prompt Caching Savings Matrix (Per 100M Tokens)

Claude Sonnet 5
Uncached: $200.00 / 100M
$20.00 / 100M
90% Discount ($180 saved)
DeepSeek V4 Flash
Uncached: $3.00 / 100M
$0.30 / 100M
90% Discount ($2.70 saved)
GPT-5.6 Sol
Uncached: $200.00 / 100M
$50.00 / 100M
75% Discount ($150 saved)

Frequently Asked Questions

Monthly token cost is calculated using the formula: ((Monthly Requests × Input Tokens per Request) / 1,000,000 × Price per 1M Input Tokens) + ((Monthly Requests × Output Tokens per Request) / 1,000,000 × Price per 1M Output Tokens). When prompt caching is active, cached input tokens receive up to a 90% discount depending on the provider.