TokenCost.io
Legal / Finance / Healthcare Production Architecture

Document & PDF Summarizer Cost Calculator

Ingesting 30 to 100-page PDF reports, balance sheets, or clinical trials to generate structured summaries. High input tokens dominate document analysis. Gemini 3.7 Flash offers a massive 1M context window with native multimodal comprehension, while DeepSeek V4 Flash offers 1.31M context at $0.03/1M.

Quick Workload Presets

Workload Parameters

326M tokens/mo
req
1k100k1M5M10M+
tok
1002,000 (RAG)8,000 (Agent)32,000 (Doc QA)64k+
tok
50 (Short JSON)500 (Chat response)2,000 (Code)4,000 (Essay)8k+
25% Cache Hit
10% Light reuse50% Standard RAG75% Agent Loop90% High caching

Monthly Economics Summary

Total Tokens326M320M In / 6M Out
Cheapest ModelDeepSeek V4 Flash$7.98 /mo
Top ReasoningClaude Fable 5$2,780.00 /mo
Prompt Caching Active (25% Cache Hit)Saves up to 90% on input tokens for Claude, DeepSeek & Gemini
DeepSeek V4 FlashDeepSeekLowest Cost

Disruptive sub-cent pricing at $0.03/1M input and 1.31M token context window.

Context: 1.3MSpeed: ~180 tpsMMLU: 89.4%
Input: $0.03/1MOutput: $0.09/1M
Input: $7.44Output: $0.540Save $2.16
$7.98/mo
$0.798 per 1k reqs
Get Keys
Llama 4 ScoutMetaSpeed Champion

Ultra-fast lightweight Llama 4 model running with 1.31M context window at 290 tokens/sec.

Context: 1.3MSpeed: ~290 tpsMMLU: 87.5%
Input: $0.11/1MOutput: $0.34/1M
Input: $35.20Output: $2.04
$37.24/mo
$3.72 per 1k reqs
Get Keys
Qwen 3.8 FlashQwenSpeed Champion

Ultra-fast 1M context model running at 170 TPS at only $0.15/1M input.

Context: 1MSpeed: ~170 tpsMMLU: 88%
Input: $0.15/1MOutput: $0.47/1M
Input: $48.00Output: $2.82
$50.82/mo
$5.08 per 1k reqs
Get Keys
Mistral Small 4MistralLowest Cost

Fast, lightweight model from Mistral with 262k context at only $0.15/1M tokens.

Context: 262.1kSpeed: ~150 tpsMMLU: 84.8%
Input: $0.15/1MOutput: $0.60/1M
Input: $48.00Output: $3.60
$51.60/mo
$5.16 per 1k reqs
Get Keys
GPT-5.6 LunaOpenAIBest Value

OpenAI’s fast workhorse model with 1.05M context at an ultra-low $0.20/1M input price.

Context: 1.1MSpeed: ~165 tpsMMLU: 88.6%
Input: $0.20/1MOutput: $1.20/1M
Input: $52.00Output: $7.20Save $12.00
$59.20/mo
$5.92 per 1k reqs
Get Keys
DeepSeek V3.2DeepSeek

Upgraded V3 architecture delivering reliable function calling and 163k context.

Context: 163.8kSpeed: ~80 tpsMMLU: 90.2%
Input: $0.26/1MOutput: $0.38/1M
Input: $64.48Output: $2.28Save $18.72
$66.76/mo
$6.68 per 1k reqs
Get Keys
Llama 4 MaverickMetaOpen Weights King

Meta’s frontier Llama 4 flagship model with 1,048,576 token context and complete weights openness.

Context: 1MSpeed: ~220 tpsMMLU: 91.2%
Input: $0.20/1MOutput: $0.80/1M
Input: $64.00Output: $4.80
$68.80/mo
$6.88 per 1k reqs
Get Keys
Gemini 3.5 Flash LiteGoogleSpeed Champion

Sub-second real-time streaming model running at 185 tokens/second with 1M context window.

Context: 1MSpeed: ~185 tpsMMLU: 86.8%
Input: $0.30/1MOutput: $2.50/1M
Input: $78.00Output: $15.00Save $18.00
$93.00/mo
$9.30 per 1k reqs
Get Keys
DeepSeek R1 (Latest)DeepSeek

Updated R1 open reasoning model with expanded 163k context and competitive math performance.

Context: 163.8kSpeed: ~50 tpsMMLU: 91.8%
Input: $0.50/1MOutput: $2.15/1M
Input: $130.00Output: $12.90Save $30.00
$142.90/mo
$14.29 per 1k reqs
Get Keys
Devstral 2 (Mistral)MistralTop Coding Engine

Specialized 2nd generation developer code model with 262k context and superior syntax precision.

Context: 262.1kSpeed: ~110 tpsMMLU: 86%
Input: $0.44/1MOutput: $2.20/1M
Input: $140.80Output: $13.20
$154.00/mo
$15.40 per 1k reqs
Get Keys
DeepSeek V4 ProDeepSeekBest Value

Frontier 1M-context flagship model from DeepSeek delivering frontier parity at $0.66/1M.

Context: 1MSpeed: ~75 tpsMMLU: 94.1%
Input: $0.66/1MOutput: $1.98/1M
Input: $163.68Output: $11.88Save $47.52
$175.56/mo
$17.56 per 1k reqs
Get Keys
ByteDance Seed-2.0-CodeByteDance

ByteDance’s high-throughput code synthesis engine with 262k context window.

Context: 262.1kSpeed: ~120 tpsMMLU: 87.2%
Input: $0.50/1MOutput: $3.00/1M
Input: $160.00Output: $18.00
$178.00/mo
$17.80 per 1k reqs
Get Keys
Gemini 3.7 FlashGoogleSpeed Champion

Google’s state-of-the-art 3.7 generation Flash model with 1,048,576 token context and native multimodal reasoning.

Context: 1MSpeed: ~150 tpsMMLU: 92.1%
Input: $0.75/1MOutput: $3.75/1M
Input: $195.00Output: $22.50Save $45.00
$217.50/mo
$21.75 per 1k reqs
Get Keys
GPT-5.4 MiniOpenAI

Compact 5.4 model offering 400k context and fast execution.

Context: 400kSpeed: ~140 tpsMMLU: 86.4%
Input: $0.75/1MOutput: $4.50/1M
Input: $195.00Output: $27.00Save $45.00
$222.00/mo
$22.20 per 1k reqs
Get Keys
Llama 3.3 70B InstructMeta

Widely deployed open model with 131k context and flat symmetric token pricing.

Context: 131.1kSpeed: ~180 tpsMMLU: 86.4%
Input: $0.71/1MOutput: $0.71/1M
Input: $227.20Output: $4.26Save $0.00000
$231.46/mo
$23.15 per 1k reqs
Get Keys
Kimi K2.7 CodeMoonshotTop Coding Engine

Moonshot’s specialized coding intelligence optimized for deep reasoning and long-context codebase understanding.

Context: 262.1kSpeed: ~90 tpsMMLU: 88.5%
Input: $0.66/1MOutput: $3.40/1M
Input: $211.20Output: $20.40
$231.60/mo
$23.16 per 1k reqs
Get Keys
xAI Grok Build 0.1xAI

Specialized code-focused Grok model offering aggressive $1.00 input / $2.00 output pricing.

Context: 256kSpeed: ~130 tpsMMLU: 87%
Input: $1.00/1MOutput: $2.00/1M
Input: $320.00Output: $12.00
$332.00/mo
$33.20 per 1k reqs
Get Keys
Mistral Medium 3.5Mistral

Mistral’s flagship model with 262k context, European hosting, and state-of-the-art multilingual fluency.

Context: 262.1kSpeed: ~90 tpsMMLU: 89.6%
Input: $1.50/1MOutput: $7.50/1M
Input: $480.00Output: $45.00
$525.00/mo
$52.50 per 1k reqs
Get Keys
GPT-5.3 CodexOpenAITop Coding Engine

OpenAI’s specialized 5.3 generation developer model fine-tuned on billions of code repositories.

Context: 400kSpeed: ~95 tpsMMLU: 92%
Input: $1.75/1MOutput: $14.00/1M
Input: $455.00Output: $84.00Save $105.00
$539.00/mo
$53.90 per 1k reqs
Get Keys
Claude Sonnet 5AnthropicMost Popular

Anthropic’s flagship 5th-generation model with 1,000,000 token context, 90% prompt caching discount, and top SWE-bench benchmarks.

Context: 1MSpeed: ~90 tpsMMLU: 93.8%
Input: $2.00/1MOutput: $10.00/1M
Input: $496.00Output: $60.00Save $144.00
$556.00/mo
$55.60 per 1k reqs
Get Keys
GPT-5.6 SolOpenAIMost Popular

OpenAI’s premier 5.6-generation flagship model with 1.05M context and unified multimodal capabilities.

Context: 1.1MSpeed: ~100 tpsMMLU: 93.4%
Input: $2.00/1MOutput: $10.00/1M
Input: $520.00Output: $60.00Save $120.00
$580.00/mo
$58.00 per 1k reqs
Get Keys
GPT-5.6 TerraOpenAIBest Reasoning

Advanced reasoning model from OpenAI combining RL thinking tokens with 1.05M token context.

Context: 1.1MSpeed: ~70 tpsMMLU: 95%
Input: $2.00/1MOutput: $12.00/1M
Input: $520.00Output: $72.00Save $120.00
$592.00/mo
$59.20 per 1k reqs
Get Keys
xAI Grok 4.6xAIBest Value

xAI’s flagship model with 500,000 context window, competitive $2.00/$6.00 pricing, and real-time knowledge.

Context: 500kSpeed: ~110 tpsMMLU: 92.8%
Input: $2.00/1MOutput: $6.00/1M
Input: $640.00Output: $36.00
$676.00/mo
$67.60 per 1k reqs
Get Keys
Qwen 3.8 MaxQwenBest Value

Alibaba’s frontier 3.8 Max model with 1M context and superior multilingual instruction compliance.

Context: 1MSpeed: ~80 tpsMMLU: 93%
Input: $2.00/1MOutput: $6.00/1M
Input: $640.00Output: $36.00
$676.00/mo
$67.60 per 1k reqs
Get Keys
Claude Sonnet 4.6Anthropic

Established 4.6 generation Sonnet with 1M context and 90% prompt caching discounts.

Context: 1MSpeed: ~85 tpsMMLU: 91.5%
Input: $3.00/1MOutput: $15.00/1M
Input: $744.00Output: $90.00Save $216.00
$834.00/mo
$83.40 per 1k reqs
Get Keys
Cohere Command ACohere

Cohere’s enterprise flagship with verifiable citations and built-in grounding for production RAG.

Context: 256kSpeed: ~70 tpsMMLU: 88%
Input: $2.50/1MOutput: $10.00/1M
Input: $800.00Output: $60.00
$860.00/mo
$86.00 per 1k reqs
Get Keys
Claude Opus 5AnthropicFrontier Flagship

Anthropic’s highest capability model with 1M context, exceptional nuance, and superior domain comprehension.

Context: 1MSpeed: ~60 tpsMMLU: 95.2%
Input: $5.00/1MOutput: $25.00/1M
Input: $1,240.00Output: $150.00Save $360.00
$1,390.00/mo
$139.00 per 1k reqs
Get Keys
Claude Fable 5AnthropicBest Reasoning

Specialized deep reasoning model from Anthropic with 1M context and extensive chain-of-thought verification.

Context: 1MSpeed: ~50 tpsMMLU: 96.5%
Input: $10.00/1MOutput: $50.00/1M
Input: $2,480.00Output: $300.00Save $720.00
$2,780.00/mo
$278.00 per 1k reqs
Get Keys

Architectural Cost Breakdown

In a typical Document & PDF Summarizer pipeline, tokens are distributed across three distinct layers:

1. System Prompt & Tool Definitions: ~400 - 800 Tokens (Cached)
2. Retrieved Knowledge / Buffer Context: ~31400 Tokens
3. Assistant Generation Response: ~600 Tokens

Prompt caching discounts apply immediately to Layer 1 and repeated multi-turn turns, reducing monthly overhead substantially.

Top Recommended Model Optimal ROI

Gemini 3.7 Flash

Google

Real-time multimodal audio/video understanding, 1M context document QA, high-speed agents

Input: $0.75/1M Output: $3.75/1M
Deploy Gemini 3.7 Flash
Runner-Up Alternative: DeepSeek V4 Flash

Disruptive sub-cent pricing at $0.03/1M input and 1.31M token context window.

Frequently Asked Questions

For this workload, we simulate an average of 10k monthly requests with 32000 input tokens and 600 output tokens per request. With 25% prompt caching active, repetitive system prompts and context documents receive significant input discounts.