Compare current LLM API pricing for GPT, Claude, Gemini, Grok, DeepSeek, Qwen, Mistral, and other major model providers.
The tables below list context windows and standard input and output prices, usually quoted per 1 million tokens.
Start with the quick comparison for current flagship models, then use the provider tables for individual model families, legacy models, image, audio, video, and other API pricing.
Quick LLM API Pricing Comparison
| Model | Context Window | Input/1M Tokens | Output/1M Tokens |
|---|---|---|---|
| GPT-5.6 Sol | ≤ 272K >272k | $5.00 $10.00 | $30.00 $45.00 |
| GPT-5.6 Terra | ≤ 272K >272k | $2.00 $4.00 | $12.00 $18.00 |
| GPT-5.6 Luna | ≤ 272K >272k | $0.20 $0.40 | $1.20 $1.80 |
| Claude Fable 5 Claude Mythos 5 | 1M | $10.00 | $50.00 |
| Claude Opus 5 | 1M | $5.00 | $25.00 |
| Claude Sonnet 5 | 1M | $2.00 | $10.00 |
| Gemini 3.1 Pro | 200K | $2.00 | $12.00 |
| Gemini 3.7 Flash | 1M | $0.75 | $3.75 |
| Nano Banana 2 | – | $0.50 | $3.00 (text and thinking) $60.00 (images) $0.067 per 1K image $0.101 per 2K image $0.151 per 4K image |
| Grok 4.6 | 500K | $2.00 | $6.00 |
| DeepSeek-V4-Pro-0813 | 1M | $0.66 (OFF-PEAK) $1.32 (PEAK) | $1.98 (OFF-PEAK) $3.96 (PEAK) |
| Qwen3.8-Max | 1M | $2.00 | $6.00 |
| GLM-5.3 | 1M | $1.40 | $4.40 |
| MiniMax M3 | ≤ 512K | $0.60 | $2.40 |
| MiniMax M3 | > 512K | $1.20 | $4.80 |
| Kimi-k3 | 1M | $3.00 | $15.00 |
| Muse Spark 1.2 | 1M | $0.10 (Contributor) $1.25 | $0.20 (Contributor) $4.25 |
How to Compare LLM API Pricing
LLM API pricing usually starts with two numbers: the cost per 1 million input tokens and the cost per 1 million output tokens. Input tokens include the prompt, system instructions, conversation history, and other context sent to the model. Output tokens are the text or reasoning tokens generated in the response.
A lower input price does not automatically mean a lower total cost. Models can have very different output rates, context pricing rules, caching discounts, and batch or off-peak rates. Some APIs also charge separately for tools, search, audio, images, video, or other modalities.
For a basic text request, the cost can be estimated as:
API cost = (input tokens ÷ 1,000,000 × input price) + (output tokens ÷ 1,000,000 × output price)
For example, a workload with long prompts but short responses depends more heavily on the input rate. A coding or reasoning workload that generates large responses can make the output rate much more important.
API Pricing by Provider
OpenAI GPT Models API Pricing
| Model | Context Window | Input/1M Tokens | Output/1M Tokens |
|---|---|---|---|
| GPT-5.6 Sol | ≤ 272K | $5.00 | $30.00 |
| GPT-5.6 Sol | >272K | $10.00 | $45.00 |
| GPT-5.6 Terra | ≤ 272K | $2.00 | $12.00 |
| GPT-5.6 Terra | >272K | $4.00 | $18.00 |
| GPT-5.6 Luna | ≤ 272K | $0.20 | $1.20 |
| GPT-5.6 Luna | >272K | $0.40 | $1.80 |
| gpt-5.5 | <272K | $5.00 | $30.00 |
| gpt-5.5 | >272K (1M Max) | $10.00 | $45.00 |
| gpt-5.5-pro | <272K | $30.00 | $180.00 |
| gpt-5.5-pro | >272K (1M Max) | $60.00 | $270.00 |
| gpt-5.4 | <272K | $2.50 | $15.00 |
| gpt-5.4 | >272K | $5.00 | $22.50 |
| gpt-5.4-pro | <272K | $30.00 | $180.00 |
| gpt-5.4-pro | >272K | $60.00 | $270.00 |
| gpt-5.4-mini | 400K | $0.75 | $4.50 |
| gpt-5.4-nano | 400K | $0.20 | $1.25 |
| gpt-5.2 | 400K | $1.75 | $14.00 |
| gpt-5.2-pro | 400K | $21 | $168 |
| gpt-5.1 | 400K | $1.25 | $10.00 |
| gpt-5 | 400K | $1.25 | $10.00 |
| gpt-5-mini | 400K | $0.25 | $2.00 |
| gpt-5-nano | 400K | $0.05 | $0.40 |
| gpt-5-pro | 400K | $15.00 | $120.00 |
| gpt-5.3-chat-latest | 400K | $1.75 | $14.00 |
| gpt-5.3-codex | – | $1.75 | $14.00 |
| gpt-5.1-codex-max | – | $1.25 | $10.00 |
| gpt-5.1-codex-mini | – | $0.25 | $6.00 |
| codex-mini-latest | – | $1.50 | $6.00 |
| gpt-5-search-api | – | $1.25 | $10.00 |
| gpt-4.1 | 1M | $2.00 | $8.00 |
| gpt-4.1-mini | 1M | $0.40 | $1.60 |
| gpt-4.1-nano | 1M | $0.10 | $0.40 |
| gpt-4o | 128K | $2.50 | $10.00 |
| gpt-4o-mini | 128K | $0.15 | $0.60 |
OpenAI Image Generation API Pricing
| Model | Input/1M Tokens | Output/1M Tokens |
|---|---|---|
| GPT-image-2 | $5.00 (text) $8.00 (image) | $30.00 |
| GPT-image-1.5 | $5.00 | $10.00 |
| GPT-image-1 | $5.00 | – |
| GPT-image-1-mini | $2.00 | – |
OpenAI Audio API Pricing
| Model | Input/1M Tokens | Output/1M Tokens |
|---|---|---|
| gpt-realtime-2.1 | $4.00 (text) $5.00 (image) $32.00 (audio) | $24.00 (text) $64.00 (audio) |
| gpt-realtime-2.1-mini | $0.60 (text) $0.80 (image) $10.00 (audio) | $2.40 (text) $20.00 (audio) |
| gpt-realtime-translate | – | $0.034 / minute |
| gpt-transcribe | – | $0.0045 / minute |
| gpt-live-transcribe | – | $0.017 / minute |
| gpt-realtime-whisper | – | $0.017 / minute |
OpenAI Reasoning Models API Pricing (Legacy)
| Model | Context Window | Input/1M Tokens | Output/1M Tokens |
|---|---|---|---|
| o4-mini | 200K | $1.10 | $4.40 |
| o4-mini-deep-research | 200K | $2.00 | $8.00 |
| o3-pro | 200K | $20.00 | $80.00 |
| o3 | 200K | $2.00 | $8.00 |
| o3-deep-research | 200K | $10.00 | $40.00 |
| o3-mini | 200K | $1.10 | $4.40 |
| o1-pro | 200K | $150.00 | $600.00 |
| o1 | 200K | $15.00 | $60.00 |
| o1-mini | 128K | $1.10 | $4.40 |
OpenAI Video Generation API Pricing
| Model | Price per second |
|---|---|
| Sora 2 | $0.10 |
| Sora 2 Pro | $0.30 |
| Sora 2 Pro (Portrait: 1024 x 1792. Landscape: 1792 x 1024) | $0.50 |
Claude API Pricing
| Model | Context Window | Input/1M Tokens | Output/1M Tokens |
|---|---|---|---|
| Claude Fable 5 Claude Mythos 5 (limited availability) | 1M | $10.00 | $50.00 |
| Claude Opus 5 | 1M | $5.00 | $25.00 |
| Claude Opus 5 Fast Mode | 1M | $10.00 | $50.00 |
| Claude Opus 4.8 | 1M | $5.00 | $25.00 |
| Claude Opus 4.8 Fast Mode | 1M | $10.00 | $50.00 |
| Claude Sonnet 5 | 1M | $2.00 | $10.00 |
| Claude Haiku 4.5 | 200K | $1.00 | $5.00 |
Gemini API Pricing
| Model | Context Window | Input/1M Tokens | Output/1M Tokens |
|---|---|---|---|
| Gemini 3.7 Flash | 1M | $0.75 | $3.75 |
| Gemini 3.6 Flash | 1M | $0.75 | $3.75 |
| Gemini 3.5 Flash | 1M | $1.50 | $9.00 |
| Gemini 3.5 Flash-Lite | 1M | $0.30 | $2.50 |
| Gemini 3.5 Live Translate | – | $3.50 or $0.0053/min (audio) | $21.00 or $0.0315/min (audio) |
| Gemini 3.1 Pro | >200K | $4.00 | $18.00 |
| Gemini 3.1 Pro | 200K | $2.00 | $12.00 |
| Gemini 3.1 Flash-Lite | 1M | $0.25 (text/image/video) $0.50 (audio) | $1.50 |
| Gemini 3.1 Flash Live | – | $0.75 (text) $3.00 or $0.005/min (audio) $1.00 or $0.002/min (image/video) | $4.50 (text) $12.00 or $0.018/min (audio) |
| Gemini 3 Pro | >200K | $4.00 | $18.00 |
| Gemini 3 Pro | 200K | $2.00 | $12.00 |
| Gemini 3 Flash | 200K | $0.50 (text / image / video) $1.00 (audio) | $3.00 |
| Gemini 2.5 Pro | >200K | $2.50 | $15.00 |
| Gemini 2.5 Pro | 200K | $1.25 | $10.00 |
| Gemini 2.5 Flash | 1M | $0.30 (text/image/video) $1.00 (audio) | $2.50 |
| Gemini 2.5 Flash-Lite | 1M | $0.10 (text/image/video) $0.50 (audio) | $0.40 |
| Gemini 2.0 Flash | 1M | $0.10 $0.70 (audio) | $0.40 |
| Gemini 2.0 Flash-Lite | 1M | $0.075 | $0.30 |
Nano Banana Pricing
| Model | Free Tier | Input/1M Tokens | Output/1M Tokens |
|---|---|---|---|
| Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) | Not available | $0.25 | $1.50 (text and thinking) $30.00 (images) $0.034 per 1K image |
| Nano Banana 2 (Gemini 3.1 Flash Image) | Not available | $0.50 | $3.00 (text and thinking) $60.00 (images) $0.067 per 1K image $0.101 per 2K image $0.151 per 4K image |
| Nano Banana 2 Pro (Gemini 3 Pro Image) | Not available | $2.00 (text) $0.0011 (image) | $12.00 (text and thinking) $120.00 (images) $0.134 per 1K/2K image $0.24 per 4K image |
Gemini Omni Pricing
| Model | Free Tier | Input/1M Tokens | Output/1M Tokens |
|---|---|---|---|
| Gemini Omni Flash Preview | Not available | $1.50 (text/image/video/audio) | $9.00 (text) $17.50 (video) |
Gemini 3.5 Live Translate
| Model | Free Tier | Input | Output |
|---|---|---|---|
| gemini-3.5-live-translate-preview | Free of charge | $3.50 or $0.0053/min | $21.00 or $0.0315/min |
Gemini 2.5 Flash Native Audio API Pricing
| Model | Free Tier | Input/1M Tokens | Output/1M Tokens |
|---|---|---|---|
| Gemini 2.5 Flash Native Audio | Not available | $0.50 (text) $3.00 (audio/video) | $2.00 (text) $12.00 (audio) |
Gemini 2.5 Flash Image Preview Pricing
| Model | Free Tier | Input/1M Tokens | Output |
|---|---|---|---|
| Gemini 2.5 Flash Image Preview | Not available | $0.30 (text/image) | $0.039 per image |
Gemini TTS Pricing
| Model | Free Tier | Input/1M Tokens | Output/1M Tokens |
|---|---|---|---|
| Gemini 3.1 Flash TTS Preview | Free of charge | $1.00 (text) | $20.00 (audio) |
| Gemini 2.5 Flash Preview TTS | Free of charge | $0.50 (text) | $10.00 (audio) |
| Gemini 2.5 Pro Preview TTS | Not available | $1.00 (text) | $20.00 (audio) |
Google Imagen 4 & 3 API Pricing
| Model | Paid Tier, per Image in USD |
|---|---|
| Imagen 4 Fast | $0.02 |
| Imagen 4 Standard | $0.04 |
| Imagen 4 Ultra | $0.06 |
| Imagen 3 | $0.03 |
Gemini Embedding API Pricing
| Model | Paid Tier, per 1M tokens in USD |
|---|---|
| gemini-embedding-2 | $0.20 (Text) $0.45 (Image $0.00012 per image) $6.50 (Audio $0.00016 per second) $12.00 (Video $0.00079 per frame) |
| gemini-embedding-001 | $0.15 |
Google Veo 3.1 API Pricing
| Model | Paid Tier, per second in USD |
|---|---|
| Veo 3.1 Standard | $0.40 (720p and 1080p) $0.60 (4k) |
| Veo 3.1 Fast | $0.15 (720p and 1080p) $0.35 (4k) |
| Veo 3.1 Light | $0.05 (720p) $0.08 (1080p) |
Grok API Pricing
| Model | Context Window | Input/1M Tokens | Output/1M Tokens |
|---|---|---|---|
| Grok 4.6 | 500K | $2.00 | $6.00 |
| Grok 4.5 | 500K | $2.00 | $6.00 |
| Grok Build 0.1 | 256K | $1.00 | $2.00 |
| Grok 4.3 | 1M | $1.25 | $2.50 |
| Grok 4.20 Multi-agent | 2M | $2.00 | $6.00 |
| Grok 4.20 Reasoning | 2M | $2.00 | $6.00 |
| Grok 4.20 | 2M | $2.00 | $6.00 |
| Imagine Image 2.0 | – | 0.01/image | 0.04/image (1K Low Quality) 0.06/image (1K) 0.06/image (2K Low Quality) 0.08/image (2K) |
| Grok Imagine Image Quality | – | 0.01/image | 0.05/image (1K) 0.07/image (2K) |
| Grok Imagine Image | – | 0.002/image | 0.02/image (1K) 0.02/image (2K) |
| Imagine Video 1.5 | – | 0.01/second | 0.08/sec (480p) 0.14/sec (720p) 0.25/sec (1080p) |
| Grok Imagine Video | – | 0.01/sec 0.002/image | 0.05/sec (480p) 0.07/sec (720p) |
| grok-voice-think-fast-2.0 | – | – | $0.08/min ($4.80/hr) audio $0.004/text input |
| grok-voice-think-fast-1.0 | – | – | $0.05/min ($3.00/hr) audio $0.004/text input |
| Text to Speech | – | – | $15.00 / 1M chars |
| Speech to Text | – | – | $0.10 / hr (REST) $0.20 / hr (Streaming) |
DeepSeek API Pricing
| Model | Context Window | Input/1M Tokens | Output/1M Tokens |
|---|---|---|---|
| DeepSeek-V4-Flash-0731 | 1M | $0.22 (OFF-PEAK) $0.44 (PEAK) | $0.66 (OFF-PEAK) $1.32(PEAK) |
| DeepSeek-V4-Pro-0813 | 1M | $0.66 (OFF-PEAK) $1.32 (PEAK) | $1.98 (OFF-PEAK) $3.96 (PEAK) |
Qwen API Pricing
| Model | Context Window | Input/1M Tokens | Output/1M Tokens |
|---|---|---|---|
| Qwen3.8-Max | 1M | $2.00 | $6.00 |
| Qwen3.7-Plus | 1M | $0.40 | $1.60 |
| Qwen3.7-Flash | 1M | $0.03 | $0.13 |
| Qwen3.5-Omni-Plus | – | $11.00 (Audio) $1.40 (Text/Image/Video) | $44.00 (Text & Audio) $8.30 (Text) |
| Qwen Image 3.0 Pro | – | $0.03/IMG | $0.075/IMG (1K) $0.04/sec (2K) |
| Wan 3.0 Video | – | – | $0.05/sec (480p) $0.10/sec (720p) $0.20/sec (1080p) |
| Wan2.7-Image-Pro | – | – | $0.075/IMG Generation |
| HappyHorse-1.1-T2V | – | – | $0.07/sec (480p) $0.14/sec (720p) $0.24/sec (1080p) |
Mistral (Premier Models) API Pricing
| Model | Context Window | Input/1M Tokens | Output/1M Tokens |
|---|---|---|---|
| Magistral Medium | 128K | $2.00 | $5.00 |
| Codestral | 32K | $0.30 | $0.90 |
| Document AI & OCR 4.1 | – | – | OCR: $4/1000 pages Document: $5/1000 pages |
| Voxtral Mini Transcribe 2 | – | – | Audio Input/min $0.003 |
| Voxtral Realtime | – | – | Audio Input/min $0.006 |
| Mistral Embed | 32k | $0.10 | – |
| Mistral Moderation 24.11 | 32k | $0.10 | |
| Magistral Small | 128K | $0.50 | $1.50 |
| Codestral Embed | 128K | $0.15 | – |
Mistral (Open Models) API Pricing
| Model | Context Window | Input/1M Tokens | Output/1M Tokens |
|---|---|---|---|
| Mistral Medium 3.5 | 256K | $1.50 | $7.50 |
| Mistral Large 3 | 256K | $0.50 | $1.50 |
| Ministral 3 – 3B | 256K | $0.10 | $0.10 |
| Ministral 3 – 8B | 256K | $0.15 | $0.15 |
| Ministral 3 – 14B | 256K | $0.20 | $0.20 |
| Pixtral Large | 128K | $2.00 | $6.00 |
| Pixtral 12B | 128K | $0.15 | $0.15 |
| Mistral Nemo | 128K | $0.15 | $0.15 |
| Mistral Small 4.0 | 128K | $0.15 | $0.60 |
| Mistral Small 3.2 | 128K | $0.10 | $0.30 |
| Devstral 2 | 128K | FREE | FREE |
| Devstral Small 2 | 128K | FREE | FREE |
| Voxtral Mini | – | $0.001 (audio) $0.04 (text) | $0.04 |
| Voxtral Small | – | $0.004 (audio) $0.1 (text) | $0.03 |
| Voxtral TTS | – | $0.016/1k Chars | |
| Mistral 7B | 32K | $0.25 | $0.25 |
| Mixtral 8x7B | 32K | $0.70 | $0.70 |
| Mixtral 8x22B | 64K | $2.00 | $6.00 |
Meta AI Muse API Pricing
| Model | Context Window | Input/1M Tokens | Output/1M Tokens |
|---|---|---|---|
| Muse Spark 1.2 (Contributor) | 1M | $0.10 | $0.20 |
| Muse Spark 1.2 | 1M | $1.25 | $4.25 |
| Muse Spark 1.1 | 1M | $1.25 | $4.25 |
Llama 4 & 3 API Pricing
| Model | Context Window | Input/1M Tokens | Output/1M Tokens |
|---|---|---|---|
| Llama 4 Scout | 10M | $0.11 | $0.34 |
| Llama 4 Maverick | 10M | $0.20 | $0.60 |
| Llama 3.3 70B Versatile | 128K | $0.59 | $0.79 |
| Llama 3.3 70B SpecDec | 8192 | $0.59 | $0.99 |
| Llama 3.3 70b Instruct | 128K | $0.23 | $0.40 |
| Llama 3.3 70b Instruct-Turbo | 128K | $0.13 | $0.40 |
| Llama 3.2 90b Vision-Instruct | 128K | $0.35 | $0.40 |
| Llama 3.2 11b Vision-Instruct | 128K | $0.055 | $0.055 |
| Llama 3.1 405B | 128K | $1.79 | $1.79 |
| Llama 3.1 70B | 128K | $0.35 | $0.40 |
| Llama 3.1 8B | 128K | $0.09 | $0.09 |
GLM API Pricing
| Model | Context Window | Input/1M Tokens | Output/1M Tokens |
|---|---|---|---|
| GLM-5.3 | 1M | $1.40 | $4.40 |
| GLM-5.2 | 1M | $1.40 | $4.40 |
| GLM-5.1 | 200K | $1.40 | $4.40 |
| GLM-5 | 200K | $1.00 | $3.20 |
| GLM-5-Turbo | 200K | $1.20 | $4.00 |
| GLM-5-Code | 200K | $1.20 | $5.00 |
| GLM-5V-Turbo | 200K | $1.20 | $4.00 |
| GLM-OCR | – | $0.03 | $0.03 |
| GLM-4.7 | 128K | $0.60 | $2.20 |
| GLM-4.7-FlashX | 128K | $0.07 | $0.40 |
| GLM-4.6 | 128K | $0.60 | $2.20 |
| GLM-4.6v | 128K | $0.30 | $0.90 |
| GLM-4.6V-FlashX | 128K | $0.004 | $0.40 |
| GLM-4.6V-Flash | 128K | FREE | FREE |
| GLM-4.5 | 128K | $0.60 | $2.20 |
| GLM-4.5v | 64K | $0.60 | $1.80 |
| GLM-4.5-X | 128K | $0.45 | $8.90 |
| GLM-4.5-Air | 128K | $0.20 | $1.10 |
| GLM-4.5-AirX | 128K | $1.10 | $4.50 |
| GLM-4.5-Flash | 128K | FREE | FREE |
Kimi K3 & K2 API Pricing
| Model | Context Window | Input/1M Tokens | Output/1M Tokens |
|---|---|---|---|
| kimi-k3 | 1M | $3.00 | $15.00 |
| kimi-k2.7-code-highspeed | 1M | $1.90 | $8.00 |
| kimi-k2.7-code | 1M | $0.95 | $4.00 |
| kimi-k2.6 | 262,144 | $0.95 | $4.00 |
| kimi-k2.5 | 262,144 | $0.60 | $3.00 |
| kimi-k2-thinking | 262,144 | $0.60 | $2.50 |
| kimi-k2-thinking-turbo | 262,144 | $1.15 | $8.00 |
| kimi-k2-0905-preview | 262,144 | $0.60 | $2.50 |
| kimi-k2-0711-preview | 131K | $0.60 | $2.50 |
| kimi-k2-turbo-preview | 262,144 | $1.15 | $8.00 |
Minimax M3 & M2.7 API Pricing
| Model | Context Window | Input/1M Tokens | Output/1M Tokens |
|---|---|---|---|
| Minimax M3 | ≤ 512K | $0.30 | $1.20 |
| Minimax M3 | > 512K | $0.60 | $2.40 |
| Minimax M2.7 | 204,800 | $0.30 | $1.20 |
| Minimax M2.7 Highspeed | 204,800 | $0.60 | $2.40 |
Minimax Video/Audio/Image Generation API Pricing
| Model | Unit Price |
|---|---|
| MiniMax H3 | $0.13/s, 2K |
| MiniMax H3 | $0.09/s, 768P |
| Music-3.0 | $0.15/up-to-5 minutes music |
| image-01 | $0.0035 per image |
Perplexity API Pricing
| Model | Context Window | Input/1M Tokens | Output/1M Tokens |
|---|---|---|---|
| Sonar | 128K | $1.00 | $1.00 |
| Sonar Pro | 200K | $3.00 | $15.00 |
| Sonar Reasoning Pro | 200K | $2.00 | $8.00 |
| Sonar Deep Research | 128K | $2.00 | $8.00 |
Cohere API Pricing
| Model | Context Window | Input/1M Tokens | Output/1M Tokens |
|---|---|---|---|
| Command A | 256K | $2.50 | $10.00 |
| Command R+ | 128K | $2.50 | $10.00 |
| Command R | 128K IN/4K OUT | $0.15 | $0.60 |
| Command R7B | 128K | $0.0375 | $0.15 |
See Also:
- What Are The Rate Limits For OpenAI API?
- Compare AI Costs: Free LLM API Price Calculator
- ChatGPT Token Limit: Free, Plus, Pro, and OpenAI API Limits
- AI Coding Plan Comparison: Prices, Limits, and Features
LLM API Pricing FAQ
How is LLM API pricing calculated?
Most text LLM APIs charge according to the number of input and output tokens processed. Multiply the input token count by the model’s input rate, multiply the output token count by its output rate, and combine the two amounts.
The final bill can also include cached tokens, long-context rates, tool calls, search or grounding requests, and multimodal inputs or outputs.
Are ChatGPT, Claude, and Gemini subscriptions the same as API pricing?
No. Consumer subscriptions and API usage use different billing systems.
A ChatGPT, Claude, or Gemini subscription pays for access to the company’s consumer application and the features included with that plan. API access is generally billed through a developer platform according to token usage or other API-specific units.
Why can long-context API requests cost more?
Long prompts consume more input tokens, so they cost more even when the per-token rate stays unchanged. Some models also apply a higher pricing tier after the request crosses a specified context threshold.
For example, current GPT-5.6 pricing applies higher rates when input exceeds 272K tokens, while some Gemini pricing tiers also distinguish requests according to prompt length.
Do cached tokens cost the same as normal input tokens?
No. Supported APIs can charge a lower rate when the model reuses eligible prompt content from a cache instead of processing the same input from scratch.
Changelog:
08/19/2026
- Added GLM 5.3.
08/16/2026
- Updated DeepSeek API pricing.
08/13/2026
- Added Mistral OCR 4.1
- Added Gemini 3.7 Flash
- Added GLM 5.3
- Added Minimax Music-3.0
08/12/2026
- Added DeepSeek-V4-Pro-0813
- Added Grok 4.6
08/11/2026
- Added Grok Imagine Image 2.0
08/06/2026
- Added Wan 3.0 Video
08/05/2026
- Added Muse Spark 1.2
08/02/2026
- Added Qwen3.8-Max
07/31/2026
- Added Minimax H3.
07/30/2026
- Updated GPT 5.6 pricing.
07/29/2026
- Added Grok Voice Think Fast 2.0
07/28/2026
- Cleanup
- Added GPT-Live-Transcribe and GPT-Transcribe
07/24/2026
- Added Claude Opus 5
07/21/2026
- Added Gemini 3.6 Flash and 3.5 Flash-Lite
07/16/2026
- Added Kimi K3.
07/09/2026
- Updated for GPT-5.6.
07/08/2026
- Added Grok 4.5.
07/06/2026
- Added gpt-realtime-2.1-mini
07/01/2026
- Added kimi-k2.7-code-highspeed
06/30/2026
- Added Claude Sonnet 5
- Added Nano Banana Lite
06/26/2026
- Added GPT 5.6
06/16/2026
- Added OCR 4
06/16/2026
- Added GLM-5.2
06/13/2026
- Updated Minimax M3
06/12/2026
- Added kimi-k2.7-code
06/09/2026
- Added Claude Fable/Mythos 5.
- Added Gemini 3.5 Live Translate.
06/07/2026
- Updated Deepseek API.
06/04/2026
- Added grok-imagine-video-1.5-preview
06/01/2026
- Added MiniMax M3
- Added Qwen3.7-Plus and HappyHorse-1.0-T2V
05/28/2026
- Added Claude Opus 4.8
05/21/2026
- Added Qwen3.7-Max
05/19/2026
- Added Gemini 3.5 Flash.
05/13/2026
- Added Claude Opus 4.7 Fast Mode
05/08/2026
- Added grok-imagine-image-quality
05/07/2026
- Added OpenAI audio models
05/01/2026
- Added Grok 4.3
04/29/2026
- Updated for Mistral Medium 3.5
04/27/2026
- Updated for DeepSeek v4
04/25/2026
- Updated
04/24/2026
- Added Deeseek V4 models
04/23/2026
- Added GPT-5.5
04/22/2026
- Added Gemini Embedding 2
04/22/2026
- Added GPT-image-2
04/21/2026
- Added Qwen3.6-Max-Preview
04/20/2026
- Added Kimi K2.6
04/18/2026
- Cleanup Grok APIs.
04/16/2026
- Added Claude Opus 4.7
04/15/2026
- Added Gemini 3.1 Flash TTS Preview
04/08/2026
- Added GLM-5.1
04/01/2026
- Added Perplexity Sonar models.
- GLM-5V-Turbo.
03/31/2026
- Added Veo 3.1 Light
03/26/2026
- Added Voxtral TTS and Gemini 3.1 Flash Live
03/23/2026
- Updated for Grok 4.2-0309
03/18/2026
- Added Minimax M2.7.
03/17/2026
- Added GPT-5.4 mini and nano
03/16/2026
- Added Mistral Small 4.0
03/16/2026
- Added GLM-5-Turbo
03/05/2026
- Added GPT-5.4
03/03/2026
- Added Gemini 3.1 Flash-Lite
- Added GPT-5.3 Instant
02/26/2026
- Added Nano Banana 2 (Gemini 3.1 Flash Image Preview)
02/19/2026
- Added Gemini Pro 3.1
02/18/2026
- Added Sonnet 4.6
02/11/2026
- Added GLM-5
02/06/2026
- Added Opus 4.6 & Voxtral Mini Transcribe 2
01/29/2026
- Added grok-imagine-video and Kimi 2.5.
12/23/2025
- Updated GLM-4.7
12/17/2025
- Added Gemini 3 Flash
12/16/2025
- Added GPT-image-1.5
12/11/2025
- Added GPT 5.2
12/09/2025
- Added Devstral 2
12/02/2025
- Added Mistral Large 3
12/01/2025
- Updated for DeepSeek-V3.2-Speciale
11/24/2025
- Updated for Claude Opus 4.5
11/20/2025
- Added Nano Banana Pro.
- Added Grok 4.1 Fast.
- Removed legacy models.
11/19/2025
- Added more models
11/18/2025
- Added Gemini 3 Pro
11/14/2025
- Added GPT-5.1
- Updated Qwen models
- Updated Kimi models
- Added Minimax M2
11/08/2025
- Added codex-mini
10/15/2025
- Update for Claude Haiku 4.5 & Veo 3.1.
10/07/2025
- Added Sora 2 and more OpenAI APIs.
09/29/2025
- Updated for Claude 4.5 and DeepSeek-V3.2
09/19/2025
- Updated price
08/21/2025
- Updated for DeepSeek-V3.1
08/20/2025
- Added GLM-4.5 and Kimi K2 models.
08/14/2025
- Added Imagen 4 Fast
08/07/2025
- Added GPT-5
08/05/2025
- Added Claude Opus 4.1
07/15/2025
- Added Voxtral model family
07/14/2025
- Added Gemini Embedding
07/11/2025
- Added Grok 4
06/26/2025
- Added o3-deep-research and o4-mini-deep-research
06/25/2025
- Added Imagen 4 Ultra and Imagen 4 Standard
06/20/2025
- Added Gemini 2.5 Flash-Lite
06/10/2025
- Added o3-pro
- Added Mistral Magistral models.
06/10/2025
- OpenAI dropped the price of o3 by 80%
05/22/2025
- Updated for Claude 4
05/21/2025
- Updated Mistral models
05/21/2025
- Added Gemini 2.5 Flash Native Audio
05/08/2025
- Added Mistral Medium 3
04/23/2025
- Added OpenAI Image Generation
04/18/2025
- Added Gemini 2.5 Flash
- Updated Cohere models
04/16/2025
- Added o4-mini
04/14/2025
- Added gpt-4.1 family
04/11/2025
- Updated DeepSeek
04/11/2025
- Added Google Imagen 3 and Veo 2.
04/10/2025
- Added Qwen models
- Added Grok 3
04/05/2025
- Added Gemini Pro 2.5
03/19/2025
- Added o1-pro
02/28/2025
- Added GPT-4.5
02/24/2025
- Added Claude 3.7
02/21/2025
- Added Grok
02/06/2025
- Added Gemini 2.0 Flash and Gemini 2.0 Flash-Lite
02/02/2025
- Added o3-mini
01/31/2025
- Added DeepSeek v3 and DeepSeek R1.
12/18/2024
- o1 in the API comes with support for function calling, developer messages, Structured Outputs, and vision capabilities.
12/07/2024
- Added Llama 3.3
11/05/2024
- Added Claude Haiku 3.5
11/03/2024
- Added Gemini 1.5 Flash-8B
10/04/2024
- Added Llama 3.2
- Added gpt-4o-realtime-preview
09/25/2024
- Updated Google Gemini
09/13/2024
- Added OpenAI’s latest model: o1.
08/07/2024
- Added gpt-4o-2024-08-06, the latest gpt-4o snapshot that supports Structured Outputs
07/25/2024
- Updated Mistral Large 2
07/24/2024
- Added Llama 3.1 405B
07/20/2024
- Added GPT-4o-mini
- Updated prices
07/13/2024
- Updated
- Added Cohere’s Command API










