AI LLM API Pricing 2026: GPT-6.1 Sol, Gemini 3.8, Claude Opus 5.5 & More

The latest API pricing for popular AI models like GPT-6.1 Sol, Gemini 3.8 Flash, Claude Fable/Mythos 5.1, Qwen 3.8 Max, Grok 4.7, Deepseek v4, GLM 5.3, Kimi K3, and more.

Compare current LLM API pricing for GPT, Claude, Gemini, Grok, DeepSeek, Kimi, GLM, Qwen, Minimax, Mistral, and other major model providers.

The tables below list context windows and standard input and output prices, usually quoted per 1 million tokens.

Start with the quick comparison for current flagship models, then use the provider tables for individual model families, legacy models, image, audio, video, and other API pricing.

Last Updated: Sep 29, 2026

Quick LLM API Pricing Comparison

ModelContext WindowInput/1M TokensOutput/1M Tokens
GPT-6 Astra
≤272K input
1.05M$10.00$50.00
GPT-6 Astra
>272K input
1.05M$20.00$75.00
GPT-6.1 Sol
≤272K input
1.05M$2.00$10.00
GPT-6.1 Sol
>272K input
1.05M$4.00$15.00
GPT-5.6 Terra
≤272K input
1.05M$2.00$12.00
GPT-5.6 Terra
>272K input
1.05M$4.00$18.00
GPT-5.6 Luna
≤272K input
1.05M$0.20$1.20
GPT-5.6 Luna
>272K input
1.05M$0.40$1.80
Claude Fable 5.1
Claude Mythos 5.1
1M$10.00$50.00
Claude Opus 5.51M$4.00$20.00
Claude Sonnet 5.51M$2.00$10.00
Gemini 3.1 Pro Preview
≤200K input
1M$2.00$12.00
Gemini 3.1 Pro Preview
>200K input
1M$4.00$18.00
Gemini 3.8 Flash1M$0.75$3.75
Grok 4.7
<200K input
500K$2.00$6.00
Grok 4.7
≥200K input
500K$4.00$12.00
DeepSeek-V4.1-Flash1M$0.15 (OFF-PEAK)
$0.30 (PEAK)
$0.66 (OFF-PEAK)
$1.32 (PEAK)
Qwen3.8-Max1M$2.00$6.00
GLM-5.31M$1.40$4.40
MiniMax M3
≤512K input
1M$0.30$1.20
MiniMax M3
>512K input
1M$0.60$2.40
Kimi-k31M$3.00$15.00
Muse Spark 1.31M$0.10 (Contributor)
$1.25
$0.20 (Contributor)
$4.25

GPT-6 Astra and GPT-5.6 use 1.05M-token context windows. 272K is the threshold for higher long-context rates. Gemini 3.1 Pro Preview uses a 1M-token context window with a 200K pricing threshold. Gemini 3.8 Flash pricing is introductory through December 31, 2026.

How to Compare LLM API Pricing

LLM API pricing usually starts with two numbers: the cost per 1 million input tokens and the cost per 1 million output tokens. Input tokens include the prompt, system instructions, conversation history, and other context sent to the model. Output tokens are the text or reasoning tokens generated in the response.

A lower input price does not automatically mean a lower total cost. Models can have very different output rates, context pricing rules, caching discounts, and batch or off-peak rates. Some APIs also charge separately for tools, search, audio, images, video, or other modalities.

For a basic text request, the cost can be estimated as:

API cost = (input tokens ÷ 1,000,000 × input price) + (output tokens ÷ 1,000,000 × output price)

For example, a workload with long prompts but short responses depends more heavily on the input rate. A coding or reasoning workload that generates large responses can make the output rate much more important.

API Pricing by Provider

GPT-6 Astra API Pricing

GPT-6 Astra costs $10 per million input tokens, $1 per million cached input tokens, $12.50 per million cache-write tokens, and $50 per million output tokens at Standard processing rates. The model supports a 1,050,000-token context window and up to 128,000 output tokens.

Astra uses higher pricing when a prompt contains more than 272,000 input tokens. Once that threshold is crossed, the higher rates apply to the full request. Input, cached input, and cache-write rates double, while the output rate increases by 50%. A request with 300,000 input tokens is therefore billed entirely at the long-context rate.

Prices are in U.S. dollars per 1 million tokens. Batch and Flex processing cost 50% of Standard rates. Fast processing costs twice the applicable Standard rate, including the higher rate for prompts above the 272K long-context threshold.

ProcessingInputCached InputCache WriteOutput
Standard, up to 272K input$10.00$1.00$12.50$50.00
Standard, over 272K input$20.00$2.00$25.00$75.00
Batch / Flex, up to 272K input$5.00$0.50$6.25$25.00
Batch / Flex, over 272K input$10.00$1.00$12.50$37.50
Fast, up to 272K input$20.00$2.00$25.00$100.00
Fast, over 272K input$40.00$4.00$50.00$150.00

OpenAI GPT Models API Pricing

ModelContext WindowInput/1M TokensOutput/1M Tokens
GPT-6.1 Sol
≤272K input
1.05M$2.00$10.00
GPT-6.1 Sol
>272K input
1.05M$4.00$15.00
GPT-6 Astra
≤272K input
1.05M$10.00$50.00
GPT-6 Astra
>272K input
1.05M$20.00$75.00
GPT-6 Astra Ultrafast
≤272K input
1.05M$60.00$300.00
GPT-6 Astra Ultrafast
>272K input
1.05M$120.00$450.00
GPT-6 Sol
≤272K input
1.05M$2.00$10.00
GPT-6 Sol
>272K input
1.05M$4.00$15.00
GPT-6 Luna
≤272K input
1.05M$0.10$0.50
GPT-6 Luna
>272K input
1.05M$0.20$0.75
GPT-5.6 Sol
≤272K input
1.05M$4.00$20.00
GPT-5.6 Sol
>272K input
1.05M$8.00$30.00
GPT-5.6 Terra
≤272K input
1.05M$2.00$12.00
GPT-5.6 Terra
>272K input
1.05M$4.00$18.00
GPT-5.6 Luna
≤272K input
1.05M$0.20$1.20
GPT-5.6 Luna
>272K input
1.05M$0.40$1.80
GPT-5.6 Cyber400K$12.50$75.00
gpt-5.5
≤272K input
1.05M$5.00$30.00
gpt-5.5
>272K input
1.05M$10.00$45.00
gpt-5.5-pro1.05M$30.00$180.00
gpt-5.4
≤272K input
1.05M$2.50$15.00
gpt-5.4
>272K input
1.05M$5.00$22.50
gpt-5.4-pro
≤272K input
1.05M$30.00$180.00
gpt-5.4-pro
>272K input
1.05M$60.00$270.00
gpt-5.4-mini400K$0.75$4.50
gpt-5.4-nano400K$0.20$1.25
gpt-5.2400K$1.75$14.00
gpt-5.2-pro400K$21$168
gpt-5.1400K$1.25$10.00
gpt-5400K$1.25$10.00
gpt-5-mini400K$0.25$2.00
gpt-5-nano400K$0.05$0.40
gpt-5-pro400K$15.00$120.00
gpt-5.3-chat-latest400K$1.75$14.00
gpt-5.3-codex–$1.75$14.00
gpt-5.1-codex-max–$1.25$10.00
gpt-5.1-codex-mini–$0.25$6.00
codex-mini-latest–$1.50$6.00
gpt-5-search-api–$1.25$10.00
gpt-4.11M$2.00$8.00
gpt-4.1-mini1M$0.40$1.60
gpt-4.1-nano1M$0.10$0.40
gpt-4o128K$2.50$10.00
gpt-4o-mini128K$0.15$0.60
From OpenAI

OpenAI Image Generation API Pricing

ModelInput/1M TokensOutput/1M Tokens
GPT-image-2.5-sunburst$5.00 (text)
$8.00 (image)
$30.00
GPT-image-2.5-flare$5.00 (text)
$8.00 (image)
$30.00
GPT-image-2$5.00 (text)
$8.00 (image)
$30.00
GPT-image-1.5$5.00$10.00
GPT-image-1$5.00–
GPT-image-1-mini$2.00–
From OpenAI

OpenAI Voice & Audio API Pricing

ModelInput/1M TokensOutput/1M Tokens
gpt-live-1–$0.05 / minute
gpt-realtime-2.1$4.00 (text)
$5.00 (image)
$32.00 (audio)
$24.00 (text)
$64.00 (audio)
gpt-realtime-2.1-mini$0.60 (text)
$0.80 (image)
$10.00 (audio)
$2.40 (text)
$20.00 (audio)
gpt-realtime-translate–$0.034 / minute
gpt-transcribe–$0.0045 / minute
gpt-live-transcribe–$0.017 / minute
gpt-realtime-whisper–$0.017 / minute
gpt-transcribe–$0.0045 / minute
From OpenAI

OpenAI Reasoning Models API Pricing (Legacy)

ModelContext WindowInput/1M TokensOutput/1M Tokens
o4-mini200K$1.10$4.40
o4-mini-deep-research200K$2.00$8.00
o3-pro200K$20.00$80.00
o3200K$2.00$8.00
o3-deep-research200K$10.00$40.00
o3-mini200K$1.10$4.40
o1-pro200K$150.00$600.00
o1200K$15.00$60.00
o1-mini128K$1.10$4.40
From OpenAI

OpenAI Video Generation API Pricing

ModelPrice per second
Sora 2$0.10
Sora 2 Pro$0.30
Sora 2 Pro (Portrait: 1024 x 1792. Landscape: 1792 x 1024)$0.50
From OpenAI

Claude API Pricing

ModelContext WindowInput/1M TokensOutput/1M Tokens
Claude Fable 5.1
Claude Mythos 5.1
(limited availability)
1M$10.00$50.00
Claude Opus 5.51M$4.00$20.00
Claude Opus 5.5 Fast Mode1M$8.00$40.00
Claude Opus 51M$5.00$25.00
Claude Opus 5 Fast Mode1M$10.00$50.00
Claude Opus 4.81M$5.00$25.00
Claude Opus 4.8 Fast Mode1M$10.00$50.00
Claude Sonnet 5.51M$2.00$10.00
Claude Haiku 4.5200K$1.00$5.00
From Anthropic

Gemini API Pricing

ModelContext WindowInput/1M TokensOutput/1M Tokens
Gemini 3.8 Flash1M$0.75$3.75
Gemini 3.7 Flash1M$0.75$3.75
Gemini 3.6 Flash1M$0.75$3.75
Gemini 3.5 Flash1M$1.50$9.00
Gemini 3.5 Flash-Lite1M$0.30$2.50
Gemini 3.5 Transcribe Live–$3.50 or $0.005/min$21.00 or $0.004/min
Gemini 3.5 Transcribe–$2.00 or $0.003/min$12.00 or $0.002/min
Gemini 3.1 Pro Preview
≤200K input
1M$2.00$12.00
Gemini 3.1 Pro Preview
>200K input
1M$4.00$18.00
Gemini 3.1 Flash-Lite1M$0.25 (text/image/video)
$0.50 (audio)
$1.50
Gemini 3.1 Flash Live–$0.75 (text)
$3.00 or $0.005/min (audio)
$1.00 or $0.002/min (image/video)
$4.50 (text)
$12.00 or $0.018/min (audio)
Gemini 3 Pro Preview (retired Mar. 9, 2026)
≤200K input
1M$2.00$12.00
Gemini 3 Pro Preview (retired Mar. 9, 2026)
>200K input
1M$4.00$18.00
Gemini 3 Flash Preview1M$0.50 (text / image / video)
$1.00 (audio)
$3.00
Gemini 2.5 Pro
≤200K input
1M$1.25$10.00
Gemini 2.5 Pro
>200K input
1M$2.50$15.00
Gemini 2.5 Flash1M$0.30 (text/image/video)
$1.00 (audio)
$2.50
Gemini 2.5 Flash-Lite1M$0.10 (text/image/video)
$0.50 (audio)
$0.40
Gemini 2.0 Flash1M$0.10
$0.70 (audio)

$0.40
Gemini 2.0 Flash-Lite1M$0.075$0.30
From Google

Nano Banana Pricing

ModelFree TierInput/1M TokensOutput/1M Tokens
Nano Banana 2 Lite
(Gemini 3.1 Flash Lite Image)
Not available$0.25$1.50 (text and thinking)
$30.00 (images)
$0.034 per 1K image
Nano Banana 2
(Gemini 3.1 Flash Image)
Not available$0.50$3.00 (text and thinking)
$60.00 (images)
$0.067 per 1K image
$0.101 per 2K image
$0.151 per 4K image
Nano Banana 2 Pro
(Gemini 3 Pro Image)
Not available$2.00 (text)
$0.0011 (image)
$12.00 (text and thinking)
$120.00 (images)
$0.134 per 1K/2K image
$0.24 per 4K image
From Google

Gemini Omni Pricing

ModelFree TierInput/1M TokensOutput/1M Tokens
Gemini Omni 1.1 FlashNot available$1.50 (text/image/video/audio)$9.00 (text)
$17.50 (video)
From Google

Gemini Live

ModelFree TierInputOutput
Gemini 3.8 Live
Gemini 3.8 Live Extended Thinking
Free of charge$0.75 (text)
$3.00 or $0.005/min (audio)
$1.00 or $0.002/min (image/video)
$4.50 (text)
$12.00 or $0.018/min (audio)
Gemini 3.5 Live TranslateFree of charge$3.50 or $0.0053/min$21.00 or $0.0315/min
From Google

Gemini 2.5 Flash Native Audio API Pricing

ModelFree TierInput/1M TokensOutput/1M Tokens
Gemini 2.5 Flash Native AudioNot available$0.50 (text)
$3.00 (audio/video)
$2.00 (text)
$12.00 (audio)
From Google

Gemini 2.5 Flash Image Preview Pricing

ModelFree TierInput/1M TokensOutput
Gemini 2.5 Flash Image PreviewNot available$0.30 (text/image)$0.039 per image
From Google

Gemini TTS Pricing

ModelFree TierInput/1M TokensOutput/1M Tokens
Gemini 3.1 Flash TTS PreviewFree of charge$1.00 (text)$20.00 (audio)
Gemini 2.5 Flash Preview TTSFree of charge$0.50 (text)$10.00 (audio)
Gemini 2.5 Pro Preview TTSNot available$1.00 (text)$20.00 (audio)
From Google

Google Imagen 4 & 3 API Pricing

ModelPaid Tier, per Image in USD
Imagen 4 Fast$0.02
Imagen 4 Standard$0.04
Imagen 4 Ultra$0.06
Imagen 3$0.03
From Google

Gemini Embedding API Pricing

ModelPaid Tier, per 1M tokens in USD
gemini-embedding-2$0.20 (Text)
$0.45 (Image $0.00012 per image)
$6.50 (Audio $0.00016 per second)
$12.00 (Video $0.00079 per frame)
gemini-embedding-001$0.15
From Google

Google Veo 3.1 API Pricing

ModelPaid Tier, per second in USD
Veo 3.1 Standard$0.40 (720p and 1080p)
$0.60 (4k)
Veo 3.1 Fast$0.15 (720p and 1080p)
$0.35 (4k)
Veo 3.1 Light$0.05 (720p)
$0.08 (1080p)
From Google

Grok API Pricing

ModelContext WindowInput/1M TokensOutput/1M Tokens
Grok 4.7
<200K input
500K$2.00$6.00
Grok 4.7
≥200K input
500K$4.00$12.00
Grok 4.6
<200K input
500K$2.00$6.00
Grok 4.6
≥200K input
500K$4.00$12.00
Grok 4.5
<200K input
500K$2.00$6.00
Grok 4.5
≥200K input
500K$4.00$12.00
Grok Build 0.1
<200K input
256K$1.00$2.00
Grok Build 0.1
<200K input
256K$2.00$4.00
Grok 4.31M$1.25$2.50
Grok 4.20 Multi-agent2M$2.00$6.00
Grok 4.20 Reasoning2M$2.00$6.00
Grok 4.202M$2.00$6.00
Imagine Image 2.0–0.01/image0.04/image (1K Low Quality)
0.06/image (1K)
0.06/image (2K Low Quality)
0.08/image (2K)
Grok Imagine Image Quality–0.01/image0.05/image (1K)
0.07/image (2K)
Grok Imagine Image–0.002/image0.02/image (1K)
0.02/image (2K)
Imagine Video 1.5–0.01/second0.08/sec (480p)
0.14/sec (720p)
0.25/sec (1080p)
Grok Imagine Video–0.01/sec
0.002/image
0.05/sec (480p)
0.07/sec (720p)
grok-voice-think-fast-2.0––$0.08/min ($4.80/hr) audio
$0.004/text input
grok-voice-think-fast-1.0––$0.05/min ($3.00/hr) audio
$0.004/text input
Text to Speech––$15.00 / 1M chars
Speech to Text
–
–$0.10 / hr (REST)
$0.20 / hr (Streaming)
From xAI

DeepSeek API Pricing

DeepSeek plans to discontinue the V4 Pro service at 12:00 Beijing Time on September 14, 2026. At that time, all requests to the Pro model will be routed to V4.1 Flash and billed at Flash’s price. After extensive internal and external testing, V4.1 Flash has comprehensively surpassed V4 Pro across all key metrics, including performance, cost, speed, and task completion time.

ModelContext WindowInput/1M TokensOutput/1M Tokens
DeepSeek-V4.1-Flash1M$0.15 (OFF-PEAK)
$0.30 (PEAK)
$0.60 (OFF-PEAK)
$1.20 (PEAK)
DeepSeek-V4-Pro-08131M$0.66 (OFF-PEAK)
$1.32 (PEAK)
$1.98 (OFF-PEAK)
$3.96 (PEAK)
From DeepSeek

Qwen API Pricing

ModelContext WindowInput/1M TokensOutput/1M Tokens
Qwen3.8-Max1M$2.00$6.00
Qwen3.8-Flash1M$0.16$0.47
Qwen3.7-Plus1M$0.40$1.60
Qwen3.7-Flash1M$0.03$0.13
Qwen3.8-Omni-Flash–$0.15$0.47
Qwen3.5-Omni-Plus–$11.00 (Audio)
$1.40 (Text/Image/Video)
$44.00 (Text & Audio)
$8.30 (Text)
Qwen Image 3.0 Pro–$0.03/IMG$0.075/IMG (1K)
$0.04/sec (2K)
Wan 3.0 Video––$0.05/sec (480p)
$0.10/sec (720p)
$0.20/sec (1080p)
Wan2.7-Image-Pro––$0.075/IMG Generation
HappyHorse-1.1-T2V––$0.07/sec (480p)
$0.14/sec (720p)
$0.24/sec (1080p)
From Qwen

Mistral (Premier Models) API Pricing

ModelContext WindowInput/1M TokensOutput/1M Tokens
Magistral Medium128K$2.00$5.00
Codestral32K$0.30$0.90
Document AI & OCR 4.1––OCR: $4/1000 pages
Document: $5/1000 pages
Voxtral Mini Transcribe 2––Audio Input/min
$0.003
Voxtral Realtime––Audio Input/min
$0.006
Mistral Embed32k$0.10–
Mistral Moderation 24.1132k$0.10
Magistral Small128K$0.50$1.50
Codestral Embed128K$0.15–
From Mistral

Mistral (Open Models) API Pricing

ModelContext WindowInput/1M TokensOutput/1M Tokens
Mistral Medium 3.5256K$1.50$7.50
Mistral Large 3256K$0.50$1.50
Ministral 3 – 3B256K$0.10$0.10
Ministral 3 – 8B256K$0.15$0.15
Ministral 3 – 14B256K$0.20$0.20
Pixtral Large128K$2.00$6.00
Pixtral 12B128K$0.15$0.15
Mistral Nemo128K$0.15$0.15
Mistral Small 4.0128K$0.15$0.60
Mistral Small 3.2128K$0.10$0.30
Devstral 2 (deprecated)256K$0.40$2.00
Devstral Small 2 (deprecated)256K$0.10$0.30
Voxtral Mini–$0.001 (audio)
$0.04 (text)
$0.04
Voxtral Small–$0.004 (audio)
$0.1 (text)
$0.03
Voxtral TTS–$0.016/1k Chars
Mistral 7B32K$0.25$0.25
Mixtral 8x7B32K$0.70$0.70
Mixtral 8x22B64K$2.00$6.00
From Mistral

Meta AI Muse API Pricing

ModelContext WindowInput/1M TokensOutput/1M Tokens
Muse Spark 1.3 (Contributor)1M$0.10$0.20
Muse Spark 1.31M$1.25$4.25
Muse Spark 1.2 (Contributor)1M$0.10$0.20
Muse Spark 1.21M$1.25$4.25
Muse Spark 1.11M$1.25$4.25
Muse Image––$0.01/image
Muse Voice Transcribe––$3.00/1,000 min
$0.18/hour
From Meta AI

Llama 4 & 3 API Pricing

ModelContext WindowInput/1M TokensOutput/1M Tokens
Llama 4 Scout10M$0.11$0.34
Llama 4 Maverick10M$0.20$0.60
Llama 3.3 70B Versatile128K$0.59$0.79
Llama 3.3 70B SpecDec8192$0.59$0.99
Llama 3.3 70b Instruct128K$0.23$0.40
Llama 3.3 70b Instruct-Turbo128K$0.13$0.40
Llama 3.2 90b Vision-Instruct128K$0.35$0.40
Llama 3.2 11b Vision-Instruct128K$0.055$0.055
Llama 3.1 405B128K$1.79$1.79
Llama 3.1 70B128K$0.35$0.40
Llama 3.1 8B128K$0.09$0.09
From Groq

GLM API Pricing

ModelContext WindowInput/1M TokensOutput/1M Tokens
GLM-5.31M$1.40$4.40
GLM-5.3 Flash1M$0.075$0.25
GLM-5.21M$1.40$4.40
GLM-5.1200K$1.40$4.40
GLM-5200K$1.00$3.20
GLM-5-Turbo200K$1.20$4.00
GLM-5-Code200K$1.20$5.00
GLM-5V-Turbo200K$1.20$4.00
GLM-OCR–$0.03$0.03
GLM-4.7128K$0.60$2.20
GLM-4.7-FlashX128K$0.07$0.40
GLM-4.6128K$0.60$2.20
GLM-4.6v128K$0.30$0.90
GLM-4.6V-FlashX128K$0.004$0.40
GLM-4.6V-Flash128KFREEFREE
GLM-4.5128K$0.60$2.20
GLM-4.5v64K$0.60$1.80
GLM-4.5-X128K$0.45$8.90
GLM-4.5-Air128K$0.20$1.10
GLM-4.5-AirX128K$1.10$4.50
GLM-4.5-Flash128KFREEFREE
From z.ai

Kimi K3 & K2 API Pricing

ModelContext WindowInput/1M TokensOutput/1M Tokens
kimi-k31M$3.00$15.00
kimi-k2.7-code-highspeed1M$1.90$8.00
kimi-k2.7-code1M$0.95$4.00
kimi-k2.6262,144$0.95$4.00
kimi-k2.5262,144$0.60$3.00
kimi-k2-thinking262,144$0.60$2.50
kimi-k2-thinking-turbo262,144$1.15$8.00
kimi-k2-0905-preview262,144$0.60$2.50
kimi-k2-0711-preview131K$0.60$2.50
kimi-k2-turbo-preview262,144$1.15$8.00
From moonshot

MiniMax M3 & M2.7 API Pricing

ModelContext WindowInput/1M TokensOutput/1M Tokens
Minimax M3
≤512K input
1M$0.30$1.20
Minimax M3
>512K input
1M$0.60$2.40
Minimax M2.7204,800$0.30$1.20
Minimax M2.7 Highspeed204,800$0.60$2.40
From Minimax

Minimax Video/Audio/Image Generation API Pricing

ModelUnit Price
MiniMax H3$0.13/s, 2K
MiniMax H3$0.09/s, 768P
Music-3.0$0.15/up-to-5 minutes music
image-01$0.0035 per image
From Minimax

Perplexity API Pricing

ModelContext WindowInput/1M TokensOutput/1M Tokens
Sonar128K$1.00$1.00
Sonar Pro200K$3.00$15.00
Sonar Reasoning Pro200K$2.00$8.00
Sonar Deep Research128K$2.00$8.00
From Perplexity

Cohere API Pricing

ModelContext WindowInput/1M TokensOutput/1M Tokens
Command A256K$2.50$10.00
Command R128K IN/4K OUT$0.15$0.60
Command R7B128K$0.0375$0.15
Transcribe––$3.75/hour/instance
Embed 4128K–$0.12 (Text)
$0.47 (Image)
Rerank 4 Fast32K–$2.00/1K searches
Rerank 4 Pro32K–$2.50/1K searches
Parse 5––$1.50/1K pages
From Cohere

See Also:

LLM API Pricing FAQ

How is LLM API pricing calculated?

Most text LLM APIs charge according to the number of input and output tokens processed. Multiply the input token count by the model’s input rate, multiply the output token count by its output rate, and combine the two amounts.

The final bill can also include cached tokens, long-context rates, tool calls, search or grounding requests, and multimodal inputs or outputs.

Are ChatGPT, Claude, and Gemini subscriptions the same as API pricing?

No. Consumer subscriptions and API usage use different billing systems.

A ChatGPT, Claude, or Gemini subscription pays for access to the company’s consumer application and the features included with that plan. API access is generally billed through a developer platform according to token usage or other API-specific units.

Why can long-context API requests cost more?

Long prompts consume more input tokens, so they cost more even when the per-token rate stays unchanged. Some models also apply a higher pricing tier after the request crosses a specified context threshold.

For example, GPT-6 Astra and GPT-5.6 apply higher rates when input exceeds 272K tokens. Gemini 3.1 Pro Preview uses higher rates above 200K input tokens.

Do cached tokens cost the same as normal input tokens?

No. Supported APIs can charge a lower rate when the model reuses eligible prompt content from a cache instead of processing the same input from scratch.

Changelog:

09/29/2026

  • Added GPT-6.1 Sol

09/28/2026

  • Added Sonnet 5.5

09/22/2026

  • Added Grok 4.7
  • Added Opus 5.5
  • Added GPT‑6 Sol and Luna

09/18/2026

  • Updated Grok models

09/17/2026

  • Added Qwen3.8-Omni-Flash

09/15/2026

  • Added Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking

09/10/2026

  • Added DeepSeek-V4.1-Flash
  • Added gpt-live-1

09/08/2026

  • Corrected context-window and long-context pricing threshold labels for GPT and Gemini models.
  • Updated MiniMax M3, Cohere Command A, and Devstral 2 API pricing.
  • Added GPT-Image-2.5 Flare and Sunburst.

09/03/2026

  • Added GPT-6 Astra

09/02/2026

  • Added Muse Voice Transcribe
  • Added Claude Fable 5.1 and Claude Mythos 5.1
  • Added Muse Spark 1.3

08/27/2026

  • Added Gemini Omni 1.1 Flash
  • Added cohere models

08/26/2026

  • Added GLM-5.3-Flash and Qwen3.8-Flash
  • Added Muse Image and Gemini 3.5 Transcribe

08/21/2026

  • Added DeepSeek-v4-flash-vision-exp.
  • Updated GPT-5.6 Sol.

08/19/2026

  • Added GLM 5.3.

08/16/2026

  • Updated DeepSeek API pricing.

08/13/2026

  • Added Mistral OCR 4.1
  • Added Gemini 3.7 Flash
  • Added GLM 5.3
  • Added Minimax Music-3.0

08/12/2026

  • Added DeepSeek-V4-Pro-0813
  • Added Grok 4.6

08/11/2026

  • Added Grok Imagine Image 2.0

08/06/2026

  • Added Wan 3.0 Video

08/05/2026

  • Added Muse Spark 1.2

08/02/2026

  • Added Qwen3.8-Max

07/31/2026

  • Added Minimax H3.

07/30/2026

  • Updated GPT 5.6 pricing.

07/29/2026

  • Added Grok Voice Think Fast 2.0

07/28/2026

  • Cleanup
  • Added GPT-Live-Transcribe and GPT-Transcribe

07/24/2026

  • Added Claude Opus 5

07/21/2026

  • Added Gemini 3.6 Flash and 3.5 Flash-Lite

07/16/2026

  • Added Kimi K3.

07/09/2026

  • Updated for GPT-5.6.

07/08/2026

  • Added Grok 4.5.

07/06/2026

  • Added gpt-realtime-2.1-mini

07/01/2026

  • Added kimi-k2.7-code-highspeed

06/30/2026

  • Added Claude Sonnet 5
  • Added Nano Banana Lite

06/26/2026

  • Added GPT 5.6

06/16/2026

  • Added OCR 4

06/16/2026

  • Added GLM-5.2

06/13/2026

  • Updated Minimax M3

06/12/2026

  • Added kimi-k2.7-code

06/09/2026

  • Added Claude Fable/Mythos 5.
  • Added Gemini 3.5 Live Translate.

06/07/2026

  • Updated Deepseek API.

06/04/2026

  • Added grok-imagine-video-1.5-preview

06/01/2026

  • Added MiniMax M3
  • Added Qwen3.7-Plus and HappyHorse-1.0-T2V

05/28/2026

  • Added Claude Opus 4.8

05/21/2026

  • Added Qwen3.7-Max

05/19/2026

  • Added Gemini 3.5 Flash.

05/13/2026

  • Added Claude Opus 4.7 Fast Mode

05/08/2026

  • Added grok-imagine-image-quality

05/07/2026

  • Added OpenAI audio models

05/01/2026

  • Added Grok 4.3

04/29/2026

  • Updated for Mistral Medium 3.5

04/27/2026

  • Updated for DeepSeek v4

04/25/2026

  • Updated

04/24/2026

  • Added Deeseek V4 models

04/23/2026

  • Added GPT-5.5

04/22/2026

  • Added Gemini Embedding 2

04/22/2026

  • Added GPT-image-2

04/21/2026

  • Added Qwen3.6-Max-Preview

04/20/2026

  • Added Kimi K2.6

04/18/2026

  • Cleanup Grok APIs.

04/16/2026

  • Added Claude Opus 4.7

04/15/2026

  • Added Gemini 3.1 Flash TTS Preview

04/08/2026

  • Added GLM-5.1

04/01/2026

  • Added Perplexity Sonar models.
  • GLM-5V-Turbo.

03/31/2026

  • Added Veo 3.1 Light

03/26/2026

  • Added Voxtral TTS and Gemini 3.1 Flash Live

03/23/2026

  • Updated for Grok 4.2-0309

03/18/2026

  • Added Minimax M2.7.

03/17/2026

  • Added GPT-5.4 mini and nano

03/16/2026

  • Added Mistral Small 4.0

03/16/2026

  • Added GLM-5-Turbo

03/05/2026

  • Added GPT-5.4

03/03/2026

  • Added Gemini 3.1 Flash-Lite
  • Added GPT-5.3 Instant

02/26/2026

  • Added Nano Banana 2 (Gemini 3.1 Flash Image Preview)

02/19/2026

  • Added Gemini Pro 3.1

02/18/2026

  • Added Sonnet 4.6

02/11/2026

  • Added GLM-5

02/06/2026

  • Added Opus 4.6 & Voxtral Mini Transcribe 2

01/29/2026

  • Added grok-imagine-video and Kimi 2.5.

12/23/2025

  • Updated GLM-4.7

12/17/2025

  • Added Gemini 3 Flash

12/16/2025

  • Added GPT-image-1.5

12/11/2025

  • Added GPT 5.2

12/09/2025

  • Added Devstral 2

12/02/2025

  • Added Mistral Large 3

12/01/2025

  • Updated for DeepSeek-V3.2-Speciale

11/24/2025

  • Updated for Claude Opus 4.5

11/20/2025

  • Added Nano Banana Pro.
  • Added Grok 4.1 Fast.
  • Removed legacy models.

11/19/2025

  • Added more models

11/18/2025

  • Added Gemini 3 Pro

11/14/2025

  • Added GPT-5.1
  • Updated Qwen models
  • Updated Kimi models
  • Added Minimax M2

11/08/2025

  • Added codex-mini

10/15/2025

  • Update for Claude Haiku 4.5 & Veo 3.1.

10/07/2025

  • Added Sora 2 and more OpenAI APIs.

09/29/2025

  • Updated for Claude 4.5 and DeepSeek-V3.2

09/19/2025

  • Updated price

08/21/2025

  • Updated for DeepSeek-V3.1

08/20/2025

  • Added GLM-4.5 and Kimi K2 models.

08/14/2025

  • Added Imagen 4 Fast

08/07/2025

  • Added GPT-5

08/05/2025

  • Added Claude Opus 4.1

07/15/2025

  • Added Voxtral model family

07/14/2025

  • Added Gemini Embedding

07/11/2025

  • Added Grok 4

06/26/2025

  • Added o3-deep-research and o4-mini-deep-research

06/25/2025

  • Added Imagen 4 Ultra and Imagen 4 Standard

06/20/2025

  • Added Gemini 2.5 Flash-Lite

06/10/2025

  • Added o3-pro
  • Added Mistral Magistral models.

06/10/2025

  • OpenAI dropped the price of o3 by 80%

05/22/2025

  • Updated for Claude 4

05/21/2025

  • Updated Mistral models

05/21/2025

  • Added Gemini 2.5 Flash Native Audio

05/08/2025

  • Added Mistral Medium 3

04/23/2025

  • Added OpenAI Image Generation

04/18/2025

  • Added Gemini 2.5 Flash
  • Updated Cohere models

04/16/2025

  • Added o4-mini

04/14/2025

  • Added gpt-4.1 family

04/11/2025

  • Updated DeepSeek

04/11/2025

  • Added Google Imagen 3 and Veo 2.

04/10/2025

  • Added Qwen models
  • Added Grok 3

04/05/2025

  • Added Gemini Pro 2.5

03/19/2025

  • Added o1-pro

02/28/2025

  • Added GPT-4.5

02/24/2025

  • Added Claude 3.7

02/21/2025

  • Added Grok

02/06/2025

  • Added Gemini 2.0 Flash and Gemini 2.0 Flash-Lite

02/02/2025

  • Added o3-mini

01/31/2025

  • Added DeepSeek v3 and DeepSeek R1.

12/18/2024

  • o1 in the API comes with support for function calling, developer messages, Structured Outputs, and vision capabilities.

12/07/2024

  • Added Llama 3.3

11/05/2024

  • Added Claude Haiku 3.5

11/03/2024

  • Added Gemini 1.5 Flash-8B

10/04/2024

  • Added Llama 3.2
  • Added gpt-4o-realtime-preview

09/25/2024

  • Updated Google Gemini

09/13/2024

  • Added OpenAI’s latest model: o1.

08/07/2024

  • Added gpt-4o-2024-08-06, the latest gpt-4o snapshot that supports Structured Outputs

07/25/2024

  • Updated Mistral Large 2

07/24/2024

  • Added Llama 3.1 405B

07/20/2024

  • Added GPT-4o-mini
  • Updated prices

07/13/2024

  • Updated
  • Added Cohere’s Command API

Leave a Reply

Your email address will not be published. Required fields are marked *

Get the latest & top AI tools sent directly to your email.

Subscribe now to explore the latest & top AI tools and resources, all in one convenient newsletter. No spam, we promise!