OpenAI API Rate Limits: RPM, TPM, Tiers, and 429 Errors (2026)

Understand OpenAI API rate limits, including RPM, TPM, RPD, TPD, usage tiers, account-specific limits, and practical fixes for 429 errors.

OpenAI API rate limits control how many requests, tokens, images, audio minutes, and batch jobs an organization or project can process within a fixed period.

The active limits depend on the model, endpoint, usage tier, organization, project, request size, and shared limit group. They are separate from model context windows, maximum output tokens, ChatGPT message caps, and monthly spend controls.

This post explains RPM, TPM, RPD, TPD, and IPM; lists verified limits for current and previous GPT, reasoning, image, video, realtime, audio, embedding, and moderation models; shows where to check account-specific limits; and covers practical fixes for 429 errors.

Last verified: September 29, 2026.

Understanding the OpenAI API Rate Limits

OpenAI measures API throughput through several independent limits. A request fails as soon as it exceeds any applicable limit.

API rate limits apply to API organizations and projects. ChatGPT plan caps, context windows, and maximum output tokens use different limits. See ChatGPT and OpenAI token limits for those values.

LimitMeaningWhat it controlsCommon trigger
RPMRequests per minuteAPI calls sent within a minuteMany short requests or high concurrency
RPDRequests per dayTotal daily API callsLarge daily job volume
TPMTokens per minuteInput and output token throughputLong prompts, large outputs, or many parallel calls
TPDTokens per dayTotal daily token throughputHigh-volume production workloads
IPMImages per minuteImage generations and editsParallel image jobs
Audio minutes per minuteAudio throughputAudio processed by eligible streaming modelsConcurrent voice sessions
Batch queue limitQueued input tokens per modelPrompt tokens waiting in Batch API jobsLarge pending batch workloads
An application can stay below TPM and hit RPM first. A long prompt can stay below RPM and hit TPM first. Batch jobs can stop accepting work after queued input tokens reach the model-specific queue limit.

Current OpenAI API Rate Limits: Quick Reference

OpenAI can assign different limits to an organization, project, long-context workload, or shared model group. Check the Platform Limits dashboard before setting production capacity or concurrency targets.

GPT-6 Astra Rate Limits

Usage tierRPMTPMBatch queue limit
FreeNot supportedNot supportedNot supported
Tier 1500500,0001,500,000
Tier 25,0001,000,0003,000,000
Tier 35,0002,000,000100,000,000
Tier 410,0004,000,000200,000,000
Tier 515,00040,000,00015,000,000,000

GPT-6.1 Sol Rate Limits

Usage tierRPMTPMBatch queue limit
FreeNot supportedNot supportedNot supported
Tier 1500500,0001,500,000
Tier 25,0001,000,0003,000,000
Tier 35,0002,000,000100,000,000
Tier 410,0004,000,000200,000,000
Tier 515,00040,000,00015,000,000,000

GPT-6 Luna Rate Limits

Usage tierRPMTPMBatch queue limit
FreeNot supportedNot supportedNot supported
Tier 1500500,0005,000,000
Tier 25,0002,000,00020,000,000
Tier 35,0004,000,00040,000,000
Tier 410,00010,000,0001,000,000,000
Tier 530,000180,000,00015,000,000,000

GPT-5.6 Sol, Terra, and Luna Synchronous Limits

Usage tierGPT-5.6 Sol RPM / TPMGPT-5.6 Terra RPM / TPMGPT-5.6 Luna RPM / TPM
FreeNot supportedNot supportedNot supported
Tier 1500 / 500,000500 / 500,000500 / 500,000
Tier 25,000 / 1,000,0005,000 / 1,000,0005,000 / 2,000,000
Tier 35,000 / 2,000,0005,000 / 2,000,0005,000 / 4,000,000
Tier 410,000 / 4,000,00010,000 / 4,000,00010,000 / 10,000,000
Tier 515,000 / 40,000,00015,000 / 40,000,00030,000 / 180,000,000

GPT-5.6 Batch Queue Limits

Batch queue limits count input tokens waiting in pending Batch jobs. Tokens leave the queue count after the related batch completes.

Usage tierGPT-5.6 SolGPT-5.6 TerraGPT-5.6 Luna
FreeNot supportedNot supportedNot supported
Tier 11,500,0001,500,0005,000,000
Tier 23,000,0003,000,00020,000,000
Tier 3100,000,000100,000,00040,000,000
Tier 4200,000,000200,000,0001,000,000,000
Tier 515,000,000,00015,000,000,00015,000,000,000

GPT Image 2.5 Rate Limits

Usage tierTPMIPM
FreeNot supportedNot supported
Tier 1100,0005
Tier 2250,00020
Tier 3800,00050
Tier 43,000,000150
Tier 58,000,000250

OpenAI Usage Tiers

OpenAI moves organizations to higher usage tiers as paid API spend increases. Higher tiers usually raise rate limits across most models. Each model keeps its own rate-limit profile.

TierQualificationMonthly usage limit
FreeUser must be in an allowed geography$100 per month
Tier 1$5 paid$100 per month
Tier 2$50 paid$500 per month
Tier 3$100 paid$1,000 per month
Tier 4$250 paid$5,000 per month
Tier 5$1,000 paid$200,000 per month

OpenAI API Free Tier Rate Limits

The API Free tier applies to eligible users in allowed geographies. Model access and exact limits depend on the API account and model.

GPT-6 Astra, GPT-6.1 Sol, GPT-6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna, GPT-4.1, GPT-4o, GPT Image 2.5, and current Realtime models list Free-tier access as Not supported in their published model tables.

The Platform Limits page shows the Free-tier models available to your account. The following specialized models have published Free-tier limits:

Model or categoryFree-tier limits
text-embedding-3-large and text-embedding-3-small100 RPM, 2,000 RPD, 40,000 TPM
omni-moderation-latest250 RPM, 5,000 RPD, 10,000 TPM
TTS-13 RPM, 200 RPD
Whisper-13 RPM, 200 RPD

OpenAI API Limits vs ChatGPT Free and Plus Limits

OpenAI API limits and ChatGPT plan limits use separate products, billing systems, and measurements. A ChatGPT subscription does not include API usage or increase an API organization’s rate limits.

Limit typeApplies toTypical measurementsWhere to check
OpenAI API rate limitsAPI organizations and projectsRPM, TPM, RPD, TPD, IPM, batch queueOpenAI Platform Limits page
API free usage tierEligible API accounts and supported modelsModel-specific API throughput and monthly usage limitOpenAI Platform Limits page
ChatGPT Free, Go, Plus, and Pro limitsChatGPT plansMessages, model access, tools, and feature usageChatGPT interface and plan documentation

For context windows, maximum output tokens, and ChatGPT plan limits, see ChatGPT and OpenAI token limits.

Where to Check Your Current OpenAI Rate Limits

Use the OpenAI Platform dashboard for account-specific values:

  1. Open OpenAI Platform Limits.
  2. Select the correct organization.
  3. Select the project used by the application.
  4. Review limits by model and endpoint.
  5. Check shared model groups, long-context limits, and project-level token limits.
  6. Use the increase request option when it appears and the workload requires more capacity.

Rate-Limit Response Headers

Production applications should log rate-limit response headers. These headers show the active request and token budget for a call.

HeaderWhat it shows
Retry-AfterMinimum wait in seconds before retrying an eligible temporary 429 error
x-ratelimit-limit-requestsMaximum request allowance in the active window
x-ratelimit-remaining-requestsRequests left in the active window
x-ratelimit-reset-requestsTime until the request allowance resets
x-ratelimit-limit-tokensMaximum token allowance in the active window
x-ratelimit-remaining-tokensTokens left in the active window
x-ratelimit-reset-tokensTime until the token allowance resets
x-ratelimit-limit-project-tokensProject-scoped token limit when one applies
x-ratelimit-remaining-project-tokensProject-scoped tokens left
x-ratelimit-reset-project-tokensTime until the project token allowance resets

How OpenAI API Rate Limits Work

  • Organization and project scope: Limits are defined at the organization and project level.
  • Model-specific capacity: Each model has its own tier table.
  • Shared model groups: Related models can draw from one RPM or TPM pool.
  • Long-context limits: Large-context requests can have a lower throughput limit.
  • Project token limits: A project-level token limit can be lower than the organization-wide model limit.
  • Monthly usage limits: OpenAI assigns an approved monthly organization usage limit. Organization and project spend controls use different settings.
  • Vector store ingestion: File and file-batch ingestion endpoints share a 300 RPM limit for each vector store ID.

Previous and Specialized OpenAI Model Rate Limits

Earlier GPT-5 Model Rate Limits

Model or model familyTier 1Tier 2Tier 3Tier 4Tier 5
GPT-5.5, GPT-5.4, GPT-5.2, GPT-5, GPT-5.3 Codex500 / 500,0005,000 / 1,000,0005,000 / 2,000,00010,000 / 4,000,00015,000 / 40,000,000
GPT-5.5 Pro50 / 50,000500 / 200,000500 / 500,0001,000 / 1,000,0002,000 / 4,000,000
GPT-5.4 Pro, GPT-5.2 Pro, GPT-5 Pro500 / 30,0005,000 / 450,0005,000 / 800,00010,000 / 2,000,00010,000 / 30,000,000
GPT-5.4 Mini, GPT-5 Mini500 / 500,0005,000 / 2,000,0005,000 / 4,000,00010,000 / 10,000,00030,000 / 180,000,000
GPT-5.4 Nano, GPT-5 Nano500 / 200,0005,000 / 2,000,0005,000 / 4,000,00010,000 / 10,000,00030,000 / 180,000,000

GPT-4.1, GPT-4o, and GPT-4 Rate Limits

ModelStatusTier 1Tier 2Tier 3Tier 4Tier 5
GPT-4.1Current500 / 30,0005,000 / 450,0005,000 / 800,00010,000 / 2,000,00010,000 / 30,000,000
GPT-4.1 miniCurrent500 / 200,0005,000 / 2,000,0005,000 / 4,000,00010,000 / 10,000,00030,000 / 150,000,000
GPT-4.1 nanoDeprecated500 / 200,0005,000 / 2,000,0005,000 / 4,000,00010,000 / 10,000,00030,000 / 150,000,000
GPT-4oCurrent500 / 30,0005,000 / 450,0005,000 / 800,00010,000 / 2,000,00010,000 / 30,000,000
GPT-4o miniCurrent500 / 200,0005,000 / 2,000,0005,000 / 4,000,00010,000 / 10,000,00030,000 / 150,000,000
GPT-4Legacy500 / 10,0005,000 / 40,0005,000 / 80,00010,000 / 300,00010,000 / 1,000,000

Earlier O-Series Reasoning Model Rate Limits

Model or model familyTier 1Tier 2Tier 3Tier 4Tier 5
o3, o3-pro, o1, o1-pro500 / 30,0005,000 / 450,0005,000 / 800,00010,000 / 2,000,00010,000 / 30,000,000
o4-mini1,000 / 100,0002,000 / 2,000,0005,000 / 4,000,00010,000 / 10,000,00030,000 / 150,000,000
o3-mini1,000 / 100,0002,000 / 200,0005,000 / 4,000,00010,000 / 10,000,00030,000 / 150,000,000
o1-mini500 / 200,0005,000 / 2,000,0005,000 / 4,000,00010,000 / 10,000,00030,000 / 150,000,000
o3-deep-research500 / 200,0005,000 / 450,0005,000 / 800,00010,000 / 2,000,00010,000 / 30,000,000
o4-mini-deep-research1,000 / 200,0002,000 / 2,000,0005,000 / 4,000,00010,000 / 10,000,00030,000 / 150,000,000

Deprecated Image Model Rate Limits

ModelStatusTier 1Tier 2Tier 3Tier 4Tier 5
GPT Image 1.5Deprecated100,000 TPM / 5 IPM250,000 / 20800,000 / 503,000,000 / 1508,000,000 / 250
GPT Image 1Deprecated100,000 TPM / 5 IPM250,000 / 20800,000 / 503,000,000 / 1508,000,000 / 250
GPT Image 1 miniDeprecated100,000 TPM / 5 IPM250,000 / 20800,000 / 503,000,000 / 1508,000,000 / 250

Deprecated Sora 2 Video API Rate Limits

The Sora API is deprecated and scheduled to shut down on September 24, 2026. These RPM values apply only to existing Sora 2 and Sora 2 Pro API integrations before shutdown.

ModelFreeTier 1Tier 2Tier 3Tier 4Tier 5
Sora 2Not supported25 RPM50 RPM125 RPM200 RPM375 RPM
Sora 2 ProNot supported10 RPM25 RPM50 RPM75 RPM150 RPM

Realtime API Rate Limits

Model groupStatusTier 1Tier 2Tier 3Tier 4Tier 5
GPT-Realtime-2.1, 2.1 mini, 2, 1.5Current200 / 40,000400 / 200,0005,000 / 800,00010,000 / 4,000,00020,000 / 15,000,000
GPT-Realtime, GPT-4o Realtime, GPT-4o mini RealtimeDeprecated200 / 40,000400 / 200,0005,000 / 800,00010,000 / 4,000,00020,000 / 15,000,000

Audio, Speech, and Transcription Rate Limits

ModelFreeTier 1Tier 2Tier 3Tier 4Tier 5
GPT-TranscribeNot supported500 / 200,0005,000 / 2,000,0005,000 / 4,000,00010,000 / 10,000,00030,000 / 150,000,000
gpt-audio-1.5Not supported500 / 30,0005,000 / 450,0005,000 / 800,00010,000 / 2,000,00010,000 / 30,000,000
GPT-4o TranscribeNot supported500 / 10,0002,000 / 100,0005,000 / 400,00010,000 / 2,000,00010,000 / 6,000,000
GPT-4o Transcribe DiarizeNot supported500 / 10,0005,000 / 100,0005,000 / 400,00010,000 / 2,000,00010,000 / 6,000,000
GPT-4o mini TranscribeNot supported500 / 50,0002,000 / 150,0005,000 / 600,00010,000 / 2,000,00010,000 / 8,000,000
Classic audio modelFreeTier 1Tier 2Tier 3Tier 4Tier 5
TTS-13 RPM / 200 RPD500 RPM2,500 RPM5,000 RPM7,500 RPM10,000 RPM
TTS-1 HDNot supported500 RPM2,500 RPM5,000 RPM7,500 RPM10,000 RPM
Whisper-13 RPM / 200 RPD500 RPM2,500 RPM5,000 RPM7,500 RPM10,000 RPM

Embedding API Rate Limits

TierRPMRPDTPMBatch Queue Limit
Free1002,00040,000–
Tier 13,000–1,000,0003,000,000
Tier 25,000–1,000,00020,000,000
Tier 35,000–5,000,000100,000,000
Tier 410,000–5,000,000500,000,000
Tier 510,000–10,000,0004,000,000,000

Moderation API Rate Limits

TierRPMRPDTPM
Free2505,00010,000
Tier 150010,00010,000
Tier 2500–20,000
Tier 31,000–50,000
Tier 42,000–250,000
Tier 55,000–500,000

Rate Limits by API Type

API typeMain limits to monitorTypical constraint
Responses APIRPM, TPM, TPD, shared limits, project token limitsAgents and multimodal workflows consume token capacity through prompts, outputs, and tool calls
Chat Completions APIRPM, TPM, TPDHigh-concurrency chat apps often hit RPM first
Embeddings APIRPM, TPM, payload sizeLarge indexing jobs consume token throughput
Images APITPM, IPMParallel generation and editing jobs hit IPM
Realtime and audio APIsSessions, RPM, TPM, audio throughput, model-specific limitsVoice products require concurrency and session control
Batch APIRequests per batch, file size, batch creation rate, queued input tokensLarge offline workloads need queue planning
Video API (deprecated)Model-specific RPMSora API access is scheduled to end on September 24, 2026
Moderation APIRPM, RPD, TPMText and image moderation uses its own model profile
Fine-tuningActive jobs, queued jobs, jobs per day, inference limitsTraining capacity and inference capacity use different limits
Vector storesIngestion requests per vector storeFile and file-batch ingestion share 300 RPM per vector store ID

Batch API Rate Limits

The Batch API processes asynchronous workloads through its own rate-limit pool. It works for evaluations, classification, embeddings, backfills, extraction, moderation, image jobs, and other jobs that can wait for completion.

Batch limitCurrent value or behavior
Requests per batchUp to 50,000 requests
Input file sizeUp to 200 MB
Embedding inputsUp to 50,000 embedding inputs across all requests in one batch
Batch creation rateUp to 2,000 batches per hour
Queued prompt tokensModel-specific limit shown in Platform settings
Completion windowDesigned to complete within 24 hours
Output-token limitNo Batch-specific output-token limit

Why You Can Get 429 Errors Below the Published Limit

OpenAI calculates token capacity from the larger of the configured maximum token value and an estimate based on request size. Set the output budget close to the expected response length.

  • A short traffic burst exceeds a smaller internal time window.
  • RPM reaches its limit before TPM.
  • TPM is exhausted through long prompts or large output budgets.
  • The request uses a project with a lower token limit.
  • Several models draw from one shared limit group.
  • A long-context request uses a dedicated throughput limit when one applies.
  • Pending batch jobs fill the model’s queued input-token allowance.
  • The organization has reached its approved monthly usage limit or a configured spend control.
  • The configured maximum output budget is much larger than the response the application needs.

How to Fix OpenAI 429 Errors

Read the HTTP status, error code, and response headers before retrying. A 429 slow_down error means the request rate increased too quickly. A 503 server_is_overloaded error points to temporary model overload.

  1. Honor Retry-After when the header is present.
  2. After a 429 slow_down error, reduce the request rate and ramp traffic back up gradually.
  3. Use exponential backoff with random jitter when Retry-After is missing.
  4. Reduce concurrency across web servers, workers, and background jobs.
  5. Shorten prompts and lower the maximum output token setting when TPM is the constraint.
  6. Check the organization, project, model, shared limit group, and monthly usage limit.
  7. Move deferred workloads to Batch API when Batch is available for the selected model and endpoint.

OpenAI SDK Retries

The official OpenAI Python SDK retries eligible connection, timeout, 409, 429, and server errors twice by default. Set the SDK retry count deliberately and avoid wrapping every request in another uncontrolled retry loop. Failed requests can continue to consume per-minute capacity.

from openai import OpenAI, RateLimitError
client = OpenAI(max_retries=5)
try:
    response = client.responses.create(
        model="gpt-6-luna",
        input="Summarize the incident report in five bullets.",
        max_output_tokens=400,
    )
    print(response.output_text)
except RateLimitError as error:
    # The configured SDK retries have already been attempted.
    # Record the error and move the job back to a controlled queue.
    print(f"Rate limit error: {error}")
HTTP / error codeWhat it usually meansFirst response
429 / slow_downTraffic increased too quickly or an active rate limit was exceededFollow Retry-After, reduce the rate, then ramp gradually
503 / server_is_overloadedTemporary model overloadFollow Retry-After and retry with increasing delay

How to Prevent Rate Limit Problems

Control concurrency. Set a global request ceiling across application servers and workers. Independent workers can exceed a shared organization or project pool even when each process appears safe.

Use a request queue. Smooth traffic spikes, prioritize user-facing jobs, and pace retries to prevent another traffic spike.

Set realistic output budgets. Use different defaults for short answers, extraction, summaries, reports, and long-form generation.

Reduce token waste. Remove duplicated instructions, unused examples, stale conversation history, oversized JSON, and irrelevant retrieved text.

Split live and offline workloads. Live requests should not compete with indexing, evaluations, analytics, or backfills in one uncontrolled queue.

Log rate-limit headers. Track remaining requests, remaining tokens, reset times, project token limits, model names, latency, and retry count.

Use Batch API for deferred work. Batch uses its own pool for eligible endpoints and models.

How to Increase OpenAI API Rate Limits

OpenAI moves organizations through usage tiers as paid API usage increases. Higher tiers usually raise capacity across most models.

Open the Platform Limits page when the workload exceeds capacity after queueing, backoff, token reduction, and Batch API use. Eligible organizations can request a limit increase from that page.

Organizations that routinely hit ramp-rate limits can evaluate Scale Tier on eligible models. Reserved Tier is available for GPT-5.6 and later models.

OpenAI Rate Limit Examples

Support chatbot: Thousands of short messages can exhaust RPM before TPM. Set concurrency controls, cache repeated answers, and queue background tool calls.

Report generator: A small number of long requests can exhaust TPM. Reduce retrieved context, generate sections individually, and lower the output budget.

Website embedding job: Large indexing work consumes token throughput. Deduplicate pages, queue chunks, and use Batch API when results do not need to return immediately.

Image generation app: Parallel jobs can exhaust IPM. Use per-user queues, visible wait states, and a global image-job ceiling.

Realtime voice assistant: Concurrent sessions consume session, token, and audio capacity. Reserve capacity for active calls and close idle sessions promptly.

Quick Troubleshooting Checklist

SymptomLikely causeFirst action
429 after many short callsRPM or burst limitReduce concurrency and queue requests
429 after long promptsTPM or project token limitShorten input and reduce output budget
429 during traffic spikesShort-window burst enforcementHonor Retry-After and use jitter
429 only in one projectProject-scoped token limit or spend controlCheck project settings and project headers
429 after switching modelsDifferent model table or shared limit groupCheck the selected model and shared pool
Batch job will not startQueued input-token limitWait for jobs to finish or reduce queued tokens
API stops after monthly usageOrganization usage limit or spend controlReview billing and limit settings
Repeated 429s after SDK retriesSustained overload or non-retryable quota issueRead the error code and move work to a queue

FAQs

Can a slow_down error happen below the published RPM and TPM limits?
Yes. OpenAI can return slow_down when traffic increases too quickly even before RPM or TPM is exhausted. Reduce the request rate, follow Retry-After when present, and increase traffic gradually.

Does every OpenAI 429 response include Retry-After?
No. Retry-After can appear on temporary rate-limit errors. Use increasing retry delays with a small random delay when the header is missing. Quota, billing, and account errors require the underlying issue to be corrected.

Do OpenAI SDKs retry rate-limit errors automatically?
Yes. OpenAI states that its official SDKs retry eligible rate-limit errors and honor Retry-After when present. The Python SDK retries eligible transient errors twice by default, and max_retries changes that count.

Related Resources

Changelog

09/29/2026

  • Added GPT-6.1 Sol RPM, TPM, and Batch queue limits.

09/23/2026

  • Added GPT-6 Sol and GPT-6 Luna RPM, TPM, and Batch queue limits.

09/03/2026

  • Added GPT-6 Astra RPM, TPM, and Batch queue limits.
  • Updated 429 slow_down and 503 overload troubleshooting.
  • Corrected GPT-4o API status and clarified the deprecated ChatGPT-4o alias.
  • Updated Sora API deprecation and shutdown timing.
  • Moved previous and specialized model tables into a dedicated reference section.

08/03/2026

  • Verified GPT-5.6 Sol, Terra, Luna, GPT Image 2, usage tiers, Batch API limits, and rate-limit headers.
  • Expanded model-family reference tables and 429 troubleshooting.

07/20/2026

  • Added GPT-5.6 Sol, GPT-5.6 Terra, and GPT-5.6 Luna.

One comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Get the latest & top AI tools sent directly to your email.

Subscribe now to explore the latest & top AI tools and resources, all in one convenient newsletter. No spam, we promise!