OpenAI API rate limits control how many requests, tokens, images, audio minutes, and batch jobs an organization or project can process within a fixed period.
The active limits depend on the model, endpoint, usage tier, organization, project, request size, and shared limit group. They are separate from model context windows, maximum output tokens, ChatGPT message caps, and monthly spend controls.
This post explains RPM, TPM, RPD, TPD, and IPM; lists verified limits for current and previous GPT, reasoning, image, video, realtime, audio, embedding, and moderation models; shows where to check account-specific limits; and covers practical fixes for 429 errors.
Last verified: September 29, 2026.
Understanding the OpenAI API Rate Limits
OpenAI measures API throughput through several independent limits. A request fails as soon as it exceeds any applicable limit.
API rate limits apply to API organizations and projects. ChatGPT plan caps, context windows, and maximum output tokens use different limits. See ChatGPT and OpenAI token limits for those values.
| Limit | Meaning | What it controls | Common trigger |
|---|---|---|---|
| RPM | Requests per minute | API calls sent within a minute | Many short requests or high concurrency |
| RPD | Requests per day | Total daily API calls | Large daily job volume |
| TPM | Tokens per minute | Input and output token throughput | Long prompts, large outputs, or many parallel calls |
| TPD | Tokens per day | Total daily token throughput | High-volume production workloads |
| IPM | Images per minute | Image generations and edits | Parallel image jobs |
| Audio minutes per minute | Audio throughput | Audio processed by eligible streaming models | Concurrent voice sessions |
| Batch queue limit | Queued input tokens per model | Prompt tokens waiting in Batch API jobs | Large pending batch workloads |
An application can stay below TPM and hit RPM first. A long prompt can stay below RPM and hit TPM first. Batch jobs can stop accepting work after queued input tokens reach the model-specific queue limit.
Current OpenAI API Rate Limits: Quick Reference
OpenAI can assign different limits to an organization, project, long-context workload, or shared model group. Check the Platform Limits dashboard before setting production capacity or concurrency targets.
GPT-6 Astra Rate Limits
| Usage tier | RPM | TPM | Batch queue limit |
|---|---|---|---|
| Free | Not supported | Not supported | Not supported |
| Tier 1 | 500 | 500,000 | 1,500,000 |
| Tier 2 | 5,000 | 1,000,000 | 3,000,000 |
| Tier 3 | 5,000 | 2,000,000 | 100,000,000 |
| Tier 4 | 10,000 | 4,000,000 | 200,000,000 |
| Tier 5 | 15,000 | 40,000,000 | 15,000,000,000 |
GPT-6.1 Sol Rate Limits
| Usage tier | RPM | TPM | Batch queue limit |
|---|---|---|---|
| Free | Not supported | Not supported | Not supported |
| Tier 1 | 500 | 500,000 | 1,500,000 |
| Tier 2 | 5,000 | 1,000,000 | 3,000,000 |
| Tier 3 | 5,000 | 2,000,000 | 100,000,000 |
| Tier 4 | 10,000 | 4,000,000 | 200,000,000 |
| Tier 5 | 15,000 | 40,000,000 | 15,000,000,000 |
GPT-6 Luna Rate Limits
| Usage tier | RPM | TPM | Batch queue limit |
|---|---|---|---|
| Free | Not supported | Not supported | Not supported |
| Tier 1 | 500 | 500,000 | 5,000,000 |
| Tier 2 | 5,000 | 2,000,000 | 20,000,000 |
| Tier 3 | 5,000 | 4,000,000 | 40,000,000 |
| Tier 4 | 10,000 | 10,000,000 | 1,000,000,000 |
| Tier 5 | 30,000 | 180,000,000 | 15,000,000,000 |
GPT-5.6 Sol, Terra, and Luna Synchronous Limits
| Usage tier | GPT-5.6 Sol RPM / TPM | GPT-5.6 Terra RPM / TPM | GPT-5.6 Luna RPM / TPM |
|---|---|---|---|
| Free | Not supported | Not supported | Not supported |
| Tier 1 | 500 / 500,000 | 500 / 500,000 | 500 / 500,000 |
| Tier 2 | 5,000 / 1,000,000 | 5,000 / 1,000,000 | 5,000 / 2,000,000 |
| Tier 3 | 5,000 / 2,000,000 | 5,000 / 2,000,000 | 5,000 / 4,000,000 |
| Tier 4 | 10,000 / 4,000,000 | 10,000 / 4,000,000 | 10,000 / 10,000,000 |
| Tier 5 | 15,000 / 40,000,000 | 15,000 / 40,000,000 | 30,000 / 180,000,000 |
GPT-5.6 Batch Queue Limits
Batch queue limits count input tokens waiting in pending Batch jobs. Tokens leave the queue count after the related batch completes.
| Usage tier | GPT-5.6 Sol | GPT-5.6 Terra | GPT-5.6 Luna |
|---|---|---|---|
| Free | Not supported | Not supported | Not supported |
| Tier 1 | 1,500,000 | 1,500,000 | 5,000,000 |
| Tier 2 | 3,000,000 | 3,000,000 | 20,000,000 |
| Tier 3 | 100,000,000 | 100,000,000 | 40,000,000 |
| Tier 4 | 200,000,000 | 200,000,000 | 1,000,000,000 |
| Tier 5 | 15,000,000,000 | 15,000,000,000 | 15,000,000,000 |
GPT Image 2.5 Rate Limits
| Usage tier | TPM | IPM |
|---|---|---|
| Free | Not supported | Not supported |
| Tier 1 | 100,000 | 5 |
| Tier 2 | 250,000 | 20 |
| Tier 3 | 800,000 | 50 |
| Tier 4 | 3,000,000 | 150 |
| Tier 5 | 8,000,000 | 250 |
OpenAI Usage Tiers
OpenAI moves organizations to higher usage tiers as paid API spend increases. Higher tiers usually raise rate limits across most models. Each model keeps its own rate-limit profile.
| Tier | Qualification | Monthly usage limit |
|---|---|---|
| Free | User must be in an allowed geography | $100 per month |
| Tier 1 | $5 paid | $100 per month |
| Tier 2 | $50 paid | $500 per month |
| Tier 3 | $100 paid | $1,000 per month |
| Tier 4 | $250 paid | $5,000 per month |
| Tier 5 | $1,000 paid | $200,000 per month |
OpenAI API Free Tier Rate Limits
The API Free tier applies to eligible users in allowed geographies. Model access and exact limits depend on the API account and model.
GPT-6 Astra, GPT-6.1 Sol, GPT-6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna, GPT-4.1, GPT-4o, GPT Image 2.5, and current Realtime models list Free-tier access as Not supported in their published model tables.
The Platform Limits page shows the Free-tier models available to your account. The following specialized models have published Free-tier limits:
| Model or category | Free-tier limits |
|---|---|
| text-embedding-3-large and text-embedding-3-small | 100 RPM, 2,000 RPD, 40,000 TPM |
| omni-moderation-latest | 250 RPM, 5,000 RPD, 10,000 TPM |
| TTS-1 | 3 RPM, 200 RPD |
| Whisper-1 | 3 RPM, 200 RPD |
OpenAI API Limits vs ChatGPT Free and Plus Limits
OpenAI API limits and ChatGPT plan limits use separate products, billing systems, and measurements. A ChatGPT subscription does not include API usage or increase an API organization’s rate limits.
| Limit type | Applies to | Typical measurements | Where to check |
|---|---|---|---|
| OpenAI API rate limits | API organizations and projects | RPM, TPM, RPD, TPD, IPM, batch queue | OpenAI Platform Limits page |
| API free usage tier | Eligible API accounts and supported models | Model-specific API throughput and monthly usage limit | OpenAI Platform Limits page |
| ChatGPT Free, Go, Plus, and Pro limits | ChatGPT plans | Messages, model access, tools, and feature usage | ChatGPT interface and plan documentation |
For context windows, maximum output tokens, and ChatGPT plan limits, see ChatGPT and OpenAI token limits.
Where to Check Your Current OpenAI Rate Limits
Use the OpenAI Platform dashboard for account-specific values:
- Open OpenAI Platform Limits.
- Select the correct organization.
- Select the project used by the application.
- Review limits by model and endpoint.
- Check shared model groups, long-context limits, and project-level token limits.
- Use the increase request option when it appears and the workload requires more capacity.
Rate-Limit Response Headers
Production applications should log rate-limit response headers. These headers show the active request and token budget for a call.
| Header | What it shows |
|---|---|
Retry-After | Minimum wait in seconds before retrying an eligible temporary 429 error |
x-ratelimit-limit-requests | Maximum request allowance in the active window |
x-ratelimit-remaining-requests | Requests left in the active window |
x-ratelimit-reset-requests | Time until the request allowance resets |
x-ratelimit-limit-tokens | Maximum token allowance in the active window |
x-ratelimit-remaining-tokens | Tokens left in the active window |
x-ratelimit-reset-tokens | Time until the token allowance resets |
x-ratelimit-limit-project-tokens | Project-scoped token limit when one applies |
x-ratelimit-remaining-project-tokens | Project-scoped tokens left |
x-ratelimit-reset-project-tokens | Time until the project token allowance resets |
How OpenAI API Rate Limits Work
- Organization and project scope: Limits are defined at the organization and project level.
- Model-specific capacity: Each model has its own tier table.
- Shared model groups: Related models can draw from one RPM or TPM pool.
- Long-context limits: Large-context requests can have a lower throughput limit.
- Project token limits: A project-level token limit can be lower than the organization-wide model limit.
- Monthly usage limits: OpenAI assigns an approved monthly organization usage limit. Organization and project spend controls use different settings.
- Vector store ingestion: File and file-batch ingestion endpoints share a 300 RPM limit for each vector store ID.
Previous and Specialized OpenAI Model Rate Limits
Earlier GPT-5 Model Rate Limits
| Model or model family | Tier 1 | Tier 2 | Tier 3 | Tier 4 | Tier 5 |
|---|---|---|---|---|---|
| GPT-5.5, GPT-5.4, GPT-5.2, GPT-5, GPT-5.3 Codex | 500 / 500,000 | 5,000 / 1,000,000 | 5,000 / 2,000,000 | 10,000 / 4,000,000 | 15,000 / 40,000,000 |
| GPT-5.5 Pro | 50 / 50,000 | 500 / 200,000 | 500 / 500,000 | 1,000 / 1,000,000 | 2,000 / 4,000,000 |
| GPT-5.4 Pro, GPT-5.2 Pro, GPT-5 Pro | 500 / 30,000 | 5,000 / 450,000 | 5,000 / 800,000 | 10,000 / 2,000,000 | 10,000 / 30,000,000 |
| GPT-5.4 Mini, GPT-5 Mini | 500 / 500,000 | 5,000 / 2,000,000 | 5,000 / 4,000,000 | 10,000 / 10,000,000 | 30,000 / 180,000,000 |
| GPT-5.4 Nano, GPT-5 Nano | 500 / 200,000 | 5,000 / 2,000,000 | 5,000 / 4,000,000 | 10,000 / 10,000,000 | 30,000 / 180,000,000 |
GPT-4.1, GPT-4o, and GPT-4 Rate Limits
| Model | Status | Tier 1 | Tier 2 | Tier 3 | Tier 4 | Tier 5 |
|---|---|---|---|---|---|---|
| GPT-4.1 | Current | 500 / 30,000 | 5,000 / 450,000 | 5,000 / 800,000 | 10,000 / 2,000,000 | 10,000 / 30,000,000 |
| GPT-4.1 mini | Current | 500 / 200,000 | 5,000 / 2,000,000 | 5,000 / 4,000,000 | 10,000 / 10,000,000 | 30,000 / 150,000,000 |
| GPT-4.1 nano | Deprecated | 500 / 200,000 | 5,000 / 2,000,000 | 5,000 / 4,000,000 | 10,000 / 10,000,000 | 30,000 / 150,000,000 |
| GPT-4o | Current | 500 / 30,000 | 5,000 / 450,000 | 5,000 / 800,000 | 10,000 / 2,000,000 | 10,000 / 30,000,000 |
| GPT-4o mini | Current | 500 / 200,000 | 5,000 / 2,000,000 | 5,000 / 4,000,000 | 10,000 / 10,000,000 | 30,000 / 150,000,000 |
| GPT-4 | Legacy | 500 / 10,000 | 5,000 / 40,000 | 5,000 / 80,000 | 10,000 / 300,000 | 10,000 / 1,000,000 |
Earlier O-Series Reasoning Model Rate Limits
| Model or model family | Tier 1 | Tier 2 | Tier 3 | Tier 4 | Tier 5 |
|---|---|---|---|---|---|
| o3, o3-pro, o1, o1-pro | 500 / 30,000 | 5,000 / 450,000 | 5,000 / 800,000 | 10,000 / 2,000,000 | 10,000 / 30,000,000 |
| o4-mini | 1,000 / 100,000 | 2,000 / 2,000,000 | 5,000 / 4,000,000 | 10,000 / 10,000,000 | 30,000 / 150,000,000 |
| o3-mini | 1,000 / 100,000 | 2,000 / 200,000 | 5,000 / 4,000,000 | 10,000 / 10,000,000 | 30,000 / 150,000,000 |
| o1-mini | 500 / 200,000 | 5,000 / 2,000,000 | 5,000 / 4,000,000 | 10,000 / 10,000,000 | 30,000 / 150,000,000 |
| o3-deep-research | 500 / 200,000 | 5,000 / 450,000 | 5,000 / 800,000 | 10,000 / 2,000,000 | 10,000 / 30,000,000 |
| o4-mini-deep-research | 1,000 / 200,000 | 2,000 / 2,000,000 | 5,000 / 4,000,000 | 10,000 / 10,000,000 | 30,000 / 150,000,000 |
Deprecated Image Model Rate Limits
| Model | Status | Tier 1 | Tier 2 | Tier 3 | Tier 4 | Tier 5 |
|---|---|---|---|---|---|---|
| GPT Image 1.5 | Deprecated | 100,000 TPM / 5 IPM | 250,000 / 20 | 800,000 / 50 | 3,000,000 / 150 | 8,000,000 / 250 |
| GPT Image 1 | Deprecated | 100,000 TPM / 5 IPM | 250,000 / 20 | 800,000 / 50 | 3,000,000 / 150 | 8,000,000 / 250 |
| GPT Image 1 mini | Deprecated | 100,000 TPM / 5 IPM | 250,000 / 20 | 800,000 / 50 | 3,000,000 / 150 | 8,000,000 / 250 |
Deprecated Sora 2 Video API Rate Limits
The Sora API is deprecated and scheduled to shut down on September 24, 2026. These RPM values apply only to existing Sora 2 and Sora 2 Pro API integrations before shutdown.
| Model | Free | Tier 1 | Tier 2 | Tier 3 | Tier 4 | Tier 5 |
|---|---|---|---|---|---|---|
| Sora 2 | Not supported | 25 RPM | 50 RPM | 125 RPM | 200 RPM | 375 RPM |
| Sora 2 Pro | Not supported | 10 RPM | 25 RPM | 50 RPM | 75 RPM | 150 RPM |
Realtime API Rate Limits
| Model group | Status | Tier 1 | Tier 2 | Tier 3 | Tier 4 | Tier 5 |
|---|---|---|---|---|---|---|
| GPT-Realtime-2.1, 2.1 mini, 2, 1.5 | Current | 200 / 40,000 | 400 / 200,000 | 5,000 / 800,000 | 10,000 / 4,000,000 | 20,000 / 15,000,000 |
| GPT-Realtime, GPT-4o Realtime, GPT-4o mini Realtime | Deprecated | 200 / 40,000 | 400 / 200,000 | 5,000 / 800,000 | 10,000 / 4,000,000 | 20,000 / 15,000,000 |
Audio, Speech, and Transcription Rate Limits
| Model | Free | Tier 1 | Tier 2 | Tier 3 | Tier 4 | Tier 5 |
|---|---|---|---|---|---|---|
| GPT-Transcribe | Not supported | 500 / 200,000 | 5,000 / 2,000,000 | 5,000 / 4,000,000 | 10,000 / 10,000,000 | 30,000 / 150,000,000 |
| gpt-audio-1.5 | Not supported | 500 / 30,000 | 5,000 / 450,000 | 5,000 / 800,000 | 10,000 / 2,000,000 | 10,000 / 30,000,000 |
| GPT-4o Transcribe | Not supported | 500 / 10,000 | 2,000 / 100,000 | 5,000 / 400,000 | 10,000 / 2,000,000 | 10,000 / 6,000,000 |
| GPT-4o Transcribe Diarize | Not supported | 500 / 10,000 | 5,000 / 100,000 | 5,000 / 400,000 | 10,000 / 2,000,000 | 10,000 / 6,000,000 |
| GPT-4o mini Transcribe | Not supported | 500 / 50,000 | 2,000 / 150,000 | 5,000 / 600,000 | 10,000 / 2,000,000 | 10,000 / 8,000,000 |
| Classic audio model | Free | Tier 1 | Tier 2 | Tier 3 | Tier 4 | Tier 5 |
|---|---|---|---|---|---|---|
| TTS-1 | 3 RPM / 200 RPD | 500 RPM | 2,500 RPM | 5,000 RPM | 7,500 RPM | 10,000 RPM |
| TTS-1 HD | Not supported | 500 RPM | 2,500 RPM | 5,000 RPM | 7,500 RPM | 10,000 RPM |
| Whisper-1 | 3 RPM / 200 RPD | 500 RPM | 2,500 RPM | 5,000 RPM | 7,500 RPM | 10,000 RPM |
Embedding API Rate Limits
| Tier | RPM | RPD | TPM | Batch Queue Limit |
|---|---|---|---|---|
| Free | 100 | 2,000 | 40,000 | – |
| Tier 1 | 3,000 | – | 1,000,000 | 3,000,000 |
| Tier 2 | 5,000 | – | 1,000,000 | 20,000,000 |
| Tier 3 | 5,000 | – | 5,000,000 | 100,000,000 |
| Tier 4 | 10,000 | – | 5,000,000 | 500,000,000 |
| Tier 5 | 10,000 | – | 10,000,000 | 4,000,000,000 |
Moderation API Rate Limits
| Tier | RPM | RPD | TPM |
|---|---|---|---|
| Free | 250 | 5,000 | 10,000 |
| Tier 1 | 500 | 10,000 | 10,000 |
| Tier 2 | 500 | – | 20,000 |
| Tier 3 | 1,000 | – | 50,000 |
| Tier 4 | 2,000 | – | 250,000 |
| Tier 5 | 5,000 | – | 500,000 |
Rate Limits by API Type
| API type | Main limits to monitor | Typical constraint |
|---|---|---|
| Responses API | RPM, TPM, TPD, shared limits, project token limits | Agents and multimodal workflows consume token capacity through prompts, outputs, and tool calls |
| Chat Completions API | RPM, TPM, TPD | High-concurrency chat apps often hit RPM first |
| Embeddings API | RPM, TPM, payload size | Large indexing jobs consume token throughput |
| Images API | TPM, IPM | Parallel generation and editing jobs hit IPM |
| Realtime and audio APIs | Sessions, RPM, TPM, audio throughput, model-specific limits | Voice products require concurrency and session control |
| Batch API | Requests per batch, file size, batch creation rate, queued input tokens | Large offline workloads need queue planning |
| Video API (deprecated) | Model-specific RPM | Sora API access is scheduled to end on September 24, 2026 |
| Moderation API | RPM, RPD, TPM | Text and image moderation uses its own model profile |
| Fine-tuning | Active jobs, queued jobs, jobs per day, inference limits | Training capacity and inference capacity use different limits |
| Vector stores | Ingestion requests per vector store | File and file-batch ingestion share 300 RPM per vector store ID |
Batch API Rate Limits
The Batch API processes asynchronous workloads through its own rate-limit pool. It works for evaluations, classification, embeddings, backfills, extraction, moderation, image jobs, and other jobs that can wait for completion.
| Batch limit | Current value or behavior |
|---|---|
| Requests per batch | Up to 50,000 requests |
| Input file size | Up to 200 MB |
| Embedding inputs | Up to 50,000 embedding inputs across all requests in one batch |
| Batch creation rate | Up to 2,000 batches per hour |
| Queued prompt tokens | Model-specific limit shown in Platform settings |
| Completion window | Designed to complete within 24 hours |
| Output-token limit | No Batch-specific output-token limit |
Why You Can Get 429 Errors Below the Published Limit
OpenAI calculates token capacity from the larger of the configured maximum token value and an estimate based on request size. Set the output budget close to the expected response length.
- A short traffic burst exceeds a smaller internal time window.
- RPM reaches its limit before TPM.
- TPM is exhausted through long prompts or large output budgets.
- The request uses a project with a lower token limit.
- Several models draw from one shared limit group.
- A long-context request uses a dedicated throughput limit when one applies.
- Pending batch jobs fill the model’s queued input-token allowance.
- The organization has reached its approved monthly usage limit or a configured spend control.
- The configured maximum output budget is much larger than the response the application needs.
How to Fix OpenAI 429 Errors
Read the HTTP status, error code, and response headers before retrying. A 429 slow_down error means the request rate increased too quickly. A 503 server_is_overloaded error points to temporary model overload.
- Honor
Retry-Afterwhen the header is present. - After a 429
slow_downerror, reduce the request rate and ramp traffic back up gradually. - Use exponential backoff with random jitter when
Retry-Afteris missing. - Reduce concurrency across web servers, workers, and background jobs.
- Shorten prompts and lower the maximum output token setting when TPM is the constraint.
- Check the organization, project, model, shared limit group, and monthly usage limit.
- Move deferred workloads to Batch API when Batch is available for the selected model and endpoint.
OpenAI SDK Retries
The official OpenAI Python SDK retries eligible connection, timeout, 409, 429, and server errors twice by default. Set the SDK retry count deliberately and avoid wrapping every request in another uncontrolled retry loop. Failed requests can continue to consume per-minute capacity.
from openai import OpenAI, RateLimitError
client = OpenAI(max_retries=5)
try:
response = client.responses.create(
model="gpt-6-luna",
input="Summarize the incident report in five bullets.",
max_output_tokens=400,
)
print(response.output_text)
except RateLimitError as error:
# The configured SDK retries have already been attempted.
# Record the error and move the job back to a controlled queue.
print(f"Rate limit error: {error}")| HTTP / error code | What it usually means | First response |
|---|---|---|
429 / slow_down | Traffic increased too quickly or an active rate limit was exceeded | Follow Retry-After, reduce the rate, then ramp gradually |
503 / server_is_overloaded | Temporary model overload | Follow Retry-After and retry with increasing delay |
How to Prevent Rate Limit Problems
Control concurrency. Set a global request ceiling across application servers and workers. Independent workers can exceed a shared organization or project pool even when each process appears safe.
Use a request queue. Smooth traffic spikes, prioritize user-facing jobs, and pace retries to prevent another traffic spike.
Set realistic output budgets. Use different defaults for short answers, extraction, summaries, reports, and long-form generation.
Reduce token waste. Remove duplicated instructions, unused examples, stale conversation history, oversized JSON, and irrelevant retrieved text.
Split live and offline workloads. Live requests should not compete with indexing, evaluations, analytics, or backfills in one uncontrolled queue.
Log rate-limit headers. Track remaining requests, remaining tokens, reset times, project token limits, model names, latency, and retry count.
Use Batch API for deferred work. Batch uses its own pool for eligible endpoints and models.
How to Increase OpenAI API Rate Limits
OpenAI moves organizations through usage tiers as paid API usage increases. Higher tiers usually raise capacity across most models.
Open the Platform Limits page when the workload exceeds capacity after queueing, backoff, token reduction, and Batch API use. Eligible organizations can request a limit increase from that page.
Organizations that routinely hit ramp-rate limits can evaluate Scale Tier on eligible models. Reserved Tier is available for GPT-5.6 and later models.
OpenAI Rate Limit Examples
Support chatbot: Thousands of short messages can exhaust RPM before TPM. Set concurrency controls, cache repeated answers, and queue background tool calls.
Report generator: A small number of long requests can exhaust TPM. Reduce retrieved context, generate sections individually, and lower the output budget.
Website embedding job: Large indexing work consumes token throughput. Deduplicate pages, queue chunks, and use Batch API when results do not need to return immediately.
Image generation app: Parallel jobs can exhaust IPM. Use per-user queues, visible wait states, and a global image-job ceiling.
Realtime voice assistant: Concurrent sessions consume session, token, and audio capacity. Reserve capacity for active calls and close idle sessions promptly.
Quick Troubleshooting Checklist
| Symptom | Likely cause | First action |
|---|---|---|
| 429 after many short calls | RPM or burst limit | Reduce concurrency and queue requests |
| 429 after long prompts | TPM or project token limit | Shorten input and reduce output budget |
| 429 during traffic spikes | Short-window burst enforcement | Honor Retry-After and use jitter |
| 429 only in one project | Project-scoped token limit or spend control | Check project settings and project headers |
| 429 after switching models | Different model table or shared limit group | Check the selected model and shared pool |
| Batch job will not start | Queued input-token limit | Wait for jobs to finish or reduce queued tokens |
| API stops after monthly usage | Organization usage limit or spend control | Review billing and limit settings |
| Repeated 429s after SDK retries | Sustained overload or non-retryable quota issue | Read the error code and move work to a queue |
FAQs
Can a slow_down error happen below the published RPM and TPM limits?
Yes. OpenAI can return slow_down when traffic increases too quickly even before RPM or TPM is exhausted. Reduce the request rate, follow Retry-After when present, and increase traffic gradually.
Does every OpenAI 429 response include Retry-After?
No. Retry-After can appear on temporary rate-limit errors. Use increasing retry delays with a small random delay when the header is missing. Quota, billing, and account errors require the underlying issue to be corrected.
Do OpenAI SDKs retry rate-limit errors automatically?
Yes. OpenAI states that its official SDKs retry eligible rate-limit errors and honor Retry-After when present. The Python SDK retries eligible transient errors twice by default, and max_retries changes that count.
Related Resources
- OpenAI API Rate Limits Guide for rate-limit behavior, headers, usage tiers, and retry guidance.
- OpenAI Model Catalog and Model Status for model availability and published per-model limits.
- OpenAI Platform Limits Dashboard for organization- and project-specific limits.
- OpenAI Batch API Guide for batch file, request, queue, and completion constraints.
- OpenAI Python SDK Reference for client configuration and retry settings.
- ChatGPT Token Limit: Free, Plus, Pro, and OpenAI API Limits for context windows, maximum output tokens, and ChatGPT plan limits.
- AI LLM API Pricing: GPT, Gemini, Claude, and More for current token pricing across major model providers.
- OpenAI API Cost Calculator for estimating request costs from token usage.
- Gemini API Free Tier Limits: Requests Per Day & Setup.
- Claude Code Pricing and Usage Limits: Free, Pro, Max, Team, and API.
- AI Coding Plan Comparison: Prices, Limits, and Features
Changelog
09/29/2026
- Added GPT-6.1 Sol RPM, TPM, and Batch queue limits.
09/23/2026
- Added GPT-6 Sol and GPT-6 Luna RPM, TPM, and Batch queue limits.
09/03/2026
- Added GPT-6 Astra RPM, TPM, and Batch queue limits.
- Updated 429
slow_downand 503 overload troubleshooting. - Corrected GPT-4o API status and clarified the deprecated ChatGPT-4o alias.
- Updated Sora API deprecation and shutdown timing.
- Moved previous and specialized model tables into a dedicated reference section.
08/03/2026
- Verified GPT-5.6 Sol, Terra, Luna, GPT Image 2, usage tiers, Batch API limits, and rate-limit headers.
- Expanded model-family reference tables and 429 troubleshooting.
07/20/2026
- Added GPT-5.6 Sol, GPT-5.6 Terra, and GPT-5.6 Luna.











This is a very good article.
Thank you all for providing great help to our users.