An Azure service that provides access to OpenAI’s GPT-3 models with enterprise capabilities.
To check the applicable limits and diagnose the errors, use this model:
- Check the deployment rate limits from responses and quota configuration
- Azure OpenAI includes rate-limit headers on every API response:
-
x-ratelimit-limit-requests -
x-ratelimit-limit-tokens -
x-ratelimit-remaining-requests -
x-ratelimit-remaining-tokens -
x-ratelimit-reset-requests -
x-ratelimit-reset-tokens -
retry-after-mson 429 responses
-
- These headers show the effective request and token limits for the deployment and how much budget remains.
- TPM can also be adjusted through quota management in the Foundry portal for high-traffic deployments, and rebalanced across deployments based on observed usage.
- Azure OpenAI includes rate-limit headers on every API response:
- Understand how concurrent users affect RPM and TPM
- Limits apply at the deployment level, not per user.
- Every incoming request contributes to the running token count for the current minute.
- The TPM calculation is based on an estimated maximum processed-token count at request time, including:
- prompt text
-
max_tokens -
best_of
- This estimate is not the same as billed tokens, so 429s can occur even when Azure Monitor token usage appears below quota.
- RPM is enforced over short windows, typically 1 or 10 seconds. A burst of simultaneous users can trigger 429s even if the total per-minute request count stays below the published RPM.
- Distinguish rate-limit errors from capacity constraints
- Not all 429s mean the same thing.
- Rate limit exceeded: messages such as "Requests to … have been limited" or "Rate limit is exceeded" indicate the deployment exceeded TPM or RPM.
- System capacity throttling: messages such as "The service is temporarily unable to process your request" or "System is experiencing high demand" indicate backend capacity pressure.
- Temporary rate limit adjustment: if configured quota has not changed but
x-ratelimit-limit-tokensis lower than the configured TPM, the service has temporarily reduced the effective rate limit to protect shared capacity. This is typically temporary. - Also check whether token usage metrics are misleading: billed token metrics only show successfully processed requests, while rate limiting is evaluated when requests arrive and can include rejected requests.
- Recommended handling for high concurrent usage
- Implement retry logic with exponential backoff.
- Honor
retry-after-mswhen present. - Queue requests and process them at a controlled rate if immediate responses are not required.
- Ramp traffic up gradually instead of sending sharp bursts.
- Rebalance quota across deployments and increase TPM on heavily used deployments.
- Spread requests across multiple deployments or regions when one deployment cannot support the required throughput.
- For workloads that require more predictable throughput and lower latency, consider Provisioned Throughput (PTU).
- Best practices to reduce rate-limit errors
- Set
max_tokensto the minimum value needed. - Set
best_ofto1unless multiple completions are required. - Reduce prompt size where possible.
- Monitor
x-ratelimit-remaining-requestsandx-ratelimit-remaining-tokensin the application and proactively throttle before 429s occur. - Use Azure Monitor metrics to review 24-hour and 7-day usage trends and identify bursty patterns.
- Test different load-increase patterns before training sessions or peak events.
- Set
- Concurrent session note for GPT-Live
- If the assistant uses the GPT-Live API, there is also a concurrent session limit per subscription by tier. For example, the default non-tiered limit is 10 concurrent sessions.
A practical troubleshooting path for the training-session scenario is:
- Capture the HTTP response headers for successful and failed requests.
- Compare
x-ratelimit-limit-*andx-ratelimit-remaining-*values during the failures. - Check the 429 error text:
- rate-limit wording -> quota/RPM/TPM issue
- high-demand or temporarily unable to process wording -> capacity issue
- Review usage trends in Azure Monitor to see whether the session created short bursts.
- If bursts are the cause, add queuing and exponential backoff, reduce
max_tokens, and distribute traffic more evenly. - If the workload regularly exceeds a single deployment’s capacity, rebalance quota, use multiple deployments or regions, or move to PTU for more predictable capacity.