Understanding Azure AI Foundry Rate Limits and Concurrent User Capacity

Mohamed, Rihan 60 Reputation points
2026-09-28T15:21:26.82+00:00

I'm investigating rate limits and capacity constraints for an AI assistant built using Azure AI Foundry.

During a training session, multiple users accessed the assistant simultaneously and some users experienced errors.

I would like to understand:

How can I check the applicable Requests Per Minute (RPM) and Tokens Per Minute (TPM) limits?

How does concurrent user activity affect these limits?

How can I identify whether errors are caused by rate limits or capacity constraints?

What are the recommended approaches to handle high concurrent usage, such as retries, request queuing, or increasing capacity?

Are there any best practices for preventing rate-limit errors during peak usage?

Any guidance or documentation would be appreciated.

Azure OpenAI in Foundry Models

1 answer

Sort by: Most helpful
  1. AI answer

    2026-09-28T15:21:46.6366667+00:00

    To check the applicable limits and diagnose the errors, use this model:

    1. Check the deployment rate limits from responses and quota configuration
      • Azure OpenAI includes rate-limit headers on every API response:
        • x-ratelimit-limit-requests
        • x-ratelimit-limit-tokens
        • x-ratelimit-remaining-requests
        • x-ratelimit-remaining-tokens
        • x-ratelimit-reset-requests
        • x-ratelimit-reset-tokens
        • retry-after-ms on 429 responses
      • These headers show the effective request and token limits for the deployment and how much budget remains.
      • TPM can also be adjusted through quota management in the Foundry portal for high-traffic deployments, and rebalanced across deployments based on observed usage.
    2. Understand how concurrent users affect RPM and TPM
      • Limits apply at the deployment level, not per user.
      • Every incoming request contributes to the running token count for the current minute.
      • The TPM calculation is based on an estimated maximum processed-token count at request time, including:
        1. prompt text
        2. max_tokens
        3. best_of
      • This estimate is not the same as billed tokens, so 429s can occur even when Azure Monitor token usage appears below quota.
      • RPM is enforced over short windows, typically 1 or 10 seconds. A burst of simultaneous users can trigger 429s even if the total per-minute request count stays below the published RPM.
    3. Distinguish rate-limit errors from capacity constraints
      • Not all 429s mean the same thing.
      • Rate limit exceeded: messages such as "Requests to … have been limited" or "Rate limit is exceeded" indicate the deployment exceeded TPM or RPM.
      • System capacity throttling: messages such as "The service is temporarily unable to process your request" or "System is experiencing high demand" indicate backend capacity pressure.
      • Temporary rate limit adjustment: if configured quota has not changed but x-ratelimit-limit-tokens is lower than the configured TPM, the service has temporarily reduced the effective rate limit to protect shared capacity. This is typically temporary.
      • Also check whether token usage metrics are misleading: billed token metrics only show successfully processed requests, while rate limiting is evaluated when requests arrive and can include rejected requests.
    4. Recommended handling for high concurrent usage
      • Implement retry logic with exponential backoff.
      • Honor retry-after-ms when present.
      • Queue requests and process them at a controlled rate if immediate responses are not required.
      • Ramp traffic up gradually instead of sending sharp bursts.
      • Rebalance quota across deployments and increase TPM on heavily used deployments.
      • Spread requests across multiple deployments or regions when one deployment cannot support the required throughput.
      • For workloads that require more predictable throughput and lower latency, consider Provisioned Throughput (PTU).
    5. Best practices to reduce rate-limit errors
      • Set max_tokens to the minimum value needed.
      • Set best_of to 1 unless multiple completions are required.
      • Reduce prompt size where possible.
      • Monitor x-ratelimit-remaining-requests and x-ratelimit-remaining-tokens in the application and proactively throttle before 429s occur.
      • Use Azure Monitor metrics to review 24-hour and 7-day usage trends and identify bursty patterns.
      • Test different load-increase patterns before training sessions or peak events.
    6. Concurrent session note for GPT-Live
      • If the assistant uses the GPT-Live API, there is also a concurrent session limit per subscription by tier. For example, the default non-tiered limit is 10 concurrent sessions.

    A practical troubleshooting path for the training-session scenario is:

    1. Capture the HTTP response headers for successful and failed requests.
    2. Compare x-ratelimit-limit-* and x-ratelimit-remaining-* values during the failures.
    3. Check the 429 error text:
      • rate-limit wording -> quota/RPM/TPM issue
      • high-demand or temporarily unable to process wording -> capacity issue
    4. Review usage trends in Azure Monitor to see whether the session created short bursts.
    5. If bursts are the cause, add queuing and exponential backoff, reduce max_tokens, and distribute traffic more evenly.
    6. If the workload regularly exceeds a single deployment’s capacity, rebalance quota, use multiple deployments or regions, or move to PTU for more predictable capacity.

    AI-generated content may be incorrect. Read our transparency notes for more information.

    Was this answer helpful?

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.