Azure Foundry subcription & latency

Rollo S 0 Reputation points
2026-09-22T00:06:57.6533333+00:00

Im using Azure Foundry with basic global subscription. And i'm noticing that my LLM calls are very slow. I dont use it much at this time, but definitely needed to boost a bit the speed to make it work for the users. What can I do to speed it up?

Foundry Tools
Foundry Tools

Formerly known as Azure AI Services or Azure Cognitive Services is a unified collection of prebuilt AI capabilities within the Microsoft Foundry platform

0 comments No comments

2 answers

Sort by: Newest
  1. Aditya Singh Rathore 195 Reputation points
    2026-09-22T11:46:51.44+00:00

    Hi @Rollo S ,

    If you're using Azure Foundry with a Standard/Global Standard deployment, I'd first check whether the latency is coming from the model itself or from the amount of input/output tokens you're sending.

    A few things that usually help:

    • Try a smaller/faster model if the workload doesn't require the larger model.
    • Reduce unnecessary context and limit the maximum output tokens.
    • Enable streaming so users start seeing the response immediately instead of waiting for the complete response.
    • Check the Foundry monitoring metrics such as Time to First Token and token processing time to see where the delay is happening.
    • If the latency is mainly due to shared capacity and you need more predictable performance, you can look at Provisioned Throughput (PTU).

    Microsoft has a good guide on this here: Azure OpenAI latency

    And for predictable throughput/latency: Provisioned Throughput

    If you're only using the service occasionally at the moment, I wouldn't move to PTU immediately. I'd first check the model, prompt size, output tokens, and streaming. Those can make a noticeable difference without changing your deployment.

    Was this answer helpful?


  2. AI answer

    2026-09-22T00:07:10.7133333+00:00

    To reduce Azure AI Foundry / Azure OpenAI latency, focus on the factors that most directly affect per-call response time:

    1. Use a lower-latency model If lowest latency is the priority, the recommended model is the latest GPT-4o mini.
    2. Reduce max_tokens A higher max_tokens value increases latency because the service reserves compute based on the requested generation size. Set max_tokens as low as possible for each call.
    3. Reduce total generated tokens Response generation is typically the slowest step. To reduce latency:
      • keep prompts shorter when possible
      • add stop sequences to prevent extra output
      • avoid multiple outputs per request by leaving n unset or setting n: 1
    4. Enable streaming Set stream: true so tokens are returned as soon as they are available. This does not reduce total completion time, but it improves time to first response and usually makes the app feel much faster to users.
    5. Separate different workloads If different request types share one endpoint, latency can increase because short requests may wait behind longer ones and cache efficiency can drop. Use separate deployments for different workloads when possible.
    6. Test batching carefully If multiple requests go to the same endpoint, batching them into a single call can reduce request overhead and may improve overall response time depending on the scenario.
    7. Review content filtering impact Content filtering improves safety but adds latency. For lower-risk workloads, evaluate whether modified content filtering policies are appropriate.
    8. Measure before changing settings Track these Azure Monitor metrics on the deployment:
      • Requests metric, split by ModelDeploymentName, to measure calls per minute
      • Processed Inference tokens metric, to measure total tokens per minute

    A practical optimization path is:

    1. Lower max_tokens.
    2. Set n: 1.
    3. Enable stream: true.
    4. Shorten prompts and add stop sequences.
    5. If latency is still too high, test a lower-latency model such as GPT-4o mini.
    6. Then review workload separation and monitoring metrics.

    AI-generated content may be incorrect. Read our transparency notes for more information.

    Was this answer helpful?

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.