Formerly known as Azure AI Services or Azure Cognitive Services is a unified collection of prebuilt AI capabilities within the Microsoft Foundry platform
Hi @Rollo S ,
If you're using Azure Foundry with a Standard/Global Standard deployment, I'd first check whether the latency is coming from the model itself or from the amount of input/output tokens you're sending.
A few things that usually help:
- Try a smaller/faster model if the workload doesn't require the larger model.
- Reduce unnecessary context and limit the maximum output tokens.
- Enable streaming so users start seeing the response immediately instead of waiting for the complete response.
- Check the Foundry monitoring metrics such as Time to First Token and token processing time to see where the delay is happening.
- If the latency is mainly due to shared capacity and you need more predictable performance, you can look at Provisioned Throughput (PTU).
Microsoft has a good guide on this here: Azure OpenAI latency
And for predictable throughput/latency: Provisioned Throughput
If you're only using the service occasionally at the moment, I wouldn't move to PTU immediately. I'd first check the model, prompt size, output tokens, and streaming. Those can make a noticeable difference without changing your deployment.