Formerly known as Azure AI Services or Azure Cognitive Services is a unified collection of prebuilt AI capabilities within the Microsoft Foundry platform
To reduce Azure AI Foundry / Azure OpenAI latency, focus on the factors that most directly affect per-call response time:
- Use a lower-latency model If lowest latency is the priority, the recommended model is the latest GPT-4o mini.
- Reduce
max_tokensA highermax_tokensvalue increases latency because the service reserves compute based on the requested generation size. Setmax_tokensas low as possible for each call. - Reduce total generated tokens
Response generation is typically the slowest step. To reduce latency:
- keep prompts shorter when possible
- add stop sequences to prevent extra output
- avoid multiple outputs per request by leaving
nunset or settingn: 1
- Enable streaming
Set
stream: trueso tokens are returned as soon as they are available. This does not reduce total completion time, but it improves time to first response and usually makes the app feel much faster to users. - Separate different workloads If different request types share one endpoint, latency can increase because short requests may wait behind longer ones and cache efficiency can drop. Use separate deployments for different workloads when possible.
- Test batching carefully If multiple requests go to the same endpoint, batching them into a single call can reduce request overhead and may improve overall response time depending on the scenario.
- Review content filtering impact Content filtering improves safety but adds latency. For lower-risk workloads, evaluate whether modified content filtering policies are appropriate.
- Measure before changing settings
Track these Azure Monitor metrics on the deployment:
- Requests metric, split by ModelDeploymentName, to measure calls per minute
- Processed Inference tokens metric, to measure total tokens per minute
A practical optimization path is:
- Lower
max_tokens. - Set
n: 1. - Enable
stream: true. - Shorten prompts and add stop sequences.
- If latency is still too high, test a lower-latency model such as GPT-4o mini.
- Then review workload separation and monitoring metrics.