Guidance on baseline threshold for HttpQueueLength metric to identify underutilized App Service Plans for downsizing

Sree Aravind M 40 Reputation points
2026-09-23T12:22:41.5833333+00:00

We are currently identifying low-utilized Azure App Service Plans to downsize them and reduce costs. A plan currently qualifies if its CPU and Memory (P90) stay under 40% over the last 30 days. To ensure safety, we want to add Http Queue Length to our downsizing logic so we do not downsize plans experiencing request delays. Could you please advise on the recommended safe threshold and aggregation type (such as Average, P90, or Max) for Http Queue Length over a 30-day window? Additionally, are there any edge cases where a low CPU plan with queue length could cause HTTP 503 errors after downsizing?

Azure App Service
Azure App Service

Azure App Service is a service used to create and deploy scalable, mission-critical web apps.


1 answer

Sort by: Newest
  1. Vinodh247-1375 44,801 Reputation points Volunteer Moderator
    2026-09-23T15:54:54.9166667+00:00

    There is currently no Microsoft-documented universal threshold for HttpQueueLength that can be used to determine whether an App Service Plan is safe to downsize. Microsoft defines the metric as the average number of HTTP requests waiting in the queue before being processed, and a rising value is generally an indication that the plan is struggling to keep up with incoming demand. [azureglossary.com], [learn.microsoft.com]

    For downsizing decisions, I would treat HttpQueueLength as a validation signal rather than a primary sizing metric.

    A practical approach would be:

    • Use P90 rather than Average over the 30-day observation period. Average can hide short periods of saturation, while Max is often dominated by isolated spikes.
    • Look for plans where HttpQueueLength remains consistently near its established baseline and does not show recurring elevation during known peak traffic periods.
    • Correlate queue length with response time, HTTP 5xx errors, request volume, and scale events before approving a downgrade.
    • Review the distribution during peak business hours separately from the overall 30-day average, since downsizing risk is usually concentrated in peak-demand windows.

    A rule such as:

    CPU P90 < 40% AND Memory P90 < 40%

    can be strengthened by adding:

    No sustained HttpQueueLength growth during peak periods and No corresponding increase in latency or HTTP 5xx errors

    rather than relying on a fixed queue length value.

    Regarding edge cases, yes, a plan can show low CPU utilisation while still experiencing queueing and potentially generating HTTP 503 responses after downsizing. Common causes include:

    • Slow downstream dependencies such as databases or external APIs
    • Thread pool starvation or connection pool exhaustion
    • Long-running synchronous requests
    • Short traffic bursts that exceed instantaneous processing capacity
    • Applications that are latency-bound rather than CPU-bound

    In these scenarios, CPU and memory metrics alone can make a plan appear underutilised even though request processing capacity is already close to its practical limit. For that reason, HttpQueueLength is best used as a workload-specific safeguard, with thresholds derived from observed application behaviour rather than from a platform-wide Azure recommendation. [bing.com], [learn.microsoft.com], [azureglossary.com]

    Help make this community better for everyone: if this answer resolved your issue, please accept it or leave an upvote. If not, share more details in a comment so we can continue the discussion and find the right solution.

    Was this answer helpful?

    0 comments No comments

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.