Container Apps: Persistent Tainting of Dedicated Nodes and Autoscaler 'Maximum Allowed Cores Exceeded' Error

David Hish 0 Reputation points
2026-09-23T15:16:32.63+00:00

Problem description

I am experiencing an issue where dedicated workload profile nodes in my Azure Container Apps environment remain tainted and unschedulable. Despite attempts to recreate the workload profile and adjust configuration, the autoscaler continues to report 'Maximum Allowed Cores exceeded' at 0/200 cores. This prevents new pods from scheduling and hinders scaling.

Environment

Azure Container Apps, Scaling, Microsoft.App/managedEnvironments, region not specified in case information

What I've already tried

I recreated the workload profile as a new pool with different VM families, increased the maximum nodes for the profile, reduced replica requests, and restarted pending executions. I also reviewed activity logs and environment usage data, but the issue persists.

Current status

I am seeking assistance to identify the cause of the tainted nodes and the autoscaler 'Maximum Allowed Cores exceeded' error, and to resolve the scheduling and scaling issues in my environment.

Azure Container Apps
Azure Container Apps

An Azure service that provides a general-purpose, serverless container platform.

0 comments No comments

2 answers

Sort by: Oldest
  1. Tejaswini Billakurthi 760 Reputation points Microsoft External Staff Moderator
    2026-09-23T16:50:04.06+00:00

    Hello @David Hish

    Thank you for reaching out to Microsoft Q & A !

    The “Maximum Allowed Cores exceeded” message should be investigated together with the environment-level usage information. Since you are seeing 0/200 cores, I would not conclude from the error message alone that the environment has consumed the 200-core limit.

    Since you have already recreated the Dedicated workload profile with different VM families, increased the maximum node count, and reduced the replica requirements, I recommend isolating the issue as follows:

    1. Verify quota and capacity separately

    Please verify:

    • The applicable Azure Container Apps/environment quota
    • The relevant Azure subscription quota in the region
    • Capacity availability for the selected VM family in that region

    Quota availability and regional VM capacity should be treated as separate checks.

    2. Capture the exact node taint

    Since the Dedicated nodes remain tainted and workloads are not scheduling, please capture the complete taint information:

    • Taint key
    • Taint value
    • Taint effect, for example NoSchedule or NoExecute, if present
    • Any provisioning or scheduling event associated with the affected node

    The presence of a taint by itself does not establish the underlying cause. The exact taint and associated events are important for determining why the node remains unschedulable.

    3. Review the Dedicated workload profile

    Verify the effective configuration of the affected Dedicated workload profile, including its VM family and minimum/maximum instance configuration, and review any provisioning or scaling errors generated when another workload-profile instance is requested.

    4. Test with a minimal workload

    After confirming the environment and workload-profile configuration, deploy a minimal workload to the affected Dedicated profile.

    If the minimal workload also remains unschedulable while the applicable quota still shows available capacity, this helps isolate the issue from the CPU/memory requirements or scaling configuration of the original application.

    5. Escalate if the minimal scenario still reproduces

    If the issue persists after recreating the Dedicated profile and can also be reproduced with a minimal workload, I recommend opening an Azure Support request so the managed environment/workload-profile provisioning can be investigated.

    Please include:

    • Container Apps environment resource ID
    • Region
    • Dedicated workload-profile name and VM family
    • Minimum and maximum instance configuration
    • Environment/workload-profile usage and quota information
    • Exact node taint key, value, and effect
    • Relevant provisioning/scheduling events and system logs
    • UTC timestamp of the failed scaling attempt
    • Complete “Maximum Allowed Cores exceeded” error

    Please remove credentials, secrets, tokens, or other sensitive information before sharing diagnostic output publicly.

    At this stage, I would avoid treating either the persistent taint or the reported 0/200 cores value alone as the confirmed root cause. Correlating the exact taint and provisioning events with the effective quota and workload-profile configuration should help narrow down the issue.

    References:

     

    Please "Accept the Answer" if this information helped you. This will help us and others in the community as well.

    Was this answer helpful?

    0 comments No comments

  2. David Hish 0 Reputation points
    2026-10-01T09:47:27.2633333+00:00

    Microsoft support identified this as a platform-side issue, not a configuration or quota problem.

    Symptoms

    • Jobs on a Dedicated workload profile stayed Pending.
    • New dedicated nodes joined but never became schedulable. System logs showed AssigningReplicaFailed with "node(s) had untolerated taint(s)", then NotTriggerScaleUp with "Maximum Allowed Cores exceeded for the Managed Environment".
    • Core usage was nowhere near the quota.
    • The Consumption profile was not affected.

    Root cause (from Microsoft) When dedicated nodes were removed during normal scale-in, a platform cleanup step did not fully release the network resources reserved for those nodes. Over many scale-in events, these stale reservations used up the capacity set aside for new dedicated nodes. New nodes could not finish their network setup, so they never became available for scheduling. The "Maximum Allowed Cores exceeded" message is a generic scale-up failure and did not indicate a real quota problem.

    Fix

    • Microsoft engineers released the stale reservations on the environment's backend. This cannot be done from the customer side.
    • As an interim step, raising the Dedicated profile's maximum node count (2 → 6 in my case) let a new node come online while the cleanup was completed.
    • After the cleanup, the minimum node count can go back to 0. Scaling up from zero worked again, and no quota reset was needed.
    • Microsoft says a platform fix for the cleanup step is rolling out and should be in production by the end of October 2026. They asked me to keep the maximum at 6 until then.

    If you see the same symptoms

    1. Move the workload to the Consumption profile in the meantime, if its CPU and memory limits allow.
    2. Open an Azure support case and reference "stale network reservations from dedicated node scale-in". Ask the support engineer to have them released for your environment.Microsoft support identified this as a platform-side issue, not a configuration or quota problem. Symptoms
      • Jobs on a Dedicated workload profile stayed Pending.
      • New dedicated nodes joined but never became schedulable. System logs showed AssigningReplicaFailed with "node(s) had untolerated taint(s)", then NotTriggerScaleUp with "Maximum Allowed Cores exceeded for the Managed Environment".
      • Core usage was nowhere near the quota.
      • The Consumption profile was not affected.
      Root cause (from Microsoft)
      When dedicated nodes were removed during normal scale-in, a platform cleanup step did not fully release the network resources reserved for those nodes. Over many scale-in events, these stale reservations used up the capacity set aside for new dedicated nodes. New nodes could not finish their network setup, so they never became available for scheduling. The "Maximum Allowed Cores exceeded" message is a generic scale-up failure and did not indicate a real quota problem. Fix
      • Microsoft engineers released the stale reservations on the environment's backend. This cannot be done from the customer side.
      • As an interim step, raising the Dedicated profile's maximum node count (2 → 6 in my case) let a new node come online while the cleanup was completed.
      • After the cleanup, the minimum node count can go back to 0. Scaling up from zero worked again, and no quota reset was needed.
      • Microsoft says a platform fix for the cleanup step is rolling out and should be in production by the end of October 2026. They asked me to keep the maximum at 6 until then.
      If you see the same symptoms
      1. Move the workload to the Consumption profile in the meantime, if its CPU and memory limits allow.
      2. Open an Azure support case and reference "stale network reservations from dedicated node scale-in". Ask the support engineer to have them released for your environment.

    Was this answer helpful?

    0 comments No comments

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.