Azure App Service PostgreSQL TCP Connections Stuck in SYN-SENT and Timing Out

Ratnakar Sriram 0 Reputation points
2026-09-30T06:02:41.78+00:00

We experienced an outbound PostgreSQL connectivity incident affecting two Linux Azure App Services (Test and Production) connecting to separate Supabase PostgreSQL projects.

During the incident:

  • Production and Test Azure App Services could not connect to their Supabase PostgreSQL Session Poolers on TCP 5432.
  • DNS resolution worked correctly.
  • Direct /dev/tcp tests from the Production App Service container timed out, independently of .NET/Npgsql/Marten.
  • During a controlled test, ss showed: 169.254.130.9:<ephemeral-port> -> 3.131.201.192:5432 remaining in SYN-SENT until the 10-second timeout.
  • All three resolved Supabase pooler IPv4 addresses failed from Azure.
  • The same Supabase services were reachable from a local machine.
  • Local application connectivity to our Test Supabase project worked while the Azure Test App Service could not connect.
  • Normal outbound HTTPS/443 from Azure worked.
  • Supabase DB banned IP list was empty.
  • Supabase network restrictions allowed 0.0.0.0/0 and ::/0.
  • App Service had no VNet integration, NSG, UDR, private endpoint, or NAT Gateway.
  • Restarting Production App Service did not resolve the problem.
  • Container socket usage was low: 29 sockets, 8 TCP in use, 0 orphaned, 0 TIME_WAIT.

Production was originally on a B2 App Service Plan. Scaling Production to a Production-tier plan restored database connectivity immediately.

However, our Test App Service remained on B1 and later recovered without any configuration or scaling changes.

After scaling Production, Azure SNAT diagnostics show approximately 128 TCP SNAT ports allocated and only ~10–15 used. Unfortunately, B2 did not expose this historical SNAT diagnostic, so we cannot see SNAT utilization during the actual outage.

Confirmed failure timestamps include approximately:

  • Sep 29 20:45 UTC – Azure TCP failed
  • Sep 29 20:53 UTC – Azure TCP failed
  • Sep 29 20:55 UTC – Azure TCP failed

Observed Azure public outbound IPv4 was 52.154.243.184.

We are trying to determine the actual root cause rather than assuming that scaling fixed it.

Questions:

  1. Does SYN-SENT until timeout in this situation indicate an App Service SNAT/egress problem, or could the same behavior result from an upstream Azure→AWS/Supabase routing issue?
  2. Is there any way to retrieve historical SNAT usage/allocation/failure telemetry for the previous B2 worker after the App Service Plan has been scaled?
  3. Can scaling the App Service Plan move the application to different worker/outbound infrastructure and therefore clear an outbound networking problem even if SNAT exhaustion was not the cause?
  4. How can we determine whether SYN packets actually left Azure and whether SYN-ACK responses returned?
  5. Given that a separate B1 Test App Service experienced the same problem and later recovered without scaling, could this indicate a transient App Service/platform or Azure-to-AWS routing issue rather than per-app SNAT exhaustion?

We would appreciate guidance on what Azure telemetry can establish the root cause retrospectively.We experienced an outbound PostgreSQL connectivity incident affecting two Linux Azure App Services (Test and Production) connecting to separate Supabase PostgreSQL projects.

Azure App Service
Azure App Service

Azure App Service is a service used to create and deploy scalable, mission-critical web apps.

0 comments No comments

1 answer

Sort by: Oldest
  1. Vinodh247-1375 44,801 Reputation points Volunteer Moderator
    2026-09-30T09:16:59.94+00:00

    Based on the evidence provided, this appears to be a TCP-level outbound connectivity issue, but the available data does not conclusively prove SNAT exhaustion.

    1. Does SYN-SENT indicate SNAT exhaustion?

    Not necessarily.

    SYN-SENT only confirms that the client initiated a TCP connection and did not receive a SYN-ACK response before timing out. This could result from:

    • SNAT exhaustion
    • App Service outbound infrastructure issues
    • Azure-to-AWS/Supabase routing problems
    • Packet drops in transit
    • Destination-side packet filtering or dropped responses

    Microsoft documents SNAT exhaustion as a common cause of intermittent outbound connectivity failures in App Service, but SYN-SENT alone cannot identify where the packet was lost. [learn.microsoft.com]

    In your case, several observations make a simple SNAT exhaustion explanation less convincing:

    • Test and Production were both affected.
    • Multiple Supabase pooler IPs failed simultaneously.
    • /dev/tcp failed independently of Npgsql/Marten.
    • Test later recovered without scaling or configuration changes.
    • HTTPS traffic continued to work.

    2. Can historical SNAT telemetry be retrieved after scaling?

    App Service provides SNAT Port Exhaustion and TCP Connections diagnostics through Diagnose and solve problems. [learn.microsoft.com], [learn.microsoft.com]

    However, if historical data for the previous B2 worker is no longer available, there is no documented way to reconstruct its exact SNAT utilisation during the incident window. Therefore, I would avoid concluding that SNAT exhaustion was confirmed.

    3. Could scaling have restored connectivity even if SNAT was not the cause?

    Yes.

    Microsoft documents that scaling between App Service tiers can change the outbound IP set and underlying infrastructure used by the application. [learn.microsoft.com], [learn.microsoft.com]

    As a result:

    Scale-up → connectivity restored

    does not automatically mean:

    Scale-up → SNAT exhaustion fixed

    The scale operation may simply have moved traffic onto different outbound infrastructure or networking paths.

    4. How can you determine whether the SYN packet actually left Azure?

    From the App Service side alone, you generally cannot prove that.

    The ss output confirms the connection attempt was created locally, but it cannot determine whether the packet successfully traversed Azure's managed outbound infrastructure.

    For future incidents, capture:

    date -u getent hosts <hostname> ss -tanp timeout 10 bash -c '</dev/tcp/<ip>/5432'

    along with packet traces where possible. The strongest evidence would come from destination-side logs showing whether the SYN was received and whether a SYN-ACK was returned.

    5. Is the Test B1 recovery significant?

    Yes.

    This is one of the stronger pieces of evidence that the issue may have been a transient networking condition rather than a problem specific to the Production app.

    The combination of:

    • impact across multiple App Services,
    • failures across multiple destination IPs,
    • successful access from outside Azure,
    • low observed connection counts,
    • and recovery of Test without scaling

    suggests a broader transient Azure/App Service outbound networking issue or Azure-to-Supabase/AWS path issue remains a credible explanation.

    Conclusion:

    Based on the available evidence, I would not attribute the incident solely to SNAT exhaustion.

    A more defensible RCA would be:

    The incident was a transient TCP-level outbound connectivity failure between Azure App Service and the Supabase PostgreSQL endpoints. The SYN-SENT state confirms that no SYN-ACK was received, but does not establish where the packet was dropped. SNAT exhaustion remains a possible but unproven hypothesis. The simultaneous impact to multiple App Services, failure across multiple destination IPs, and recovery of the Test application without scaling are also consistent with a transient App Service networking or Azure-to-Supabase/AWS routing issue.

    Help make this community better for everyone: if this answer resolved your issue, please accept it or leave an upvote. If not, share more details in a comment so we can continue the discussion and find the right solution

    Was this answer helpful?

    0 comments No comments

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.