Azure App Service is a service used to create and deploy scalable, mission-critical web apps.
Based on the evidence provided, this appears to be a TCP-level outbound connectivity issue, but the available data does not conclusively prove SNAT exhaustion.
1. Does SYN-SENT indicate SNAT exhaustion?
Not necessarily.
SYN-SENT only confirms that the client initiated a TCP connection and did not receive a SYN-ACK response before timing out. This could result from:
- SNAT exhaustion
- App Service outbound infrastructure issues
- Azure-to-AWS/Supabase routing problems
- Packet drops in transit
- Destination-side packet filtering or dropped responses
Microsoft documents SNAT exhaustion as a common cause of intermittent outbound connectivity failures in App Service, but SYN-SENT alone cannot identify where the packet was lost. [learn.microsoft.com]
In your case, several observations make a simple SNAT exhaustion explanation less convincing:
- Test and Production were both affected.
- Multiple Supabase pooler IPs failed simultaneously.
-
/dev/tcpfailed independently of Npgsql/Marten. - Test later recovered without scaling or configuration changes.
- HTTPS traffic continued to work.
2. Can historical SNAT telemetry be retrieved after scaling?
App Service provides SNAT Port Exhaustion and TCP Connections diagnostics through Diagnose and solve problems. [learn.microsoft.com], [learn.microsoft.com]
However, if historical data for the previous B2 worker is no longer available, there is no documented way to reconstruct its exact SNAT utilisation during the incident window. Therefore, I would avoid concluding that SNAT exhaustion was confirmed.
3. Could scaling have restored connectivity even if SNAT was not the cause?
Yes.
Microsoft documents that scaling between App Service tiers can change the outbound IP set and underlying infrastructure used by the application. [learn.microsoft.com], [learn.microsoft.com]
As a result:
Scale-up → connectivity restored
does not automatically mean:
Scale-up → SNAT exhaustion fixed
The scale operation may simply have moved traffic onto different outbound infrastructure or networking paths.
4. How can you determine whether the SYN packet actually left Azure?
From the App Service side alone, you generally cannot prove that.
The ss output confirms the connection attempt was created locally, but it cannot determine whether the packet successfully traversed Azure's managed outbound infrastructure.
For future incidents, capture:
date -u getent hosts <hostname> ss -tanp timeout 10 bash -c '</dev/tcp/<ip>/5432'
along with packet traces where possible. The strongest evidence would come from destination-side logs showing whether the SYN was received and whether a SYN-ACK was returned.
5. Is the Test B1 recovery significant?
Yes.
This is one of the stronger pieces of evidence that the issue may have been a transient networking condition rather than a problem specific to the Production app.
The combination of:
- impact across multiple App Services,
- failures across multiple destination IPs,
- successful access from outside Azure,
- low observed connection counts,
- and recovery of Test without scaling
suggests a broader transient Azure/App Service outbound networking issue or Azure-to-Supabase/AWS path issue remains a credible explanation.
Conclusion:
Based on the available evidence, I would not attribute the incident solely to SNAT exhaustion.
A more defensible RCA would be:
The incident was a transient TCP-level outbound connectivity failure between Azure App Service and the Supabase PostgreSQL endpoints. The
SYN-SENTstate confirms that no SYN-ACK was received, but does not establish where the packet was dropped. SNAT exhaustion remains a possible but unproven hypothesis. The simultaneous impact to multiple App Services, failure across multiple destination IPs, and recovery of the Test application without scaling are also consistent with a transient App Service networking or Azure-to-Supabase/AWS routing issue.
Help make this community better for everyone: if this answer resolved your issue, please accept it or leave an upvote. If not, share more details in a comment so we can continue the discussion and find the right solution