SQL Database Resource Health Events — Unplanned Unavailability on September 28, 2026

Mike-7243 0 Reputation points
2026-09-28T16:55:39.72+00:00

Problem description

I am experiencing an unplanned availability incident with my Azure SQL Database. Around 15:36–15:37 UTC on September 28, 2026, application connections were refused, and error 9001 ('database is not currently available') was received by all application instances. Despite resource health indicating 'Available' and no health events being reported, the applications experienced about 14 minutes of failed logins.

Environment

Azure SQL Database, Hyperscale, HS_Gen5_40 compute, incident window 2026-09-28 15:37:30–15:52 UTC, US WEST

What I've already tried

I checked Resource Health, which showed 'Available' with no errors. I reviewed the guidance focusing on Resource Health, including login success/failure checks every 1–2 minutes, and examined automated diagnostics for login failures and platform reconfiguration events. Diagnostics indicated a brief reconfiguration and blocking sessions during the incident window. However, these diagnostics do not fully explain the ~14-minute login refusal period.

Current status

I am seeking a detailed explanation of the cause of the unavailability, clarification on the relationship between platform reconfiguration and application impact, and guidance on how to detect and mitigate similar incidents more effectively in the future.

Azure SQL Database
0 comments No comments

1 answer

Sort by: Newest
  1. Abinesh Magudeeswaran 230 Reputation points Student Ambassador
    2026-09-28T17:20:32.4833333+00:00

    The reported behavior is consistent with a transient Azure SQL Database connectivity/reconfiguration event, but the information provided is not sufficient to identify the exact platform root cause.

    Resource Health for Azure SQL Database is based on observed login success/failure telemetry and is updated approximately every 1–2 minutes. An Available state means Resource Health did not detect system-related login failures at the threshold required to change the health state; it does not mean that every individual connection attempt succeeded during the entire interval.

    Azure SQL documentation also specifically notes that reconfigurations can occur as transient conditions, including those triggered by load balancing or software/hardware failures. During maintenance or reconfiguration, existing connections can be terminated and new connections can temporarily fail. Microsoft recommends robust connection retry logic for production applications.

    For this incident, I would correlate the following data using UTC timestamps:

    Application connection failures and error 9001

    Azure SQL Resource Health history

    Azure Activity Log/resource-health events

    Azure SQL diagnostic logs

    Availability metric

    Any platform reconfiguration/failover events

    Client-side connection retry behavior

    Resource Health history is particularly useful because Azure SQL reports downtime reasons there when they are available. The reported health information can also have a delay relative to the actual login failures.

    For future detection, consider configuring Azure Monitor alerts for the Azure SQL Database Availability metric and Resource Health/Activity Log events rather than relying only on the portal's current Resource Health state. Azure SQL supports availability monitoring and alerts, while Resource Health supports Activity Log-based alerting.

    For a roughly 14-minute period of refused connections, I would open a Microsoft support request if the Resource Health history does not provide a corresponding explanation. Provide the database/server resource ID, region, exact UTC incident window (15:37:30–15:52 UTC), error 9001 samples, correlation/request IDs where available, and the diagnostic evidence showing the reconfiguration.

    Microsoft would need the service-side telemetry to determine whether the reconfiguration was the direct cause of the connection failures or whether another platform event occurred concurrently. I would avoid concluding from the client diagnostics alone that the reconfiguration was the root cause.

    Was this answer helpful?


Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.