Azure SQL Change Event Streaming: producer stops delivering to one Event Hubs partition after a transient outage and never reconnects

Eugenio Rossetto 0 Reputation points
2026-09-22T20:54:48.7866667+00:00

Environment

  • Azure SQL Database, General Purpose (reproduced on both serverless and provisioned compute), CES preview enabled.

Destination: Event Hubs Standard namespace, Kafka endpoint (mynamespace.servicebus.windows.net:9093/myhub), @destination_type = N'AzureEventHubs', Microsoft Entra managed-identity credential, @partition_key_scheme = N'Table', JSON encoding, 4 partitions, 14 tables in one stream group.

Scenario

After a transient connectivity interruption, delivery to one partition fails while the other partitions keep delivering normally. If the interruption is short (~8 minutes in one captured episode), the producer recovers on its own. If it lasts longer (~35 minutes in a contrasting episode with the identical error signature), the producer never recovers: it retries the same batch every 5 seconds against the dead connection indefinitely (29 hours in our worst case), batch_start_lsn in sys.dm_change_feed_log_scan_sessions and the checkpoint (seqno) stay frozen, and the healthy partitions receive the same window re-published as duplicates (~1,300 msg/h). The only recovery is sp_disable_change_event_stream plus a full group rebuild. We have documented 14 episodes of this over four weeks.

Errors (repeating every ~5 s in sys.dm_change_feed_errors; two variants)

23692 Delivery of the message failed due to a timeout...
23679 Unable to reach the configured destination.
23658 Change Event Streaming encountered a SQL exception.
22782 Failed to output batch data for table N in partition XXXX
22783 Failed to process commit work item
22777 Failure reported in table group / database processing

Troubleshooting done

Event Hubs healthy throughout every episode: zero ServerErrors/UserErrors/ThrottledRequests, namespace Active. We enabled namespace diagnostic logs (KafkaCoordinatorLogs, KafkaUserErrorLogs, OperationalLogs, DiagnosticErrorLogs) and captured an episode live: while CES reported error 23679, Event Hubs logged zero errors and no connection activity from the producer. Other connections from the same producer deliver to the same namespace during the failure, and rebuilding the group with identical configuration and credential restores delivery instantly, so network, firewall and authentication are ruled out. Reproduced on serverless (0.5 and 1 vCore minimum) and provisioned (GP_Gen5_2) compute. No Resource Health or activity-log events in any window. No CDC, replication, Synapse Link or mirroring on the database. Followed the configure guide and its limitations section. Failures concentrate on one partition per period (same partition id across consecutive incidents); onsets show no time-of-day pattern and historically cluster around the first writes after idle stretches.

Question

Is this a known defect in the CES preview producer's connection recovery, and is there a supported way to make the producer re-establish a failed partition connection without sp_disable_change_event_stream (which destroys the group, the table enrollment and the error ring buffer, and skips the wedged partition's undelivered events past the checkpoint reset)?

We have full captures (error ring buffers, scan sessions, checkpoints, Event Hubs diagnostic logs) for several episodes and can share them.

Azure SQL Database

1 answer

Sort by: Newest
  1. Mladen Andzic - Msft 11 Reputation points Microsoft Employee
    2026-10-01T10:11:09.9666667+00:00

    Hi @Eugenio Rossetto , greetings from the SQL product group. I'd suggest opening a support case so the feature team can investigate it. I'm sure they'll get to the bottom of it.

    Was this answer helpful?

    0 comments No comments

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.