Azure Communication Services Call Automation media streaming misses the first sentences after call connection

AI MSI 20 Reputation points
2026-09-25T00:29:57.8466667+00:00

We are using Azure Communication Services Call Automation with the .NET SDK and realtime media streaming over a WebSocket.

Our backend starts media streaming when the call is created/answered and receives the MediaStreamingStarted event. However, the realtime WebSocket does not consistently capture the first few sentences spoken by the person immediately after the call connects.

Example behavior:

14:00:43 Call connected

14:00:44 MediaStreamingStarted

14:00:45 Customer begins speaking

14:00:51 First speech received by realtime STT: "you have reached [phone number of the person]"

The first few words or sentences before the received STT are missing, In some calls, the first captured stt begins only after several seconds of speech.

We use:

  • Azure Communication Services Call Automation
  • .NET SDK
  • StartMediaStreaming = True
  • Authenticated realtime media WebSocket transport URI
  • Azure Speech realtime transcription
  • CallConnected and MediaStreamingStarted events
  • PSTN calls and transferred/group calls

We would like to understand:

  1. Does ACS begin sending media only after the MediaStreamingStarted event?
  2. Can audio packets spoken immediately after CallConnected be dropped while the WebSocket connection is being established?
  3. Is there a recommended buffering or synchronization pattern to guarantee that the first customer speech is captured?
  4. Should media streaming be started during CreateCall/AnswerCall, or only after CallConnected?
  5. For group calls and transfers, when is customer audio guaranteed to be available?
  6. Is there a recommended delay before starting speech recognition or realtime STT?
  7. Are there ACS diagnostics or correlation IDs that can confirm whether the missing audio was never sent by ACS or was lost by our WebSocket receiver?

The missing audio is especially important because the first sentence may contain the voicemail greeting or the automated call-screening prompt used to classify the call.

What is the correct ACS implementation pattern for reliably capturing the complete audio from the beginning of a PSTN call?

Azure Communication Services

1 answer

Sort by: Newest
  1. Rukshan edirisinghe 910 Reputation points
    2026-09-25T04:54:36.0433333+00:00

    Hi @AI MSI

    Your timeline gives it away, so let me go point by point.

    1. Yes, ACS only sends audio after the WebSocket is up. It sends an AudioMetadata packet first, then AudioData. There's no pre-buffer, so speech before that first packet is gone.
    2. Yes, audio during socket setup is lost. Keep the gap small: startMediaStreaming true inside CreateCall/AnswerCall, and a warm, always-on endpoint that accepts instantly.
    3. Start your STT recognizer when the AudioMetadata packet arrives, not on the MediaStreamingStarted webhook. That webhook often lands after audio is already flowing, so waiting for it drops packets on your side. Queue audio until the recognizer is ready, then flush.
    4. Start streaming at Create/Answer. Calling the start API after CallConnected adds a round trip and loses more.
    5. Group calls and transfers: a participant's audio is there once they've joined. Use the Unmixed channel type to get per-participant audio with participantRawID so you can isolate the customer.
    6. No delay before STT. Your 6 second gap is far bigger than socket setup, which usually means the recognizer started late or you're only logging final Recognized results. Use the interim Recognizing events.
    7. To prove where audio went, log the timestamp and silent flag of the first AudioData packet against CallConnected. Non-silent packets before your first STT result means the loss is yours. First packet arriving late means socket setup. callConnectionId and correlationId tie into the Call Automation operational logs if Microsoft needs to check.

    If this helped, please click Accept Answer so others building voice bots can find it.

    References: https://learn.microsoft.com/en-us/azure/communication-services/how-tos/call-automation/audio-streaming-quickstart https://learn.microsoft.com/en-us/azure/communication-services/concepts/call-automation/audio-streaming-concept

    Was this answer helpful?

    0 comments No comments

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.