An Azure communication platform for deploying applications across devices and platforms.
Hi @AI MSI
Your timeline gives it away, so let me go point by point.
- Yes, ACS only sends audio after the WebSocket is up. It sends an AudioMetadata packet first, then AudioData. There's no pre-buffer, so speech before that first packet is gone.
- Yes, audio during socket setup is lost. Keep the gap small: startMediaStreaming true inside CreateCall/AnswerCall, and a warm, always-on endpoint that accepts instantly.
- Start your STT recognizer when the AudioMetadata packet arrives, not on the MediaStreamingStarted webhook. That webhook often lands after audio is already flowing, so waiting for it drops packets on your side. Queue audio until the recognizer is ready, then flush.
- Start streaming at Create/Answer. Calling the start API after CallConnected adds a round trip and loses more.
- Group calls and transfers: a participant's audio is there once they've joined. Use the Unmixed channel type to get per-participant audio with participantRawID so you can isolate the customer.
- No delay before STT. Your 6 second gap is far bigger than socket setup, which usually means the recognizer started late or you're only logging final Recognized results. Use the interim Recognizing events.
- To prove where audio went, log the timestamp and silent flag of the first AudioData packet against CallConnected. Non-silent packets before your first STT result means the loss is yours. First packet arriving late means socket setup. callConnectionId and correlationId tie into the Call Automation operational logs if Microsoft needs to check.
If this helped, please click Accept Answer so others building voice bots can find it.
References: https://learn.microsoft.com/en-us/azure/communication-services/how-tos/call-automation/audio-streaming-quickstart https://learn.microsoft.com/en-us/azure/communication-services/concepts/call-automation/audio-streaming-concept