Azure Container Apps: responses over the environment-internal path are truncated after the first HTTP/2 window (65,535 bytes)

Nayer Girgis 0 Reputation points
2026-09-23T07:08:27.7566667+00:00

In a VNet-integrated Azure Container Apps environment with Consumption and Dedicated workload profiles, some HTTP responses are truncated when they cross an environment-internal path. The client receives a status line and headers with the expected Content-Length, followed by only part of the body before the connection closes. Comparable responses served through public ingress without an internal hop arrive complete.

Pacing application writes at approximately 4 MB/s prevents the truncation, but we would like to remove that workaround.

Setup

  • App A: Internal ingress, target port 8000, HTTP transport; gunicorn 21.2 on Python 3.11, with 4 workers and 4 threads per worker.

App B: External ingress, target port 80, HTTP transport; nginx 1.27 proxies to http://<app-a>:80 over HTTP/1.1.

Dapr is not enabled.

Test clients are a VM in the same VNet using curl 8.5.0 and Python urllib inside App A.

Observed behavior

A JSON response written in one piece with gzip and Content-Length: 84727 is consistently cut after 64,963 body bytes. That is close to the default 65,535-byte HTTP/2 flow-control window, although we have not confirmed that flow control is the cause.

For file responses with Content-Length, the cutoff varies between requests. For a 292,458-byte file, observed received lengths include 101,106, 137,446, 146,937, and 183,562 bytes.

In our Content-Length tests, bodies below 65,535 bytes arrive complete.

Through nginx, a truncated response ends about one second after it starts, and nginx logs upstream prematurely closed connection while reading upstream. Direct app-to-app requests fail sooner.

The upstream application logs HTTP 200 and no write error.

Reproductions

App to app, without nginx: From inside App A, request http://<app-a>:80/<static-file> with Accept-Encoding: identity. Three requests for a 292,458-byte file return HTTP 200 and Content-Length: 292458, but Python receives only 183,562, 137,446, and 101,106 bytes (IncompleteRead). For a 1,489,225-byte file, one request receives 269,158 bytes.

VM through App B to App A: A VNet VM requests App B’s public endpoint, which proxies to App A. For the 292,458-byte file, received lengths include 146,937, 146,226, and 179,877 bytes. Curl reports error 18, “transfer closed with N bytes remaining.”

Range request: Range: bytes=0-199999 returns HTTP 206 and Content-Length: 200000, but only 106,024 bytes arrive in three of three tests. Range: bytes=0-59999 completes.

JSON API: A gzip response with Content-Length: 84727 is cut at 64,963 bytes in two of two tests. A 56,634-byte response completes.

nginx as the upstream server: From inside App A, requests for a 4,080,366-byte static file served locally by App B’s nginx receive 1,589,002, 1,417,485, and 1,773,711 bytes. The same URL through App B’s public ingress completes in three of three tests.

TCP ingress: After switching App A’s ingress transport to TCP, a request from inside the environment to http://<app-a>:8000/... still truncates.

Other tests

Chunked, uncompressed responses can also truncate at varying points, sometimes below 65,535 bytes.

Chunked gzip responses streamed by an on-the-fly compressor complete at every tested size.

Chunked gzip responses written in one burst from memory complete in three of eight tests; the other five truncate at about 64.9 KB.

Explicit Content-Encoding: identity and changing writes from 64 KB to 4 KB do not resolve it.

Ingress transport settings auto, http, and tcp, and gunicorn worker models gthread (4×8 and 4×4) and sync (4×1), show the same behavior.

The issue reproduces when the environment is idle. Replicas use about 1 GB of their 4 GB allocation at peak.

Pacing writes in 8 KB pieces, with a 2 ms pause after each piece once 60 KB have been sent, has made every probe complete.

Could an internal ingress or proxy flow-control or buffering issue explain these results? Is there a platform-side fix or a supported environment or ingress setting that would allow full-speed writes without truncating responses?In a VNet-integrated Azure Container Apps environment with Consumption and Dedicated workload profiles, some HTTP responses are truncated when they cross an environment-internal path. The client receives a status line and headers with the expected Content-Length, followed by only part of the body before the connection closes. Comparable responses served through public ingress without an internal hop arrive complete.

Pacing application writes at approximately 4 MB/s prevents the truncation, but we would like to remove that workaround.

Setup

App A: Internal ingress, target port 8000, HTTP transport; gunicorn 21.2 on Python 3.11, with 4 workers and 4 threads per worker.

App B: External ingress, target port 80, HTTP transport; nginx 1.27 proxies to http://<app-a>:80 over HTTP/1.1.

Dapr is not enabled.

Test clients are a VM in the same VNet using curl 8.5.0 and Python urllib inside App A.

Observed behavior

A JSON response written in one piece with gzip and Content-Length: 84727 is consistently cut after 64,963 body bytes. That is close to the default 65,535-byte HTTP/2 flow-control window, although we have not confirmed that flow control is the cause.

For file responses with Content-Length, the cutoff varies between requests. For a 292,458-byte file, observed received lengths include 101,106, 137,446, 146,937, and 183,562 bytes.

In our Content-Length tests, bodies below 65,535 bytes arrive complete.

Through nginx, a truncated response ends about one second after it starts, and nginx logs upstream prematurely closed connection while reading upstream. Direct app-to-app requests fail sooner.

The upstream application logs HTTP 200 and no write error.

Reproductions

App to app, without nginx: From inside App A, request http://<app-a>:80/<static-file> with Accept-Encoding: identity. Three requests for a 292,458-byte file return HTTP 200 and Content-Length: 292458, but Python receives only 183,562, 137,446, and 101,106 bytes (IncompleteRead). For a 1,489,225-byte file, one request receives 269,158 bytes.

VM through App B to App A: A VNet VM requests App B’s public endpoint, which proxies to App A. For the 292,458-byte file, received lengths include 146,937, 146,226, and 179,877 bytes. Curl reports error 18, “transfer closed with N bytes remaining.”

Range request: Range: bytes=0-199999 returns HTTP 206 and Content-Length: 200000, but only 106,024 bytes arrive in three of three tests. Range: bytes=0-59999 completes.

JSON API: A gzip response with Content-Length: 84727 is cut at 64,963 bytes in two of two tests. A 56,634-byte response completes.

nginx as the upstream server: From inside App A, requests for a 4,080,366-byte static file served locally by App B’s nginx receive 1,589,002, 1,417,485, and 1,773,711 bytes. The same URL through App B’s public ingress completes in three of three tests.

TCP ingress: After switching App A’s ingress transport to TCP, a request from inside the environment to http://<app-a>:8000/... still truncates.

Other tests

Chunked, uncompressed responses can also truncate at varying points, sometimes below 65,535 bytes.

Chunked gzip responses streamed by an on-the-fly compressor complete at every tested size.

Chunked gzip responses written in one burst from memory complete in three of eight tests; the other five truncate at about 64.9 KB.

Explicit Content-Encoding: identity and changing writes from 64 KB to 4 KB do not resolve it.

Ingress transport settings auto, http, and tcp, and gunicorn worker models gthread (4×8 and 4×4) and sync (4×1), show the same behavior.

The issue reproduces when the environment is idle. Replicas use about 1 GB of their 4 GB allocation at peak.

Pacing writes in 8 KB pieces, with a 2 ms pause after each piece once 60 KB have been sent, has made every probe complete.

Could an internal ingress or proxy flow-control or buffering issue explain these results? Is there a platform-side fix or a supported environment or ingress setting that would allow full-speed writes without truncating responses?

Azure Container Apps
Azure Container Apps

An Azure service that provides a general-purpose, serverless container platform.


1 answer

Sort by: Most helpful
  1. Allan Solomon Mejia 10,145 Reputation points
    2026-09-23T19:17:48.85+00:00

    Hello @Nayer Girgis

    Based on the results, this doesn’t match a documented response-size limit. Azure Container Apps documentation doesn’t expose a supported setting for modifying the HTTP/2 flow-control window, response buffering, or maximum response-body size. Premium ingress provides configurable request limits and idle timeouts, but it doesn’t document a control for this failure pattern.

    The proximity of one cutoff to 65,535 bytes isn’t enough to confirm HTTP/2 flow control as the cause, particularly because:

    • Other responses terminate at varying sizes.
    • Chunked responses can also fail.
    • The problem persists with auto, http, and tcp transport.
    • The same content succeeds through the public ingress path.
    • Pacing the writes changes the outcome.

    This evidence suggests a failure within the environment-internal ingress or connection path, but I cannot find verified Microsoft documentation identifying it as a known issue or providing a supported application-side fix.

    Enable the environment’s ContainerAppHTTPLogs diagnostic category and correlate the failed requests by timestamp, revision, replica, upstream host, connection ID, and ingress/Envoy instance. Microsoft documents these logs as being emitted by the managed ingress layer specifically for diagnosing request traffic and disconnects.

    Given the consistent reproduction across different servers, worker models, encodings, and transport settings, open an Azure support request and provide:

    • Environment and region
    • App and revision names
    • Exact UTC timestamps
    • HTTP ingress logs and system logs
    • Client traces showing expected and received byte counts
    • nginx’s upstream prematurely closed connection entries
    • A minimal reproducible container or endpoint
    • Results from both the internal and public ingress paths

    Write pacing may remain a temporary mitigation, but don’t treat it as a platform-supported resolution.

    References:

    Ingress in Azure Container Apps

    Configure ingress for an Azure Container Apps environment

    Use premium ingress in Azure Container Apps

    Application logging in Azure Container Apps

    Monitor logs with Log Analytics

    Diagnose and solve problems in Azure Container Apps


    Help make this community better for everyone: If this answer helped or resolved your issue, please accept it or upvote it. If not, share more details in a comment so we can continue the discussion and find the right solution. Thank you.

    Was this answer helpful?

    0 comments No comments

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.