In a VNet-integrated Azure Container Apps environment with Consumption and Dedicated workload profiles, some HTTP responses are truncated when they cross an environment-internal path. The client receives a status line and headers with the expected Content-Length, followed by only part of the body before the connection closes. Comparable responses served through public ingress without an internal hop arrive complete.
Pacing application writes at approximately 4 MB/s prevents the truncation, but we would like to remove that workaround.
Setup
- App A: Internal ingress, target port 8000, HTTP transport; gunicorn 21.2 on Python 3.11, with 4 workers and 4 threads per worker.
App B: External ingress, target port 80, HTTP transport; nginx 1.27 proxies to http://<app-a>:80 over HTTP/1.1.
Dapr is not enabled.
Test clients are a VM in the same VNet using curl 8.5.0 and Python urllib inside App A.
Observed behavior
A JSON response written in one piece with gzip and Content-Length: 84727 is consistently cut after 64,963 body bytes. That is close to the default 65,535-byte HTTP/2 flow-control window, although we have not confirmed that flow control is the cause.
For file responses with Content-Length, the cutoff varies between requests. For a 292,458-byte file, observed received lengths include 101,106, 137,446, 146,937, and 183,562 bytes.
In our Content-Length tests, bodies below 65,535 bytes arrive complete.
Through nginx, a truncated response ends about one second after it starts, and nginx logs upstream prematurely closed connection while reading upstream. Direct app-to-app requests fail sooner.
The upstream application logs HTTP 200 and no write error.
Reproductions
App to app, without nginx: From inside App A, request http://<app-a>:80/<static-file> with Accept-Encoding: identity. Three requests for a 292,458-byte file return HTTP 200 and Content-Length: 292458, but Python receives only 183,562, 137,446, and 101,106 bytes (IncompleteRead). For a 1,489,225-byte file, one request receives 269,158 bytes.
VM through App B to App A: A VNet VM requests App B’s public endpoint, which proxies to App A. For the 292,458-byte file, received lengths include 146,937, 146,226, and 179,877 bytes. Curl reports error 18, “transfer closed with N bytes remaining.”
Range request: Range: bytes=0-199999 returns HTTP 206 and Content-Length: 200000, but only 106,024 bytes arrive in three of three tests. Range: bytes=0-59999 completes.
JSON API: A gzip response with Content-Length: 84727 is cut at 64,963 bytes in two of two tests. A 56,634-byte response completes.
nginx as the upstream server: From inside App A, requests for a 4,080,366-byte static file served locally by App B’s nginx receive 1,589,002, 1,417,485, and 1,773,711 bytes. The same URL through App B’s public ingress completes in three of three tests.
TCP ingress: After switching App A’s ingress transport to TCP, a request from inside the environment to http://<app-a>:8000/... still truncates.
Other tests
Chunked, uncompressed responses can also truncate at varying points, sometimes below 65,535 bytes.
Chunked gzip responses streamed by an on-the-fly compressor complete at every tested size.
Chunked gzip responses written in one burst from memory complete in three of eight tests; the other five truncate at about 64.9 KB.
Explicit Content-Encoding: identity and changing writes from 64 KB to 4 KB do not resolve it.
Ingress transport settings auto, http, and tcp, and gunicorn worker models gthread (4×8 and 4×4) and sync (4×1), show the same behavior.
The issue reproduces when the environment is idle. Replicas use about 1 GB of their 4 GB allocation at peak.
Pacing writes in 8 KB pieces, with a 2 ms pause after each piece once 60 KB have been sent, has made every probe complete.
Could an internal ingress or proxy flow-control or buffering issue explain these results? Is there a platform-side fix or a supported environment or ingress setting that would allow full-speed writes without truncating responses?In a VNet-integrated Azure Container Apps environment with Consumption and Dedicated workload profiles, some HTTP responses are truncated when they cross an environment-internal path. The client receives a status line and headers with the expected Content-Length, followed by only part of the body before the connection closes. Comparable responses served through public ingress without an internal hop arrive complete.
Pacing application writes at approximately 4 MB/s prevents the truncation, but we would like to remove that workaround.
Setup
App A: Internal ingress, target port 8000, HTTP transport; gunicorn 21.2 on Python 3.11, with 4 workers and 4 threads per worker.
App B: External ingress, target port 80, HTTP transport; nginx 1.27 proxies to http://<app-a>:80 over HTTP/1.1.
Dapr is not enabled.
Test clients are a VM in the same VNet using curl 8.5.0 and Python urllib inside App A.
Observed behavior
A JSON response written in one piece with gzip and Content-Length: 84727 is consistently cut after 64,963 body bytes. That is close to the default 65,535-byte HTTP/2 flow-control window, although we have not confirmed that flow control is the cause.
For file responses with Content-Length, the cutoff varies between requests. For a 292,458-byte file, observed received lengths include 101,106, 137,446, 146,937, and 183,562 bytes.
In our Content-Length tests, bodies below 65,535 bytes arrive complete.
Through nginx, a truncated response ends about one second after it starts, and nginx logs upstream prematurely closed connection while reading upstream. Direct app-to-app requests fail sooner.
The upstream application logs HTTP 200 and no write error.
Reproductions
App to app, without nginx: From inside App A, request http://<app-a>:80/<static-file> with Accept-Encoding: identity. Three requests for a 292,458-byte file return HTTP 200 and Content-Length: 292458, but Python receives only 183,562, 137,446, and 101,106 bytes (IncompleteRead). For a 1,489,225-byte file, one request receives 269,158 bytes.
VM through App B to App A: A VNet VM requests App B’s public endpoint, which proxies to App A. For the 292,458-byte file, received lengths include 146,937, 146,226, and 179,877 bytes. Curl reports error 18, “transfer closed with N bytes remaining.”
Range request: Range: bytes=0-199999 returns HTTP 206 and Content-Length: 200000, but only 106,024 bytes arrive in three of three tests. Range: bytes=0-59999 completes.
JSON API: A gzip response with Content-Length: 84727 is cut at 64,963 bytes in two of two tests. A 56,634-byte response completes.
nginx as the upstream server: From inside App A, requests for a 4,080,366-byte static file served locally by App B’s nginx receive 1,589,002, 1,417,485, and 1,773,711 bytes. The same URL through App B’s public ingress completes in three of three tests.
TCP ingress: After switching App A’s ingress transport to TCP, a request from inside the environment to http://<app-a>:8000/... still truncates.
Other tests
Chunked, uncompressed responses can also truncate at varying points, sometimes below 65,535 bytes.
Chunked gzip responses streamed by an on-the-fly compressor complete at every tested size.
Chunked gzip responses written in one burst from memory complete in three of eight tests; the other five truncate at about 64.9 KB.
Explicit Content-Encoding: identity and changing writes from 64 KB to 4 KB do not resolve it.
Ingress transport settings auto, http, and tcp, and gunicorn worker models gthread (4×8 and 4×4) and sync (4×1), show the same behavior.
The issue reproduces when the environment is idle. Replicas use about 1 GB of their 4 GB allocation at peak.
Pacing writes in 8 KB pieces, with a 2 ms pause after each piece once 60 KB have been sent, has made every probe complete.
Could an internal ingress or proxy flow-control or buffering issue explain these results? Is there a platform-side fix or a supported environment or ingress setting that would allow full-speed writes without truncating responses?