GPT-5.4 (Responses API) rejects dense prompts above 272k tokens with context_length_exceeded, but accepts English prompts up to 1M

Chris Cook 0 Reputation points
2026-09-24T10:39:19.1833333+00:00

Environment

  • Model: gpt-5.4, Azure AI Foundry, region Sweden Central, Responses API (/openai/v1/responses), streaming
  • reasoning.effort: none, small max_output_tokens

Problem

Some requests above 272k input tokens fail with:

{"type":"error","error":{"type":"invalid_request_error","code":"context_length_exceeded","message":"Your input exceeds the context window of this model. Please adjust your input and try again.","param":"input"},"sequence_number":2}

The error arrives inside the HTTP 200 stream, after response.created and response.in_progress. Whether a request fails does not depend on the token count. It depends on how many characters the prompt has per token.

Synthetic prompts, same deployment, token counts from the OpenAI token-count endpoint:

Prompt Tokens Characters Chars/token Chars ÷ 4 Result
English 285,008 1,725,828 6.06 431,457 accepted
Dense ASCII 285,018 424,223 1.49 106,056 context_length_exceeded
Japanese 284,996 471,727 1.66 117,932 context_length_exceeded
Chinese 285,004 393,590 1.38 98,398 context_length_exceeded
Japanese 699,986 1,158,607 1.66 289,652 accepted

A Japanese prompt of 285k tokens is rejected, but a Japanese prompt of 700k tokens is accepted. In production we also see English prompts accepted up to ~536k tokens on the same deployment. The same rejected prompt is accepted by OpenAI's API directly.

All results agree with this: the service estimates the prompt size as characters ÷ 4 to choose between a backend with a 272k limit and a long-context backend. Text with few characters per token (CJK, dense data) is estimated too low, goes to the 272k backend, and is rejected there when the real token count is above 272k.

Reproduction script (Node 18+, no dependencies, dry run by default): https://gist.github.com/zirkelc/7222464315271b3f43cf6f9488e80b2f

Questions

  1. Is this a known issue? Is a fix planned?
  2. Is there a supported way to make sure a request goes to the long-context backend (a header, a deployment type or a setting)?

A similar issue for gpt-4.1 (context_length_exceeded at ~300k tokens) was fixed in July 2025: https://learn.microsoft.com/en-us/answers/questions/2280883/

Azure OpenAI in Foundry Models

1 answer

Sort by: Most helpful
  1. Gulnaz Mushtaq 490 Reputation points MVP
    2026-09-25T17:37:16.1933333+00:00

    Based on your findings, it doesn't look like a GPT-5.4 context-window limitation but an Azure-side routing bug.

    It looks like the correlation between your chars/4 is a solid indication that Azure's routing and size estimation might be mistakenly relying on character-based estimation for certain inputs. To address this, the best course of action would be to open a support ticket with Azure. Be sure to include your reproduction steps, token-count measurements, and request IDs, while clearly outlining the suspected chars/4 routing issue.

    References:

    https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure

    https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/provisioned-throughput-sizing

    Was this answer helpful?

    0 comments No comments

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.