gpt-chat-latest (2026-05-05) intermittently answers an unrelated, invented question instead of the request sent

NK 5 Reputation points
2026-09-27T21:24:30.0666667+00:00

Problem: Since 2026-09-25 we have seen chat-completion calls return HTTP 200 with content that has nothing to do with the prompt, reading like the answer to a different user's generic question. Three production examples:

  1. 2026-09-25 ~20:46 UTC: asked to write a roleplay system prompt, the model returned an essay comparing embedded hardware platforms (MCU/MPU/SoC/FPGA/DSP).
  2. 2026-09-25 ~23:30 UTC (json_schema structured output): the reply's message field said "Arya Stark does not die in HBO's Game of Thrones…" in answer to a business scenario description.
  3. 2026-09-27 11:18 UTC: asked to summarize coaching insights, the model returned generic troubleshooting advice for "Access denied" errors. Request IDs: 3f494da2-28e9-4125-b3dd-dd19ff32e2be and 05bd650f-bbd5-44f4-8d77-eab3e93d9b40.

Prompt token counts were normal and complete in every case. Replaying the same inputs usually produces correct output, so the failure is intermittent.

Questions: (1) Is this a known issue with gpt-chat-latest 2026-05-05? (2) Can you inspect the two request IDs above and confirm what was served? (3) Is there any possibility of responses being generated against another request's context?

Azure OpenAI in Foundry Models

1 answer

Sort by: Newest
  1. Taz 10,046 Reputation points MVP Volunteer Moderator
    2026-09-29T07:23:19.19+00:00

    Hi NK,

    The behavior you describe is not expected, but there is an important detail with this deployment.

    Microsoft lists gpt-chat-latest (2026-05-05) as a Preview version with a retirement date of August 5, 2026. The current model list shows newer gpt-chat-latest versions, including 2026-06-24 and 2026-08-06.

    I would therefore first check the actual model version currently attached to the deployment rather than relying only on the deployment name. Microsoft documents that the deployed model version can be verified in Foundry or through the deployments API.

    For your other question, Chat Completions is stateless and the application supplies the conversation history with each request. Microsoft also states that customer prompts and completions are not available to other customers.

    The two request IDs cannot be inspected from Microsoft Q&A. They need to be provided to Azure support, since apim-request-id is specifically intended for troubleshooting.

    Given that the same request normally succeeds on retry, I would capture the exact request body, deployment/model version, timestamp, response headers, and apim-request-id for every failure and provide those to Microsoft Support for service-side investigation.

    Was this answer helpful?

    1 person found this answer helpful.

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.