What a Translating LLM Gateway Drops - Thinking Blocks, Signatures, Cache Breakpoints, Tool Call IDs, and Citations on Each Route Between the Messages, Chat Completions, Responses, and Converse APIs
First Published:
Last Updated:
max_tokens in the Messages format with max_output_tokens in the Responses format, and stop_sequences with stop in the Chat Completions format. However, conversation history contains elements that aren't included in a simple parameter mapping. These include the thinking block from the previous turn, along with its signature, encrypted reasoning, prompt cache breakpoints, tool call IDs, and citations, as well as stream events. Depending on the conversion route, these elements may be carried over directly, converted into different fields, or simply dropped with no error and no warning.The LLM gateway, agentgateway, changed in version 1.6.0 (released on October 2, 2026) to prioritize OpenAI Responses over OpenAI Chat Completions when converting requests from Anthropic Messages, for providers that support both formats. The release notes for version 1.6.0 state that converting to Responses will result in the loss of the thinking history, while converting to Chat Completions will preserve it. The agentgateway documentation also states that agentgateway drops any features of the Messages format that Responses cannot represent from the converted request, with no error and no warning. However, it's not just the gateway that drops these elements. According to Anthropic's documentation, when sending a request containing a thinking block from Claude Sonnet 5.5 with an account other than the one that created it (or an account it is linked to), the API will drop that block before it reaches the model, and the request will still succeed. Anthropic's documentation on "preserved thinking" also cautions that libraries, proxies, and gateways that modify the history should be aware that those modifications count as edits.
This article details what information is carried, converted, dropped, or rejected for each type of data as the same conversation history passes through a conversion route. It also identifies who makes these decisions (the conversion gateway, clients or libraries that modify the history, and the receiving provider), and describes what is visible from the caller's perspective. The information is based on documentation, API references, release notes, and changelogs from each provider, as well as pull requests on the agentgateway GitHub repository, verified on October 9, 2026. The gateway was not run locally for this article, and this article is neither a product comparison nor a recommendation of any particular option.
Related articles on this site:
- Which MCP Clients a Remote MCP Server Lets In - Client ID Metadata Documents, Pre-Registration, Deprecated Dynamic Client Registration, and the Allowlists Layered on Top
- Where AI Agent Authorization Standards Stand - RFCs, IETF Working Group Documents, Individual Drafts, and Specifications from Other Bodies
- LLM API Parameter Compatibility Reference - Anthropic, OpenAI, Google Gemini, and Amazon Bedrock
- Mid-Conversation Changes in the Claude API - System Messages, Tool Changes, Per-Message Effort, and What the Prompt Cache and Preserved Thinking Keep
- Amazon Bedrock Inference Endpoints and API Keys - What bedrock-runtime and bedrock-mantle Each Accept, Which IAM Action Each Call Checks, and What CloudTrail Records
- Amazon Bedrock Converse API Deep Dive - Unified Model Interface, Conversation State, Tool Use, and Streaming
- Anthropic Claude API Prompt Caching and Token Efficiency Guide - Cache Breakpoints, Batch Processing, and Context Engineering
- Anthropic Claude API Errors Reference - HTTP Status Codes, Error Types, Causes, and Official Solutions
- Controlling How Much a Model Thinks - Why the Same Request Means Different Things Across Model Generations, and What the Output Cap Actually Counts
- Practical Guide to Deploying Claude Code and Claude Desktop Behind a Corporate Proxy - Domains, MSIX, NTLM/Kerberos, and VPN Coexistence
- Claude Code on Pay-As-You-Go API Billing - Anthropic API, Amazon Bedrock, and Google Cloud Vertex AI
- Managed Agent Runtimes - What Claude Managed Agents, the OpenAI Agents API, Amazon Bedrock, and GKE Agent Substrate Each Run for You, and What Stays With You
- Anthropic Claude Model Release Timeline - Model Family Tree, Capability Evolution, and Platform Availability
- OpenAI GPT Model Release Timeline - Model Lineage, ChatGPT and Codex Milestones, and Platform Availability
Table of Contents
- 1. The Scope of This Article and the Date It Was Verified
- 2. How to Read the Tables — Hops, Outcome Words, and What the Caller Sees
- 3. Where Each of the Four Formats Keeps Each Kind of Information
- 4. What Happens on a Gateway's Conversion Routes — agentgateway 1.6.0
- 5. What Providers Drop and Reject
- 6. What Is Lost When a Client, a Library, or a Pass-Through Gateway Rewrites the Request
- 7. Combinations the Target Format Does Not Support
- 8. Ways to Notice That Something Was Dropped
- 9. Where the Sources Disagree and Where No Statement Was Found
- 10. Frequently Asked Questions about Format Conversion in LLM Gateways
- 11. Summary
- 12. References
1. The Scope of This Article and the Date It Was Verified
This section defines the scope of this article and specifies the date of verification. Furthermore, it outlines the sources consulted, the terminology used in this article, and what topics will not be addressed.1.1 What Is Not Carried Even When Parameter Names Map
The LLM API Parameter Compatibility Reference lists the request parameters for Anthropic Messages, OpenAI Chat Completions and Responses, Google Gemini, and Amazon Bedrock Converse, mapping them by name and value range. Section 13 of the document highlights eight potential compatibility pitfalls, all stemming from differences in request parameter names, value ranges, placement, and presence. The article explicitly states that comparisons of proxy and gateway products are outside the scope.It does not cover pricing, rate limits, or benchmark scores; per-model exhaustive default tables; SDK-language-specific syntax; or a comparison of proxy or gateway products.
This article focuses on the content within conversation histories and responses. It does not address parameter name mapping. Specifically, it covers six categories:
- Thinking and Signature: The reasoning generated by the model in the previous turn, along with its associated
signature. - Encrypted Reasoning: Anthropic's
redacted_thinking, OpenAI Responses'encrypted_content, and Bedrock Converse'sredactedContent. - Cache Breakpoints: Markers that show the boundary for reusing the prompt cache.
- Tool Call IDs: IDs that link the model's tool calls with the tool results returned by the client.
- Citations: Information indicating the source location associated with text in the reply.
- Stream Event Granularity: The unit in which the above information is delivered within a streaming response.
This article primarily focuses on three conversion routes: from Anthropic Messages requests to other provider formats – specifically, to Responses, Chat Completions, and Converse. Section 4.6 also addresses one reverse route, from a Chat Completions client to an Anthropic Messages provider. Outside of these six core categories, this article briefly mentions parameters (such as
stop_sequences) and server tools in tools that the agentgateway documentation also says are dropped.1.2 The Verification Date and the Sources Read
This article's information was verified on October 9, 2026. At the time of verification, the latest versions were agentgateway 1.6.0 (released on GitHub on October 2, 2026, at 16:26:26 UTC) and Claude Code 2.1.295 (October 8, 2026). agentgateway's GitHub release list has no version after 1.6.0 yet. Claude Code 2.1.296 was released after the check, on October 9, 2026, at 19:28:59 UTC. Its CHANGELOG entry was reviewed on October 10, 2026, and it did not include any items related to the gateway hint headers orANTHROPIC_BASE_URL (checked 2026-10-10). The three places in this article that depend on sources that changed after the check (the Claude Code version in this paragraph, the signature compatibility in Section 3.2, and the system and developer messages in Section 5.3) were read again on October 10, 2026, and each is marked checked 2026-10-10. The models in scope appear in their providers' release notes and changelogs with these dates: Claude Sonnet 5.5 (September 28, 2026), Claude Haiku 5.5 (October 7, 2026), GPT-6 Astra (September 3, 2026), and GPT-6.1 Sol (September 29, 2026).The materials consulted are as follows, categorized by provider:
- agentgateway: The agentgateway documentation, specifically the Messages page (latest and 1.5.x), Chat completions, Amazon Bedrock, Custom providers, Anthropic, and Claude Code pages (latest). Also, GitHub releases (1.6.0), pull requests #3363, #3491, #3496, #3516, and the source code for GitHub's v1.6.0 tag (located in
crates/llm/src/conversion/, includingresponses.rs,completions.rs,bedrock.rs, andmod.rs). - Anthropic: The Anthropic documentation, including the Thinking, Preserved thinking, Prompt caching, Citations, Streaming messages, Handle tool calls, OpenAI SDK compatibility, What's new in Claude Sonnet 5.5, What's new in Claude Haiku 5.5, the Migration guide for Claude Haiku 5.5, Troubleshooting thinking, and each page of the Claude Platform release notes.
- Claude Code: The Claude Code documentation, including the Other LLM gateways page and the Claude Code gateway compatibility guide, as well as the CHANGELOG and GitHub releases.
- OpenAI: The OpenAI API documentation, including the Migrate to the Responses API, Reasoning models, Using GPT-6, Prompt caching, Function calling, Streaming API responses, Changelog, and the Chat and Responses sections of the API reference.
- Amazon Bedrock: The Amazon Bedrock user guide, including the Inference using Converse API, Prompt caching for faster model inference, Responses API, Chat Completions API, and the model cards (GPT-6 Sol, GPT-6.1 Sol, GPT-6 Luna) pages, as well as the API reference for
ContentBlock,ContentBlockDelta,ReasoningTextBlock,ToolUseBlock, andToolResultBlock.
This article discusses the conversion routes of the gateway, based on the information available in the agentgateway documentation and release notes. The agentgateway itself is an implementation that documents, with versions, what its conversions drop in both the documentation and release notes. This article does not cover conversions within other gateways. Neither the gateway nor the models were run locally; all information comes from publicly available documentation.
1.3 Terminology Used in This Article
Words with the same spelling may refer to different entities depending on the format or product. This article tells them apart as follows:- thinking: This refers to the reasoning the model writes before a reply or between tool calls. In Anthropic Messages, it's represented as a
thinkingblock; in OpenAI Responses, it's an item labeled "reasoning"; and in Bedrock Converse, it's areasoningContentblock. Thereasoning_contentthat agentgateway uses in its conversion to Chat Completions is the field that agentgateway's docs describe for self-hosted engines that report reasoning asreasoning_content. agentgateway carries the signature along this route asreasoning_signature. - signature: This is the
signatureassociated with thethinkingblock. Descriptions of this field vary in different documentation (see Section 3.2). - LLM gateway: This is a proxy placed between the client and the model provider, responsible for relaying requests. Some gateways convert the format, while others pass it through unchanged. The Claude apps gateway (mentioned frequently in Claude Code's CHANGELOG) and the Gateway in Amazon Bedrock AgentCore are distinct from the LLM gateways discussed in this article.
- cache: This refers to the prompt cache. A response cache, where a gateway stores and returns replies, is out of scope.
- ID: This is the ID of a tool call. The
idof an OpenAI Responses item, andprevious_response_idused to connect conversations, are treated as separate from the ID of a tool call (see Section 3.3). - history and reply: History refers to past turns that the client resends in a request. A reply is what the model returns in response to that request. Even when dealing with the same information, the conversion results may differ between the history and the reply.
- buffer and stream: A buffered reply is a complete reply returned as a single body. A streamed reply is a reply returned incrementally as a series of events.
- route: This refers to the direction of a format conversion. This article names the source and the target, as in the route from Messages to Responses.
1.4 What This Article Does Not Cover
The LLM API Parameter Compatibility Reference covers the mapping of parameter names and their value ranges, as well as the names and placement of cache breakpoints. Section 11 of that reference lists Anthropic'scache_control, Bedrock's cachePoint, and OpenAI's prompt_cache_breakpoint. This article only covers whether breakpoints are carried on a conversion route.For the Claude API, Mid-Conversation Changes in the Claude API details the impact of mid-conversation changes on the thinking blocks and prompt cache. That article explains two types of checks related to the prompt cache and preserved thinking, as well as how thinking blocks are lost when the model is switched. This article will not recreate the mechanisms of those checks, and will present the results as rows at the provider hop.
Amazon Bedrock Inference Endpoints and API Keys covers which APIs Amazon Bedrock's
bedrock-runtime and bedrock-mantle each accept, and the IAM checks. Amazon Bedrock Converse API Deep Dive describes the usage of Converse, and Anthropic Claude API Prompt Caching and Token Efficiency Guide details the design of cache_control breakpoints. Anthropic Claude API Errors Reference documents the types of 400 errors, and Controlling How Much a Model Thinks explains the meaning of "effort".The "Other LLM gateways" section for Claude Code states that Anthropic does not endorse, maintain, or audit third-party gateway products, and does not support sending Claude Code through a gateway to models other than Claude. Placing Claude Code behind a gateway that converts the Messages format to OpenAI's format and then sending it to other models is a usage that this statement does not support.
Anthropic doesn't endorse, maintain, or audit third-party gateway products, and doesn't support routing Claude Code to non-Claude models through any gateway.
Practical Guide to Deploying Claude Code and Claude Desktop Behind a Corporate Proxy and Claude Code on Pay-As-You-Go API Billing cover setting up Claude Code behind a gateway. Managed Agent Runtimes details the infrastructure for running managed agents. The Claude and OpenAI model timeline articles cover model release dates.
How a remote MCP server decides which MCP clients it lets in is covered in Which MCP Clients a Remote MCP Server Lets In, and the stages of the RFCs and drafts related to AI agent authorization are covered in Where AI Agent Authorization Standards Stand. Both articles borrow this article's way of reading the tables (Section 2).
This article does not cover the Google Gemini format or OpenAI's Decisions API (released in beta on October 6, 2026). It also does not cover comparisons of gateway products, configuration procedures, or pricing.
2. How to Read the Tables — Hops, Outcome Words, and What the Caller Sees
This section defines the terms and columns used in the tables in this article. The tables represent hops, describe what happens using four outcome words, and depict what is visible to the caller using four visibility words.2.1 Hops
Requests travel through several hops as they move from the client to the model. This article uses the following four hops. TheHop column in the table lists the English names for each hop.Client: This encompasses the application, its SDK, frameworks, and agent harnesses like Claude Code. It assembles the history, sends it, receives responses, and adds them to the next turn's history.Gateway: This is a proxy located between the client and the provider. It may convert the format of the request, or simply pass it through unchanged.Provider: This refers to the API that receives the request. It includes Anthropic's Claude API, OpenAI API, Amazon Bedrock, and self-hosted engines located beyond the gateway.Model: This is the final destination. In the sources this article reads, the provider's API is responsible for dropping blocks. TheModelhop is used solely to indicate what data ultimately reaches the model.

2.2 Outcome Words (What Happens)
TheWhat Happens column in the table indicates what occurs with the information at that hop, using one of the following four words. The verb used in the documentation determines the term.Carried: Use this term when the documentation states that the information is carried, preserved, kept, covered, or allowed through, and does not specify a particular destination field.Converted: Use this term when the documentation specifies a destination field or format.Dropped: Use this term when the documentation states that the information is dropped, omitted, ignored, removed, stripped, discarded, lost, not kept, or not returned, or that the result has no such item.Rejected: Use this term when the documentation indicates that the request fails due to an error (rejected, returns a400status, or results in an error). Include the status code if the documentation gives one.
2.3 What the Caller Sees
TheWhat the Caller Sees column indicates whether the caller will see a result, using one of the following four terms:No error and no warning: The documentation states that no errors or warnings are generated, that the request is successful, or that the result is silent.Reported only on request: The caller is only notified when they specifically request it. For example, theinput_transformationssection of the response will only display blocks if a beta header is sent.Error: An error response comes back. If the documentation gives the status code and message, the cell includes them.The source does not say.: The documentation does not state what the caller will see.
This article only states that blocks are silently dropped when the combination is
Dropped and No error and no warning. If the documentation only mentions that blocks are dropped, but does not specify whether an error occurred, the combination should be Dropped and The source does not say. For rows labeled Carried and Converted, as nothing is lost, enter — in this column. However, if the documentation describes what the caller sees, use that description instead.2.4 Common Columns, and How Status and Versions Are Written
The tables presenting results will each have the following six columns in common:Hop: The hop that causes the information loss (see Section 2.1).Decided By: The component and version responsible for that hop. Examples include agentgateway 1.6.0's conversion to Responses, or the Claude API's model check. If the check of another hop determines a result (for example, when a provider's check rejects a gateway rewrite), document that check as well.What Happens: The outcome described in Section 2.2, along with relevant documentation (or a summary thereof).What the Caller Sees: The observed behavior described in Section 2.3, along with relevant documentation (or a summary thereof).Status and Version: The version and date, along with the current status (e.g.,GA,beta), and the name of any beta headers. Also include the date of verification.Where the Source Says So: The documentation that supports the information in that row.
Columns that have the same value across all rows of a table should be listed only once, before the table itself. For example, the tables for each route in Section 4 all have
Hop as Gateway and Decided By as the conversion of agentgateway 1.6.0. Therefore, these two should be listed before the tables. The tables in Sections 7 and 8, which list combination constraints and methods of detection, have content different from the result tables and will therefore have columns tailored to their specific purpose.Versions and dates should be presented as pairs, in the format
<product> <version> (<YYYY-MM-DD>). For example, agentgateway 1.6.0 (2026-10-02) or Claude Code 2.1.288 (2026-10-02). The dates should reflect the release time (UTC) as shown on GitHub. Verification dates should be written as checked 2026-10-09 (checked 2026-10-10 for the three places listed in Section 1.2).2.5 When the Source Does Not Say, and When Sources Disagree
In theWhat Happens and What the Caller Sees columns, only include information explicitly stated in the documentation (or a summary thereof). If the documentation does not say something, write The source does not say. Before writing this, search the full text of the relevant documentation, as well as any linked pages and other pages from the same provider (reference guides, migration guides, changelogs, release notes, model cards), using the subject of the row and the terms drop, ignore, preserve, carry, omit, not supported, and silently. If the documentation only provides a general rule and does not specifically mention the route in question, do not write The source does not say. Instead, include the general rule and a note indicating that the route is not explicitly mentioned.When there are discrepancies between different pieces of documentation, list both descriptions with their respective sources, without favoring either. This article collects such disagreements in Section 9.
3. Where Each of the Four Formats Keeps Each Kind of Information
Before discussing the conversion routes, this section will outline how each of the four formats represents five types of information (excluding cache breakpoints). Information is dropped in a conversion if a field from the source format does not have a corresponding field in the target format, or if the gateway is not converting it.3.1 Fields by Kind of Information
The following table lists the fields that each provider's docs and API reference use for each kind of information. Section 11 of the LLM API Parameter Compatibility Reference covers the names and placement of cache breakpoints, so this table leaves them out.| Information | Anthropic Messages | OpenAI Chat Completions | OpenAI Responses | Amazon Bedrock Converse |
|---|---|---|---|---|
| Thinking and Signature | thinking block's thinking and signature | No corresponding field found in the API reference to carry the full reasoning text (the reasoning_effort parameter and reasoning_tokens in usage are available). | Reasoning item. Does not return the full text; a summary (summary) is returned only when requested. | reasoningContent's reasoningText's text and signature. |
| Encrypted Reasoning | redacted_thinking block's data | No corresponding field found in the API reference. | Reasoning item's encrypted_content (included by default when store is false or within ZDR organizations). | reasoningContent's redactedContent. |
| Tool Call ID | tool_use block's id and tool_result's tool_use_id | Each id in the assistant's tool_calls and tool_call_id in the tool's message. | call_id in the function_call item and call_id in the function_call_output. A separate id exists for the item itself. | toolUse's toolUseId and toolResult's toolUseId (between 1 and 64 characters, pattern [a-zA-Z0-9_.:-]+). |
| Citations | text block's citations (including cited_text) | annotations on the assistant message (url_citation, when the web search tool is used). | annotations in the output text (including url_citation, file_citation, etc.). | Block within citationsContent (when document citations are enabled). |
| Stream Events | Within content_block_delta: text_delta, input_json_delta, thinking_delta, signature_delta, citations_delta | chat.completion.chunk's delta | Typed events (including response.output_text.delta) | Within contentBlockDelta in ConverseStream: text, reasoningContent, citation, toolUse. |
The two lines of Chat Completions results are from searching the OpenAI API reference page for Chat, obtained on October 9, 2026, using the terms
reasoning, encrypted, and redacted. Only reasoning_effort and reasoning_tokens were found; no fields related to the actual reasoning content or encrypted reasoning were identified.The OpenAI guide for Reasoning models states that responses generated in stateless mode will include the
encrypted_content property within reasoning items by default. It also notes that the API accepts reasoning.encrypted_content in include for compatibility, but that it is not strictly necessary.When you create a response in stateless mode, reasoning items in the response's `output` array include an `encrypted_content` property by default. Stateless mode applies when `store` is `false` or when your organization uses Zero Data Retention (ZDR). The API still accepts the legacy `reasoning.encrypted_content` value in `include` for compatibility, but doesn't require it.
3.2 Three Descriptions of the Signature
Anthropic's documentation, Amazon Bedrock's user guide, and Bedrock's API reference describe the signature on the thinking block in different ways.Anthropic's "Thinking" page describes a signature as an encrypted copy of the full reasoning, passed back unchanged in multi-turn and tool-use conversations.
Each thinking block also carries a `signature` field, an encrypted copy of the full reasoning that you pass back unchanged in multi-turn and tool-use conversations
The same page states that signature values are compatible across platforms, but the next sentence adds an exception: thinking blocks from Claude Sonnet 5.5 and Claude Haiku 5.5 can only be used in the account that created them or in an account linked to that account (checked 2026-10-10).
`signature` values are compatible across platforms (the Claude API, Amazon Bedrock, and Google Cloud). Values generated on one platform work on another, except that thinking blocks from Claude Sonnet 5.5 and Claude Haiku 5.5 work only in the account that produced them or in an account linked to it.
This exception was added to the page after this article was verified on October 9, 2026, and it corresponds to the account check described in Section 5.1.
Bedrock's user guide, on the "Inference using Converse API" page, describes the
signature within reasoningText as a hash of all messages in a conversation, designed to prevent tampering with the reasoning. It further states that signatures and all previous messages must be included in subsequent Converse requests, and that any change to a message will result in an error response.The `signature` field is a hash of all the messages in the conversation and is a safeguard against tampering of the reasoning used by the model. You must include the signature and all previous messages in subsequent `Converse` requests. If any of the messages are changed, the response throws an error.
Bedrock's API reference, in the
ReasoningTextBlock, describes a signature as a token that verifies that the model generated the reasoning text.A token that verifies that the reasoning text was generated by the model. If you pass a reasoning block back to the API in a multi-turn conversation, include the text and its signature unmodified.
While these three explanations use different phrasing to describe what a signature contains, this article does not attempt to determine which is correct. All three, however, emphasize the requirement to send the signature unchanged.
3.3 A Tool Call Can Have More Than One ID
In OpenAI Responses, the tool call item has two identifiers:id and call_id. In the examples in the Function calling guide, the id of the function_call item begins with fc_, while the call_id begins with call_. The client uses the call_id when returning the tool's result. The "Migrate to the Responses API" guide states that in Responses, tool calls and their outputs are separate items, linked by the call_id.In Responses, tool calls and their outputs are two distinct types of Items that are correlated using a `call_id`.
In Anthropic Messages, the
tool_use block has a single id, and the tool_result includes a tool_use_id that references it. Therefore, the gateway that converts Messages history to Responses must decide whether to map the id from tool_use to the call_id or the id of the item in Responses. Section 4.3 covers how agentgateway decides. The previous_response_id, used to connect conversations, is a separate ID that refers to the response itself, and is distinct from the tool call ID.4. What Happens on a Gateway's Conversion Routes — agentgateway 1.6.0
This section outlines the conversion routes, detailing them route by route. This article focuses on agentgateway because its documentation clearly identifies the specific routes and versions where information is dropped, and documents these in the documentation and release notes. Other gateways may not behave in the same way.4.1 Which Route a Messages Request Takes
The agentgateway can route Anthropic Messages requests to providers that do not support the Messages format. It converts the request to match the provider's format and then converts the reply back into the Messages format. The format used for conversion is the first format in the following order that the provider supports. The order differs between versions 1.5.x and the latest (1.6.0) as documented.| Order | agentgateway 1.5.x | agentgateway 1.6.0 |
|---|---|---|
| 1 | messages (no conversion) | messages (no conversion) |
| 2 | completions (converted to Chat Completions) | responses (converted to Responses) |
| 3 | responses (converted to Responses) | completions (converted to Chat Completions) |
| 4 | Bedrock Converse (converted to Converse) | Bedrock Converse (converted to Converse) |
The release notes for version 1.6.0 highlight this change as a key consideration before upgrading.
Agentgateway now prefers OpenAI Responses when it converts Anthropic Messages requests. In 1.5, it preferred Chat Completions. This affects OpenAI, Ollama, Groq, Hugging Face, and xAI providers; Azure models other than Claude; and custom providers that advertise both Responses and Chat Completions.
The Responses conversion drops extended-thinking history, while the Chat Completions conversion preserves it.
The latest documentation indicates that for providers supporting both formats, Responses is prioritized. It also notes that this conversion drops thinking history.
When a provider supports both `responses` and `completions`, agentgateway prefers Responses. The Responses conversion drops extended-thinking history. The Chat Completions conversion preserves it.
The documentation mentions an environment variable as a temporary workaround to prioritize Chat Completions. This variable applies to all providers and is planned for removal in version 1.7.
For providers that support both formats, the `AGENTGATEWAY_MESSAGES_PREFER_COMPLETIONS` environment variable provides a temporary workaround for Responses conversion bugs. To prefer Chat Completions, set the variable to `true` on the agentgateway process. The variable applies to every provider and is planned for removal in version 1.7.
The documentation advises using the
custom provider for servers compatible with OpenAI (the agentgateway page for Claude Code shows a different configuration; Section 9.4). If the server does not support /v1/responses, only completions should be declared. It also says to use the same configuration to keep the thinking history. The Messages page notes that Bedrock providers only support Converse, so Messages requests are always converted to Converse. Conversely, the Amazon Bedrock page explains that agentgateway chooses whether to send requests to Bedrock's Runtime or Mantle endpoints, depending on the model. It adds that models that resolve to Mantle accept the formats in their tags, except anthropic.claude* models, which always take the Anthropic Messages format (Section 9.2). The description of the route to Converse in this article refers to sending requests to Runtime.4.2 Results by Route at a Glance
The following table summarizes the key results from Sections 4.3 through 4.5, presenting one row per type of information. The tables for each section include details regarding server tools intools, the specification of thinking for models such as gpt-5.3, and encrypted reasoning returned to the Chat Completions client. Each cell comes from the documentation for agentgateway 1.6.0 (latest, checked 2026-10-09). However, the "Responses" column for the rows related to tool call IDs comes from a commit message in pull request #3491, while the rows concerning cache breakpoints come from the source code for tag v1.6.0 (Sections 4.3 to 4.5). Parenthetical notes within cells provide additional context for the results. The notation No error and no warning indicates that the documentation explicitly states this condition. When Dropped appears without No error and no warning in its parentheses, the documentation does not describe what the calling side observes.| Information | Messages to Responses | Messages to Chat Completions | Messages to Converse |
|---|---|---|---|
| Thinking in the history | Dropped (No error and no warning) | Converted (reasoning_content) | Converted (reasoningContent.reasoningText) |
| Signature on thinking in the history | Dropped (with the block; No error and no warning) | Converted (reasoning_signature; only when the turn has one signed block; otherwise Dropped) | Converted (the signature in reasoningContent.reasoningText) |
redacted_thinking in the history | Dropped (No error and no warning) | Dropped | Converted (reasoningContent.redactedContent) |
| Thinking in a buffered reply | Dropped (the reply has no thinking block) | Converted (from reasoning_content to thinking) | Converted (from the reasoning text to thinking) |
| Encrypted or withheld reasoning in a buffered reply | Dropped (the docs say the reasoning output is dropped; encrypted reasoning is not named) | Converted (reasoning the engine withholds becomes a signature-only thinking block) | Converted (redacted_thinking, keeping the content) |
| Thinking in a streamed reply | Dropped | Converted (thinking_delta, and signature_delta when there is a signature) | Converted (the reasoning text) |
| Encrypted or withheld reasoning in a streamed reply | Dropped (same as above) | Converted (reasoning the engine withholds becomes a signature-only block) | Dropped (arrives as a thinking block with [REDACTED] and cannot be replayed) |
| Cache breakpoints in the history | Converted (prompt_cache_breakpoint when the model name starts with gpt- and has a major.minor version of 5.6 or later; otherwise Dropped; from the source code) | Converted (same as above; from the source code) | Converted (cachePoint, up to four; from the source code) |
| Tool call IDs in the history | Converted (call_id; no item id) | The source does not say. | The source does not say. |
Citations on text blocks in the history | Dropped (the text is kept; No error and no warning) | General rule only (Section 9.1) | The source does not say. |
| Citations in a buffered reply | Converted (URL citations only; cited_text and encrypted_index are empty) | General rule only (Section 9.1) | The source does not say. |
| Citations in a streamed reply | Dropped | General rule only (Section 9.1) | The source does not say. |
| Document, search-result, and server-tool blocks in the history | Dropped (No error and no warning) | The source does not say. | Dropped |

The source does not say. are the result of searching the agentgateway documentation, specifically the "Messages," "Chat completions," "Amazon Bedrock," "Custom providers," and "Anthropic" pages, as well as the release notes for version 1.6.0, using the terms cache, id, call_id, tool_call_id, toolUseId, citation, document, search_result, and server tool. No statement was found about tool call IDs, about citations on the Converse route, or about document, search-result, and server-tool blocks on the Chat Completions route. The search terms correspond to the subjects of the rows, and the relevant lines were read using the verbs described in Section 2.5. Regarding citations on the "Chat Completions" route, the release notes for version 1.6.0 state a general rule for transferring citations between Anthropic Messages and OpenAI formats, but there is no specific mention of the route itself (Section 9.1). The body of pull request #3496 begins with the line Align to completion path compatibility. but does not specifically mention citations on the "Chat Completions" route or the handling of blocks on that route.4.3 The Messages-to-Responses Route
The latest documentation states that agentgateway drops Messages features that the Responses format cannot represent from the converted request, with no error and no warning to the client. The docs list the following four as dropped.A Messages feature that the Responses format cannot represent is dropped from the converted request, with no error and no warning to the client. These features are dropped this way:
- Thinking and redacted-thinking history, so the model loses its prior reasoning on each turn
- Citations on text blocks in the message history. The text itself is kept.
- Document, search-result, and server-tool content blocks, and content blocks of a type that agentgateway does not recognize, including these parts of a tool result
- Server tools, such as web search, in the `tools` list
Regarding differences when converting a reply back to the Messages format, the documentation mentions three points. Two of these are detailed below; the third concerns how to handle failures in streaming responses.
- The reasoning output of the model is dropped, so the reply has no `thinking` block.
- A buffered reply keeps each URL citation as a `web_search_result_location` citation with the source `url` and `title`. The `cited_text` and `encrypted_index` fields are empty strings, because the Responses format does not return them. File citations and `logprobs` are dropped. A streamed reply has no citations.
The docs list what this route covers (a common agent subset) as including: text and system instructions, images, tool definitions and selection, the history of assistant tool calls and tool results, structured output, reasoning effort, prompt cache breakpoints, and reporting of streaming and usage. The documentation does not specify what a prompt cache breakpoint will become after the conversion.
In the source code for GitHub version 1.6.0, both the conversion from Messages to Responses (
responses.rs) and the conversion from Messages to Chat Completions (completions.rs) use the same evaluation function. This function only adds a prompt_cache_breakpoint to the part with cache_control when the model name begins with gpt- and the version that follows is in major.minor form (such as 5.6) and is 5.6 or higher. Otherwise, it is not added. A model name whose version has no minor part (such as gpt-6-sol) does not meet this condition. The completions.rs file contains the following comment, which is a description in the source code, not a description in the documentation. A commit in pull request #3516 (titled "llm: messages to responses better cache control"; listed in the 1.6.0 release notes; merged on September 16, 2026) also says to use it only on supported models.// Explicit prompt-cache breakpoints are accepted only by GPT 5.6 and newer models.
The handling of tool call IDs is not described in the documentation. A commit within pull request #3491 (listed in the 1.6.0 release notes; merged on September 15, 2026), specifically the commit
33c3339, states that the tool_use.id from Messages should be preserved as a call_id, which refers to the call itself, and that the optional item id should be omitted. The stated reason for this is to avoid HTTP 400 errors during replay. This is a message in the commit, and not a description in the agentgateway documentation.Messages tool_use.id identifies the call, not the Responses item. Preserve it as call_id and omit the optional item id to avoid HTTP 400 on replay.
The results of this route use the common columns. Each row indicates that
Hop is Gateway, Decided By refers to agentgateway 1.6.0 (2026-10-02)'s conversion to Responses, and Status and Version corresponds to the latest documentation (checked 2026-10-09). However, the rows related to tool call IDs come from the commit in pull request #3491, while the rows indicating cache breakpoints come from the source code for version 1.6.0.| Information | What Happens | What the Caller Sees | Where the Source Says So |
|---|---|---|---|
Thinking and redacted_thinking in the history | Dropped. The model loses its prior reasoning on each turn. | No error and no warning | agentgateway docs (Messages) |
Citations on text blocks in the history | Dropped. The text itself is kept. | No error and no warning | agentgateway docs (Messages) |
| Document, search-result, and server-tool blocks in the history, and blocks of types not recognized by agentgateway | Dropped. This includes items within tool results. | No error and no warning | agentgateway docs (Messages) |
Server tools within tools (e.g., web search) | Dropped. | No error and no warning | agentgateway docs (Messages) |
| ID of tool calls within the history | Converted. The tool_use.id is maintained as call_id, and the id within the item is omitted. | — | Commit message from pull request #3491 |
| Cache breakpoints in the history | Converted. If the model name starts with gpt- and has a major.minor version of 5.6 or later, the label prompt_cache_breakpoint is added. Otherwise, it is dropped (the documentation only states that this is within the scope of handling). | The source does not say. | agentgateway docs (Messages), v1.6.0 source (responses.rs) |
| Thinking in a reply (buffered or streamed) | Dropped. The reply does not contain a thinking block. | The source does not say. (Only that the reply does not contain a thinking block is known.) | agentgateway docs (Messages) |
| Citations in a buffered reply | Converted. URL citations become web_search_result_location. cited_text and encrypted_index are empty strings. File citations and logprobs are Dropped. | The source does not say. | agentgateway docs (Messages) |
| Citations in a streamed reply | Dropped. A streamed reply has no citations. | The source does not say. | agentgateway docs (Messages) |
The documentation indicates that when using the same route,
stop_sequences and top_k are also dropped, and the warning section documents this behavior.The Responses format has no equivalent for `stop_sequences` or `top_k`. Agentgateway accepts both fields and drops them, with no error and no warning to the client.
This concerns parameters, and it's a detail outside of the six types discussed in this article. According to Section 13, item 8 of the LLM API Parameter Compatibility Reference,
stop and other parameters are silently dropped when moving from Chat Completions to Responses.4.4 The Messages-to-Chat Completions Route
The latest documentation says that the conversion to Chat Completions carries thinking history in both directions. However, it clarifies that this transfer is made possible by a self-hosted engine that returns reasoning asreasoning_content and also accepts reasoning_content sent back within assistant messages.The Chat Completions conversion carries extended-thinking history in both directions, so a thinking session on a converted route keeps its prior reasoning from one turn to the next. Self-hosted engines that report reasoning as `reasoning_content` also accept it back on an assistant message, which is what makes the carryover possible.
On the sending side, the assistant's
thinking block from the history is sent as reasoning_content. On the receiving side, the reasoning_content of a buffered reply becomes the thinking block preceding the text block. In streaming mode, the thinking block begins with a thinking_delta event, and a signature_delta is added when the engine sends its signature. If the engine withholds reasoning, both the buffered reply and the streamed reply will deliver a thinking block containing only empty text and the signature.The documentation lists three scenarios where this round trip does not occur. The first is when a single turn contains two or more
thinking blocks with signatures.The thinking text is joined, but no `reasoning_signature` is sent. The signature is carried only when the turn has exactly one signed block.
The second is the presence of
redacted_thinking blocks.Dropped, because it holds nothing that an OpenAI-compatible engine can replay.
The third is when the request is converted to Responses, and, as described in Section 4.3, the thinking history is dropped with no error and no warning. The documentation states that the first two scenarios result in data loss, but it does not specify whether an error will occur. This article's table lists
What the Caller Sees as The source does not say.The documentation also describes a scenario where the thinking specification is dropped. Some models reject Chat Completions requests that simultaneously specify reasoning effort and tools. The documentation provides
gpt-5.3 as an example. Requests with tools sent to that model are sent with reasoning_effort: "none", and any thinking that the client asks for through thinking or output_config.effort is dropped.Certain models, such as `gpt-5.3`, reject a Chat Completions request that sets both a reasoning effort and tools. In the Chat Completions conversion, a request with tools to one of these models is sent with `reasoning_effort: "none"`, and any thinking that the client asked for through `thinking` or `output_config.effort` is dropped.
The documentation does not list any other models that exhibit this behavior, aside from the example of
gpt-5.3.The results of this route use the common columns. Each row indicates that
Hop is Gateway, Decided By refers to agentgateway 1.6.0 (2026-10-02)'s conversion to Chat Completions, and Status and Version corresponds to the latest documentation (checked 2026-10-09). However, the row indicating cache breakpoints comes from the source code for version 1.6.0.| Information | What Happens | What the Caller Sees | Where the Source Says So |
|---|---|---|---|
| Thinking in the history | Converted. Sent as reasoning_content. A turn made of thinking alone is also sent. | — | agentgateway docs (Messages) |
| Signature on thinking in the history | Converted. Sent as reasoning_signature, only when the turn has exactly one signed block. If there are two or more, the text is combined, and reasoning_signature is not sent (dropped). | The source does not say. | agentgateway docs (Messages) |
redacted_thinking in the history | Dropped, because it holds nothing an OpenAI-compatible engine can replay. | The source does not say. | agentgateway docs (Messages) |
| Thinking in a buffered reply | Converted. The reasoning_content becomes a thinking block ahead of the text block. | — | agentgateway docs (Messages) |
| Thinking in a streamed reply | Converted. Starts with a thinking_delta event, and if a signature exists, a signature_delta is added. | — | agentgateway docs (Messages) |
| Reasoning the engine withholds, in a reply (buffered or streamed) | Converted. Becomes a thinking block with empty text and only a signature. | — | agentgateway docs (Messages) |
Thinking requested in a request with tools (some models such as gpt-5.3) | Dropped. Sent with reasoning_effort: "none". | The source does not say. | agentgateway docs (Messages) |
| Cache breakpoints in the history | Converted. If the model name starts with gpt- and has a major.minor version of 5.6 or later, a prompt_cache_breakpoint is added. Otherwise, it is dropped (no description in the documentation). | The source does not say. | v1.6.0 source (completions.rs) |
4.5 The Messages-to-Converse Route (Bedrock Runtime)
The latest documentation says that the conversion to Bedrock Converse carries thinking history in both directions, including encrypted reasoning.The Bedrock Converse conversion carries extended-thinking history in both directions, including encrypted reasoning.
On the sending side, agentgateway converts the history's thinking blocks into blocks containing a signed
reasoningContent.reasoningText. The redacted_thinking block becomes a reasoningContent.redactedContent block, allowing the client to replay the encrypted reasoning previously returned by Bedrock. agentgateway drops document, search-result, and server-tool blocks.On the receiving side, Bedrock's reasoning text becomes a thinking block. In a buffered reply, Bedrock's encrypted reasoning becomes a
redacted_thinking block, preserving its content for the next turn. In a streamed reply, the content is not preserved.A streamed reply does not keep the payload. The encrypted reasoning arrives as a `thinking` block with the text `[REDACTED]`, which cannot be replayed.
The Amazon Bedrock page for agentgateway provides a table illustrating how the same encrypted reasoning comes back to the client in a buffered reply, depending on the client's API. For clients using
/v1/messages, it's returned as a redacted_thinking block. For a /v1/responses client, it comes back as a reasoning item with encrypted_content. In either case, if the client resends it, it becomes reasoningContent.redactedContent in the Bedrock request. Regarding clients using /v1/chat/completions, the documentation states:Omitted, because the Chat Completions format has no field for encrypted reasoning. Signed reasoning text still arrives in `reasoning_content` and `reasoning_signature`.
The cell for replay on the next turn says
Not possible.Regarding cache breakpoints, the conversion from Messages to Converse (within
bedrock.rs's from_messages) at the v1.6.0 tag inserts cachePoint markers at each position that carries cache_control, limiting the number to four. Unlike conversions to Responses and Chat Completions, it does not filter by model name. This is a detail from the source code, and not a statement in the official documentation.The results of this route use the common columns. Each row indicates that
Hop is Gateway, Decided By reflects a conversion to Converse by agentgateway 1.6.0 (2026-10-02), and Status and Version refers to the latest documentation (checked 2026-10-09). However, the row indicating cache breakpoints comes from the source code for version v1.6.0.| Information | What Happens | What the Caller Sees | Where the Source Says So |
|---|---|---|---|
| Thinking and its signature in the history | Converted. Becomes reasoningContent.reasoningText with a signature. | — | agentgateway docs (Messages) |
redacted_thinking in the history | Converted. Becomes reasoningContent.redactedContent. | — | agentgateway docs (Messages) |
| Document, search-result, and server-tool blocks in the history | Dropped. | The source does not say. | agentgateway docs (Messages) |
| Cache breakpoints in the history | Converted. Inserts a cachePoint at the location with cache_control. Up to four entries. Does not filter by model name (no description in the documentation). | — | v1.6.0 source (bedrock.rs) |
| Encrypted reasoning in a buffered reply | Converted. Becomes redacted_thinking while preserving the content. | — | agentgateway docs (Messages, Amazon Bedrock) |
| Reasoning text in a reply | Converted. Becomes a thinking block. | — | agentgateway docs (Messages) |
| Encrypted reasoning in a streamed reply | Dropped. Arrives as a thinking block with the text [REDACTED], and cannot be replayed. | The source does not say. (Arrives as a thinking block with [REDACTED] in the reply.) | agentgateway docs (Messages, Amazon Bedrock) |
| Encrypted reasoning in a buffered reply to a Chat Completions client | Dropped. Chat Completions format does not have a field for encrypted reasoning. | The source does not say. | agentgateway docs (Amazon Bedrock) |
4.6 The Reverse Route — From a Chat Completions Client to Messages
The agentgateway also converts requests from the Chat Completions client and sends them to the Anthropic Messages provider. According to the agentgateway's Chat completions page, this reverse route involves carrying the reasoning history between turns. On the sending side, assistant messages containing a non-emptyreasoning_signature along with reasoning_content are replayed as signed thinking blocks. A reasoning_content without a signature is not sent, as the provider will reject unsigned thinking blocks. On the receiving side, the signature from the reply's thinking block comes back as the reasoning_signature. If a buffered reply contains two or more thinking blocks, agentgateway combines those blocks into a single reasoning_content and does not send the reasoning_signature.The same page includes a note regarding clients that discard signatures.
Send the signature back along with the reasoning. A client that keeps `reasoning_content` but discards `reasoning_signature` loses its thinking history on the next turn, with no error and no warning.
One reason why the history of thinking is silently dropped along this route is clients that discard the
reasoning_signature. As mentioned above, if a buffered reply contains two or more thinking blocks, the gateway also does not send the reasoning_signature. Section 6.1 collects the patterns in which the client is the cause.4.7 What Changed from 1.5 to 1.6.0 — From a 400 to Silence, and From Chat Completions to Responses
The documentation for agentgateway 1.5.x states that if a Messages feature cannot be represented at all in the Responses format, it will fail before sending the request upstream, returning a400 with an unsupported conversion message (the 1.5.x documentation is still published on the verification date). The first item listed is the history of thinking and redacted_thinking.A Messages feature that the Responses format cannot represent at all fails before the upstream request is sent, and returns a `400` with an `unsupported conversion` message. These features fail this way:
- Thinking and redacted-thinking history
- Document, search-result, and server-tool content blocks
- Tool results that are not text
The latest (1.6.0) documentation, as mentioned in Section 4.3, states that the same thinking history is now dropped without generating any errors or warnings. The body of pull request #3496 (merged on September 15, 2026), listed in the 1.6.0 release notes, states that the thinking history, previously rejected, is now skipped so that the request succeeds. It also notes that the thinking data is still lost. This is from the pull request description, not the agentgateway documentation itself.
Thinking history: Previously rejected; now skipped so the request succeeds. Thinking is still lost: preserving summaries and encrypted reasoning remains the TODO part of #3479
The same description also states that regarding citations, agentgateway previously rejected even empty
citations: [], and now keeps the text while ignoring the citation metadata.Citations: Previously, even citations: [] caused rejection. Now we preserve the text and ignore citation metadata, including in tool results.
The following table sets the two versions of the documentation side by side, for the same information in the conversion to Responses.
| Information | agentgateway 1.5.x documentation | agentgateway 1.6.0 documentation (latest) |
|---|---|---|
History of thinking and redacted_thinking | Rejected (400, unsupported conversion) – fails before being sent upstream. | Dropped (No error and no warning). |
| Document, search-result, and server-tool blocks in the history | Rejected (400, unsupported conversion). | Dropped (No error and no warning). |
| Tool results that are not text | Rejected (400, unsupported conversion). | Image tool results are in the subset the conversion covers. Blocks of an unrecognized type are Dropped. |
| Which route a request takes | completions comes before responses. A provider that supports both is not affected by the Responses conversion. | responses comes before completions. |
The change from 1.5.x to 1.6.0 results in two different behaviors, depending on the format declared by the provider.
One type is providers that declare
responses but not completions. The 1.5.x documentation states that requests to these providers are converted to responses. In the 1.5.x documentation, sending the history of thinking to these providers results in a 400 rejection. In the 1.6.0 documentation, the history is dropped without any error or warning. While the thinking history still doesn't reach the model, the difference lies in whether the calling application can detect this.Because `completions` comes before `responses`, a provider that advertises both is unaffected by the Responses conversion. That conversion applies to a provider that advertises `responses` and not `completions`.
The other type is providers that declare both formats. The 1.6.0 release notes list providers such as OpenAI, Ollama, Groq, Hugging Face, xAI, as well as Azure models (excluding Claude) and
custom providers that declare both formats. In 1.5.x, requests to these providers went through a conversion to Chat Completions. The 1.5.x documentation does not describe what happens to the history of thinking on that route. The body of pull request #3363 (merged on September 11, 2026, after 1.5.0, the only stable 1.5.x release, had come out on August 27, 2026), listed in the 1.6.0 release notes, states that before this pull request's change, the request conversion dropped the assistant's thinking blocks. This is from the pull request description, not the agentgateway documentation, and the description does not say whether an error occurred.Before this change none of the three carried anything: the response translation had no thinking block to build, and the request translation dropped assistant `thinking` blocks on the way out.
"The three" in the description are the buffered reply, the streamed reply, and the history sent back, which the description shows in its verification.
In 1.6.0, the conversion to Chat Completions carries the history of thinking, but requests to these providers are now converted to
responses by default, and the history of thinking is dropped without any error or warning. According to the 1.6.0 documentation and the body of pull request #3363, with these providers, under the default route, the history of thinking does not reach the model in either version. What changes when upgrading to 1.6.0 is the route the requests take, and whether the documentation states that the history is dropped.For providers that declare only
responses, describing this gateway as "dropping thinking" or "rejecting it" without specifying a version contradicts the documentation for the other version.5. What Providers Drop and Reject
This section addresses issues that can occur independently of the gateway's conversion process. The receiving provider's API checks incoming blocks, either dropping blocks the model cannot use or rejecting the request. Mid-Conversation Changes in the Claude API discusses Anthropic's mechanism for the "preserved thinking" checks. This article presents the results of those checks as rows at the hop.5.1 Results at the Provider Hop
The following table lists results, whereHop is Provider in every row, and the verification date is October 9, 2026.| Decided By | What Happens | What the Caller Sees | Status and Version | Where the Source Says So |
|---|---|---|---|---|
| Claude API model check (a thinking block the current model cannot read) | Dropped. Removed from that request before it reaches the model. Applies to every account; the check started with Claude Fable 5.1. | No error and no warning. If the thinking-binding-controls-2026-08-01 header is sent, Reported only on request (model_binding_mismatch). | checked 2026-10-09 | Preserved thinking |
| Claude API account check (Claude Sonnet 5.5 and Claude Haiku 5.5 thinking blocks sent by an account other than the one that produced them or a linked account) | Dropped. The request succeeds, and the model answers without that block's reasoning. | No error and no warning (the request is successful). With the Claude API and Google Cloud, if no header is sent, it is silent; if a header is sent, it is reported only on request (organization_binding_mismatch). | Claude Sonnet 5.5 (2026-09-28), Claude Haiku 5.5 (2026-10-07) | Preserved thinking, What's new in Claude Sonnet 5.5 |
| Claude API prefix check (accounts created on or after August 31, 2026, 00:00 UTC) | Rejected (400). If the beta header is sent and prefix_mismatch_behavior is set to "drop_block", the blocks are Dropped (with Claude Sonnet 5.5 and Claude Haiku 5.5, this can only be specified when thinking is adaptive). | Error (bound to a different conversation). With "drop_block", Reported only on request (prefix_binding_mismatch). | Claude Fable 5.1, Claude Opus 5.5, Claude Sonnet 5.5, Claude Haiku 5.5 | Preserved thinking |
| Claude API prefix check (accounts created before that date) | Carried. Unless the request sets prefix_mismatch_behavior, blocks that fail the check still reach the model. | Reported only on request (if a header is sent, thinking_mismatch_allowed). | Claude Fable 5.1, Claude Opus 5.5, Claude Sonnet 5.5, Claude Haiku 5.5 | Preserved thinking |
| Amazon Bedrock Converse (a request that changes an earlier message in a conversation with signed reasoning) | Rejected. The user guide states that modifying any message will result in an error, but no status code is provided. | Error (no status code or message given). | checked 2026-10-09 | Inference using Converse API |
| OpenAI API (requests that change the model family) | Dropped. Incompatible reasoning is removed from the model's context. This is also true when reasoning.context is all_turns. | The source does not say. (The response's reasoning.context shows the effective mode.) | checked 2026-10-09 | Reasoning models |
Anthropic's documentation describes these checks as characteristics of the Claude models, and separately details platform-specific differences (whose blocks a model reads, and whether a drop appears in
input_transformations). The rows in the table for the Claude API come from descriptions specific to the Claude API.Regarding the model check, the "Preserved thinking" page states that when the current model cannot read a block, the API drops it from that request without error.
If the current model can't read a block, the API drops it from that request without an error.
For the account check, the same page says:
Thinking blocks that Claude Sonnet 5.5 or Claude Haiku 5.5 produces work only in the account that produced them, or in an account linked to it. When another account sends one of these blocks, the API drops the block before the model sees it, and the request succeeds. The model answers without the reasoning in the dropped blocks. Blocks from earlier models aren't affected.
In both the Claude API and Google Cloud, dropped blocks only appear in the response's
input_transformations when a header is sent. Otherwise, the fact that blocks were dropped is not indicated.On the Claude API and Google Cloud, with the `thinking-binding-controls-2026-08-01` beta header, the response lists each dropped block in `input_transformations` as a `thinking_dropped` entry with `reason: "organization_binding_mismatch"`. Without the header, the drop is silent.
For the accounts on which the prefix check is enforced by default, the same page says:
The model check applies to every account. The API enforces the prefix check by default for accounts created on or after August 31, 2026, 00:00 UTC.
The "What's new in Claude Sonnet 5.5" page says that the prefix check for Claude Sonnet 5.5 is enforced by default for accounts created on or after August 31, 2026, 00:00 UTC, across the Claude API, Amazon Bedrock, and Google Cloud.
Regarding configurations where the account check is effective, the Claude Haiku 5.5 migration guide names code that saves conversations and resends them from a different account, citing, for example, a service that serves multiple customers from a single conversation store.
This affects code that stores conversations and replays them through a different account, for example a service that serves several customers from one conversation store. Replay each conversation through the account that produced it.
The OpenAI "Reasoning models" guide states that when switching model families, the API omits incompatible reasoning from the model's context.
When you switch model families, the API omits incompatible reasoning from the model's context, even when `reasoning.context` is `all_turns`.
5.2 A Dropped Block Also Moves the Prompt Cache Prefix
Anthropic's prompt caching documentation says that when the API drops a thinking block, the cached prefix for that request changes from that block's position onward.When the API drops a Claude Fable 5.1, Claude Mythos 5.1, Claude Opus 5.5, Claude Sonnet 5.5, or Claude Haiku 5.5 thinking block that isn't preserved on that request (for example, one you replay to a model that can't read it), the cached prefix changes from that block's position onward on that request.
The occurrence of missing thinking blocks and cache invalidation can appear to be separate issues, but they sometimes stem from the same underlying cause. Refer to Section 8 for methods to identify these issues.
5.3 What Anthropic's OpenAI SDK Compatibility Layer Ignores
Anthropic provides a compatibility layer that lets the OpenAI SDK call Claude. The OpenAI SDK compatibility page describes this layer as primarily intended for testing and comparing model capabilities, and not as a long-term or production-ready solution for most use cases.This compatibility layer is primarily intended to test and compare model capabilities, and is not considered a long-term or production-ready solution for most use cases.
The same page lists features that are not supported, noting that most unsupported fields are silently ignored rather than generating errors.
Most unsupported fields are silently ignored rather than producing errors.
Regarding this article's six kinds of information and related aspects, the page further states: The OpenAI SDK does not return Claude's thinking. It does not support prompt caching (Anthropic's SDK supports it). The
strict parameter for function calling is ignored. System and developer messages, regardless of their position within the conversation, are combined into the top-level system field (checked 2026-10-10). The page also lists the reasoning_effort parameter in requests among the ignored fields.| Decided By | What Happens | What the Caller Sees | Where the Source Says So |
|---|---|---|---|
| Anthropic OpenAI SDK Compatibility Layer (regarding thinking in the reply) | Dropped. The docs say that the OpenAI SDK does not return Claude's thinking. | The source does not say. | OpenAI SDK compatibility |
| Anthropic OpenAI SDK Compatibility Layer (regarding most unsupported fields) | Dropped (ignored). | No error and no warning (silently ignored rather than producing errors). | OpenAI SDK compatibility |
| Anthropic OpenAI SDK Compatibility Layer (regarding system and developer messages) | Converted. Combined into the top-level system field, regardless of their position within the conversation (checked 2026-10-10). | — | OpenAI SDK compatibility |
This table, too, indicates that
Hop is Provider in every row, and Status and Version refers to the documentation as of the verification date (checked 2026-10-09). However, the row regarding system and developer messages was written based on the page as of October 10, 2026, as the page was updated after this article's check. A table in Section 7 lists the lack of support for prompt caching.6. What Is Lost When a Client, a Library, or a Pass-Through Gateway Rewrites the Request
This section addresses scenarios where data loss occurs as a result of modifying the history, even without changing the data format. This includes modifications by clients and their SDKs, as well as by pass-through gateways that simply relay information without altering its format.6.1 Two Ways a Client Silently Drops Information
Anthropic's "Thinking" page describes code that filters reply content using types. It notes that when round-tripping replies that use tools, filtering solely by thethinking type silently drops redacted_thinking blocks, disrupting the multi-turn protocol.If your code filters content blocks by type (for example, `block.type == "thinking"`) when round-tripping responses with tool use, also include `redacted_thinking` blocks. Filtering on `block.type == "thinking"` alone silently drops `redacted_thinking` blocks and breaks the multi-turn protocol described in Preserving thinking blocks.
The Troubleshooting thinking page from the same provider describes a situation where a request to return tool results fails with a
400 error, and the message includes the following text.`thinking` or `redacted_thinking` blocks in the latest assistant message cannot be modified
The same page first lists as the cause the code that drops
redacted_thinking content based on type filtering.This error happens when the assistant message you send back differs from the one the API returned, most often because your code filters content blocks by type and drops `redacted_thinking` blocks, or rebuilds the assistant message instead of echoing it.
In essence, the drop originates in the client's code, and no immediate error is reported. However, if the dropped block is in the most recent assistant message, the API will reject the next request with a
400 error.Another scenario, highlighted in Section 4.6, concerns the "Chat completions" page for agentgateway. Clients that retain
reasoning_content but discard reasoning_signature lose the history of thinking in the next turn, without any error or warning. (As noted in Section 4.6, the gateway sometimes does not send the signature itself.) Both of these issues can occur solely due to the client's code, even when there are no problems with the gateway or the provider.6.2 Rewrites Count as Edits — The Preserved Thinking Docs' Note
Anthropic's "Preserved thinking" page includes a section dedicated to libraries, proxies, and gateways. It states that a library, proxy, or gateway sits between someone else's history and the API, so its rewriting counts as an edit, and that its users cannot see or fix these rewrites.A library, proxy, or gateway sits between someone else's history and the API, so its own rewrites count as edits, and its users can't see or fix them.
That section of the page lists four points.
- Pass through what you don't recognize. Forward the caller's `anthropic-beta` values and `thinking.block_binding` unchanged, and return `input_transformations` to them. An options schema that rejects unknown keys stops your users from choosing `"drop_block"`.
- Leave a `role: "system"` message where the caller put it. Moving it into the top-level `system` field changes `system` on that request and invalidates every thinking block in the conversation.
- To turn tool use off for a request, send `tool_choice: {"type": "none"}`. Don't remove `tools`.
- Don't hide the 400. If your code catches it, strips thinking, and retries on the caller's behalf, log that it did: their history is still edited, and the model loses its earlier reasoning on every later request.
This is a note directed at those creating libraries and gateways. It does not describe specific instances where a particular client or gateway has done so. When rewriting occurs, the outcome follows the process outlined in Section 5.1 regarding prefix checks. For new accounts, a
400 error is returned, and with the "drop_block" designation, the blocks are dropped. For older accounts, the block will reach the model unless prefix_mismatch_behavior is specified.6.3 Gateways That Strip Headers or cache_control
The Claude Code gateway compatibility guide in Claude Code Docs outlines what a gateway should pass through when using Claude Code behind an ANTHROPIC_BASE_URL gateway. It states that features paired with beta headers and request body fields will return a 400 error if only the headers are stripped, and only turn off (referred to in the documentation as "quietly") when both are missing.A gateway that strips the header while passing the body, or forwards an Anthropic-format body to an upstream with a different schema, produces hard `400` errors; only when both halves are absent together does the feature turn off quietly.
In the Feature pass-through table on the same page, the rows that concern this article's six kinds of information are the interleaved thinking row and the prompt caching row. (There is also a row regarding a
400 error when the upstream system doesn't accept the thinking setting field, but this relates to parameters.) Features like interleaved thinking, which rely solely on beta headers and do not include body fields, will silently cease to function if the headers are stripped.Silently unavailable when the header is stripped; the upstream never sees the capability request
Regarding the prompt cache row, the "Symptom when broken" column begins with
No error, and describes the behavior as follows. Its remediation cell says to pass through cache_control without modification, regardless of its location, and not to convert block-form system or message content into simple strings.visible as high `input_tokens` with little or no cache activity in `usage`
6.4 A Client That Removes Thinking and Retries After a 400
The same guide also describes what Claude Code does when the upstream system rejects a request. If the thinking signature is rejected, Claude Code removes earlier thinking blocks from the request, retries, and keeps them out of every later request.When the upstream rejects a thinking signature, including with a `400` whose message says the block is `bound to a different conversation`, Claude Code removes earlier thinking blocks from the request, retries, and keeps them out of every later request. New responses still include thinking
The guide continues, stating that rejections related to
bound to a different conversation originate from the preserved thinking check. It also notes that gateways that rewrite system, tools, or previous messages can be the cause of these rejections. Combining these two points from the same guide reveals that gateway rewrites can trigger a provider's 400 error, and the client's resend turns it into the loss of thinking on every later request. The client is the recipient of the 400 error. What is subsequently displayed to the user is not clear from this guide. The "Don't hide the 400" item quoted in Section 6.2 asks libraries, proxies, and gateways to log instances where they perform the same recovery action (removing thinking after receiving a 400 and resending the request on behalf of the caller).6.5 Patterns Visible in the Claude Code CHANGELOG
The changelog for Claude Code lists bug fixes where information was dropped or rejected. Each entry represents a correction, stating that the issue was resolved in that specific version. Here are three examples from different versions. 2.1.281 and 2.1.282 name a proxy or gateway. The web search results in 2.1.282 are outside this article's six kinds of information. 2.1.287 is the client's own processing (a model switch), with no gateway involved.- 2.1.281 (2026-09-23): Fixed an issue where, when a proxy or gateway properly closed a stream, incomplete responses were incorrectly displayed as complete without warning.
Fixed responses cut short by a proxy or gateway that closes the stream cleanly being shown as complete with no warning, and tool calls running twice on duplicated stream events
- 2.1.282 (2026-09-24): Resolved an issue where conversations whose history holds web search results that the API could not decrypt resulted in all requests returning a 400 error. An example of this involved turns answered through a third-party gateway.
Fixed every request failing with a 400 error in conversations whose history holds web search results the API cannot decrypt (for example, from a turn answered through a third-party gateway)
- 2.1.287 (2026-10-01): Fixed an issue related to switching between Opus 5.5 and Sonnet 5.5, where previous MCP tool announcements were rewritten, which could drop earlier extended thinking.
Fixed switching between Opus 5.5 and Sonnet 5.5 (`/model`, `opusplan`) rewriting earlier MCP tool announcements, which could drop earlier extended thinking
6.6 Results for Clients and Pass-Through Gateways
Sections 6.1 to 6.4 are laid out in the common columns.| Hop | Decided By | What Happens | What the Caller Sees | Status and Version | Where the Source Says So |
|---|---|---|---|---|---|
| Client | Code that filters reply blocks on block.type == "thinking" alone when round-tripping replies with tool use (the Claude API's check decides the result) | Dropped (redacted_thinking). If the dropped block is part of the latest assistant message, the next request is Rejected (400). | No error and no warning (at the point of dropping; the documentation states it silently drops). If it's part of the latest assistant message, the next request results in an Error (400, cannot be modified). | checked 2026-10-09 | Thinking, Troubleshooting thinking |
| Client | A Chat Completions client that discards reasoning_signature (agentgateway's reverse route) | Dropped (previous thinking history for the next turn). | No error and no warning. | agentgateway 1.6.0 (2026-10-02) | agentgateway docs (Chat completions) |
| Client | Claude Code after a 400 on a signature | Dropped (earlier thinking blocks). Removes them, retries, and keeps them out of every later request. | The source does not say. (Claude Code receives the 400, and what the user sees is not stated.) | checked 2026-10-09 | Claude Code gateway compatibility guide |
| Gateway | Proxy or gateway that modifies system, tools, or previous messages (a library belongs to the Client hop; the Claude API's prefix check decides the result) | Rejected (for new accounts, it's a 400). With "drop_block", subsequent thinking blocks are Dropped. For older accounts, it is Carried by default (Section 5.1). | Error (for new accounts, it's a 400, bound to a different conversation). With "drop_block", it is Reported only on request (prefix_binding_mismatch). | checked 2026-10-09 | Preserved thinking, Claude Code gateway compatibility guide |
| Gateway | Gateway that moves role: "system" messages to the top-level system. The Claude API's prefix check decides the result. | Rejected or Dropped (every thinking block in the conversation becomes invalid, with the result of the prefix check in Section 5.1) | Error (a 400 on new accounts; see the prefix check rows in Section 5.1) | checked 2026-10-09 | Preserved thinking |
| Gateway | Gateway that strips only beta headers (header-only functionality). | Dropped (requests for features like interleaved thinking). | No error and no warning (Silently unavailable). | checked 2026-10-09 | Claude Code gateway compatibility guide |
| Gateway | Gateway that strips beta headers and passes through the body fields. | Rejected (400). | Error (400). The message varies depending on the feature. | checked 2026-10-09 | Claude Code gateway compatibility guide |
| Gateway | A gateway that does not forward cache_control unchanged, or that converts block-form system or message content into strings | Dropped (cache breakpoints) | No error and no warning (the guide's "Symptom when broken" column starts with No error). High input_tokens in usage and little cache activity. | checked 2026-10-09 | Claude Code gateway compatibility guide |
7. Combinations the Target Format Does Not Support
This section lists combinations where, according to the sources, the target format is not compatible with the model or endpoint. The documentation for agentgateway says that, to keep the history of thinking, a provider should declarecompletions without responses (see Section 4.1). Conversely, the OpenAI documentation notes that some newer models may not support tool calls with Chat Completions. Section 3.3 of Amazon Bedrock Inference Endpoints and API Keys details the differences in the bedrock-runtime Responses API.| Decided By | Combination | What the Source Says | Where the Source Says So |
|---|---|---|---|
| OpenAI API (GPT-5.4 and later) | Chat Completions, tool calls, reasoning_effort is not none | Not supported | Migrate to the Responses API |
| OpenAI API (GPT-6 Astra, GPT-6.1 Sol) | Chat Completions, tool calls | Not supported. Responses is required for tool calls. | Using GPT-6, Reasoning models |
| OpenAI API (GPT-6 Sol, GPT-6 Luna) | Chat Completions, function calls | Supported only with reasoning_effort: "none". | Using GPT-6 |
| OpenAI API (GPT-6 Astra) | reasoning_effort is none | Returns an HTTP 400 error. | Reasoning models |
| agentgateway 1.6.0 (converting to Chat Completions) | Certain models (e.g., gpt-5.3), requests with tools | Sends the request with reasoning_effort: "none" and drops the thinking that was asked for. | agentgateway docs (Messages) |
Amazon Bedrock's bedrock-runtime Responses API | background=true | Rejected with a 400 error. Requests are always synchronous. | Responses API (Bedrock user guide) |
Amazon Bedrock's bedrock-runtime Responses API | Server-side tools (including web search) | Not supported. Client-side tools are supported. | Responses API (Bedrock user guide) |
| Amazon Bedrock's Converse (GPT-6 Sol, GPT-6.1 Sol, GPT-6 Luna) | Explicit cache breakpoints (cachePoint) | Not supported. Converse (bedrock-runtime) only supports implicit caching. Responses and Chat Completions support explicit caching with bedrock-runtime and bedrock-mantle, and InvokeModel supports it with bedrock-runtime. | Model cards |
| Anthropic's OpenAI SDK compatibility layer | Prompt caching | Not supported | OpenAI SDK compatibility |
OpenAI has two descriptions listed below:
Starting with GPT-5.4, Chat Completions does not support tool calling with `reasoning_effort` values other than `none`.
GPT-6 Astra and GPT-6.1 Sol support Chat Completions, but tool calling requires Responses. GPT-6 Sol and GPT-6 Luna support function calling in Chat Completions only with `reasoning_effort: "none"`. Use Responses for reasoning with tools.
The Bedrock model cards for GPT-6 Sol, GPT-6.1 Sol, and GPT-6 Luna describe Converse caching in the same words.
Converse supports implicit caching on the `bedrock-runtime` endpoint. The native `cachePoint` field isn't supported for explicit caching with this model.
OpenAI's documentation refers to GPT-5.4 and later, while the agentgateway documentation provides an example using
gpt-5.3. These statements are about different models, and this article does not treat either as wrong.Combining the facts presented in this section leads one to want to extrapolate how agentgateway's conversion to Chat Completions functions when sending tool-enabled conversations to GPT-6 models. However, the agentgateway documentation only provides examples using
gpt-5.3, and does not describe behavior with GPT-6 models. This article will not document the combined results. It will also not document the combination of the fourth row in the table (where sending none to GPT-6 Astra results in a 400 error) and the fifth row (where agentgateway sends none to some models). The agentgateway documentation does not list any models besides gpt-5.3 that are subject to the behavior described in the fifth row. What can be stated is that OpenAI's documentation indicates that GPT-6 Astra and GPT-6.1 Sol require Responses for tool calls, while the agentgateway documentation states that its conversion to Responses drops the thinking history. These are separate descriptions from different sources.8. Ways to Notice That Something Was Dropped
This section lists, in the scope of the documentation, methods for observing instances where something was dropped. It does not state that there are no methods for detecting such drops. Where no method was found, this section says so and gives the search scope. Section 7.4 of Mid-Conversation Changes in the Claude API discusses the Claude API's cache diagnostics and the use ofinput_transformations. When a request passes through a gateway, the "Preserved thinking" page asks the gateway to return input_transformations to the caller (Section 6.2).| Mechanism | What It Shows | Condition | Where the Source Says So |
|---|---|---|---|
Claude API response's input_transformations | The position and reason for each dropped thinking block (model_binding_mismatch, organization_binding_mismatch, prefix_binding_mismatch), and blocks that reached the model despite failing the check (thinking_mismatch_allowed) | When sending the beta header thinking-binding-controls-2026-08-01. organization_binding_mismatch applies to both the Claude API and Google Cloud. | Preserved thinking |
Claude API response's usage | Token count for cache reads and writes. The Claude Code guide states that when cache_control is dropped, it is indicated by a high input_tokens value and little cache activity. | In every response. | Prompt caching (Anthropic), Claude Code gateway compatibility guide |
| OpenAI Prompt Cache Diagnostics | Cache reuse compared with a previous response, and the reasons for cache misses | Responses API, supported models from GPT-5.6 on. Made generally available on September 8, 2026. | Changelog |
| A way for agentgateway to report what a conversion dropped | No statement found in the range checked | — | agentgateway docs, release notes for version 1.6.0 |
The entries for agentgateway were the result of searching the documentation (docs) pages for "Messages," "Chat completions," "Amazon Bedrock," "Custom providers," and "Anthropic," as well as the release notes for version 1.6.0, using the terms
warn, log, header, notify, and metric. No descriptions were found indicating that what a conversion dropped is reported to the calling code or logs. This does not mean that agentgateway lacks any means of notification.Regarding OpenAI's response
reasoning.context, it indicates the effective mode (either current_turn or all_turns). The guide for Reasoning models states that when switching model families, the API omits any non-compatible reasoning even with all_turns, and it does not explicitly state that reasoning.context is used to indicate that such reasoning has been omitted. This article does not consider this a means of awareness.Concerning OpenAI's Prompt Cache Diagnostics, the changelog entry dated September 8, 2026, reads:
Prompt Cache Diagnostics is now generally available in the Responses API for GPT-5.6 and later supported models.
Claude Code has gateway hint headers (enabled by setting
CLAUDE_CODE_GATEWAY_HINT_HEADERS=1; version 2.1.283, released on September 25, 2026, added x-claude-code-prompt-id to them). The CHANGELOG describes this as a feature that allows the gateway to aggregate requests serving a single user's prompt. However, it does not mention any mechanism for reporting drops, so this article does not consider it a means of awareness.9. Where the Sources Disagree and Where No Statement Was Found
This section lists the areas where the documents disagree and the areas where no corresponding information was found. This article does not assume that any of these discrepancies are due to errors in any of the documents.9.1 agentgateway Citations — The Release Notes and the Docs
The new features section in the 1.6.0 release notes states that Messages conversion has become more complete, and that citations are now carried over during conversions between Anthropic Messages and OpenAI formats.More complete Messages conversion: Conversion between Anthropic Messages and OpenAI formats now carries citations, refusals, strict tool schemas, reasoning effort, and images in tool results. Provider context-overflow errors are translated so Claude Code can compact and retry.
The latest documentation states that during conversion to Responses, citations on the
text blocks of the history are dropped (Section 4.3). On the reply side, the URL citations of a buffered reply are maintained as web_search_result_location. These two descriptions have different scopes. The release notes do not specify the direction (history or reply) or the route (Responses or Chat Completions) for carrying citations. The documentation, however, details how, for the Responses route, history citations are dropped while the URL citations from a buffered reply are converted, specifying this on a per-direction basis. There is no mention in the documentation regarding citations for the Chat Completions route. The table in Section 4.2 of this article follows the documentation, listing the information by route and direction, and including only the general rule from the release notes for the Chat Completions column.9.2 agentgateway's Route to Bedrock — The Messages Page and the Amazon Bedrock Page
The Messages page says that a Bedrock provider supports only Converse and that Messages requests are always converted to Converse.A `bedrock` provider supports only Converse, so agentgateway always converts Messages requests to that format.
The Amazon Bedrock page says that agentgateway chooses, per model, whether to send a request to Bedrock's Runtime or Mantle endpoint, from the model's tags and the
bedrockEndpointPreference setting. Runtime supports the Converse and Invoke APIs, while Mantle supports the native APIs of OpenAI and Anthropic. The default, runtimePreferred, uses Runtime for all models except those with only the mantle tag. The 1.6.0 release notes also mention Bedrock Mantle routing as a new feature. The Amazon Bedrock page adds that a model that resolves to Mantle accepts the formats in its tags, except anthropic.claude* models, which always take the Anthropic Messages format. The two pages differ in how they describe the format of Messages requests routed to Mantle. The section on routing to Converse (4.5 of this article) is a description of requests sent to Runtime.9.3 Descriptions of the Signature
As mentioned in Section 3.2, the description ofsignature varies across three sources: Anthropic's Thinking page (an encrypted copy of the full reasoning), the Bedrock user guide (a hash of all messages in a conversation), and the Bedrock API reference (a token to verify that the text was generated by the model). This article will continue to list these three descriptions separately.9.4 agentgateway's Provider for an OpenAI-Compatible Server — The Messages, Custom Providers, and Claude Code Pages
The latest Messages page says to use theopenAI provider only for the OpenAI API and, for other OpenAI-compatible servers, to use a custom provider with the formats that the server supports (Section 4.1). The Custom providers page says that OpenAI-compatible endpoints can be used with provider: openai and a customized baseUrl, but it advises preferring provider: custom. The agentgateway page for Claude Code gives an example configuration that sends Claude Code's requests to an OpenAI-compatible server such as vLLM, and the example uses provider: openAI with a local baseUrl. This article does not treat any of these pages as wrong.9.5 Where No Statement Was Found
No descriptions were found for the following items in the range checked. The main points follow. In each cell of the tables,The source does not say. indicates that the relevant documentation was searched without finding a corresponding description. This does not mean that the information is absent or unsupported.- Regarding the handling of tool call IDs and citations during the conversions to Chat Completions and to Converse, and of document, search-result, and server-tool blocks during the conversion to Chat Completions, in the agentgateway documentation (Section 4.2 documents the search terms and scope). What a cache breakpoint becomes after the conversion was not described in the documentation for any of the three routes (for the Responses route, the documentation only lists it as within scope); this article supplemented this information by examining the source code for version 1.6.0.
- A way for agentgateway to tell the caller what a conversion dropped (Section 8).
- The handling of citations within Amazon Bedrock's OpenAI-compatible endpoints (Responses and Chat Completions). A search of the Bedrock user guide's Responses API page and Chat Completions API page for
citationandannotationfound no statement. - Specific fields in the OpenAI Chat Completions API reference, namely fields for reasoning text and encrypted reasoning (Section 3.1).
Regarding OpenAI's announcement from September 22, 2026, "Better prompt caching for GPT-6," the content could not be retrieved (received a 403 error when attempting to access it via curl). The information in this article regarding OpenAI's caching practices relies on the "Prompt caching" guide and the changelog. The changelog entry from July 9, 2026, states that explicit prompt caching control was added in GPT-5.6, and the "Prompt caching" guide says that explicit cache breakpoints are supported from GPT-5.6 on.
10. Frequently Asked Questions about Format Conversion in LLM Gateways
This article addresses common questions that arise when switching models behind the LLM gateway, or when the gateway is introduced and the behavior of the conversation changes.Q1. Does thinking disappear when a request passes through an LLM gateway?
The results vary depending on the route and version. According to the documentation for agentgateway 1.6.0, the route from Messages to Responses drops the thinking history, while the routes from Messages to Chat Completions and from Messages to Converse convert it into a different field and carry it. In the documentation for 1.5.x, requests to providers that only declareresponses will receive a 400 error when including the thinking history, while requests to providers declaring both formats will go through the route to Chat Completions. Even with gateways that pass data through without changing the format, if the system or tools fields are altered, on Claude models that perform prefix checks, the provider's check may reject or drop the thinking history on new accounts (see Sections 5.1 and 6.2).Q2. If Chat Completions is declared in agentgateway, is thinking history always preserved?
No. Requests to providers that also declareresponses go through the route to Responses in 1.6.0 (see Section 4.1). The documentation indicates that the ability to carry over the thinking history on the Chat Completions route depends on the self-hosted engine that returns the reasoning as reasoning_content and also accepts it. If a turn contains two or more signed blocks, the signature will not be sent. redacted_thinking is dropped. In tool-enabled requests for certain models, such as gpt-5.3, the requested thinking is dropped. OpenAI's documentation specifies that GPT-6 Astra and GPT-6.1 Sol's tool calls require Responses (see Section 7).Q3. If there is no error, can you assume nothing was dropped?
No. The documentation states that agentgateway's route to Responses, the Claude API's model and account checks, and Anthropic's OpenAI SDK compatibility layer (for most unsupported fields) all drop information without producing errors. In the Claude API, by sending a beta header, the dropped blocks will appear in the response'sinput_transformations (see Section 8).Q4. The prompt cache stopped working after a gateway was added. What should you suspect?
Based on the information in this article, there are two potential causes: dropped cache breakpoints and dropped thinking blocks. Claude Code's documentation describes the symptoms of a gateway that is droppingcache_control as exhibiting no errors, high input_tokens usage, and little cache activity. Anthropic's Prompt caching page states that when the API drops a thinking block from one of the models listed in Section 5.2, the prefix of the cache for that block and subsequent blocks changes (Sections 5.2 and 6.3).Q5. Returning a tool result through the Responses conversion got a 400. What is the cause?
Based on the information in this article, it's not possible to pinpoint a single cause. The commit message in pull request #3491 for agentgateway states that by maintaining thetool_use.id from Messages as the call_id in Responses, and avoiding assigning the optional item id, it's possible to avoid the HTTP 400 error that occurs on replay. With OpenAI Responses, tool calls have separate id and call_id fields, and the call_id links each result to its call (see Sections 3.3 and 4.3). The 1.5.x documentation also states that, in the conversion to Responses, requests containing thinking history or tool results that are not text fail with a 400 (see Section 4.7).Q6. Does converting to Bedrock Converse carry all encrypted reasoning?
No. According to the agentgateway documentation, the route to Converse converts the history'sredacted_thinking to redactedContent, while maintaining the encrypted reasoning from a buffered reply as redacted_thinking. However, in streamed replies, the content is not preserved; it arrives as a thinking block with [REDACTED] and cannot be replayed. The Chat Completions client does not receive the encrypted reasoning (see Section 4.5).Q7. Can Claude Code's gateway hint headers detect what was dropped?
The sources do not describe them as a way to detect drops. The CHANGELOG mentions thatx-claude-code-prompt-id was added in version 2.1.283 (2026-09-25), and describes it as a mechanism for the gateway to consolidate requests that pertain to a single prompt. The method for identifying dropped blocks in the Claude API is through the input_transformations included when sending beta headers (Section 8).11. Summary
In the conversion between LLM API formats, information not included in the parameter mapping table is handled differently depending on the route. The documentation for agentgateway 1.6.0 (October 2, 2026) indicates that the route from Messages to Responses silently drops historical thinking andredacted_thinking, historical citations, and blocks such as document blocks, without generating any errors or warnings. The route from Messages to Chat Completions converts thinking into reasoning_content before transmitting it. However, the signature is not sent when a single turn contains two or more signed thinking blocks, and redacted_thinking is always dropped (Section 4.4). The route from Messages to Converse converts signed thinking and encrypted reasoning. However, encrypted reasoning in streamed replies is dropped.Two things concerning the thinking history changed between versions 1.5 and 1.6.0. When sending historical thinking to providers that only declare
responses, the 1.5.x documentation says the request is rejected with a 400 error, and the 1.6.0 documentation says the history is silently dropped. For providers that declare both formats (such as OpenAI), requests from Messages went through the Chat Completions route in version 1.5, and go through the Responses route by default in version 1.6.0, where the historical thinking is silently dropped. The body of a pull request listed in the 1.6.0 release notes states that, before that pull request's change, the conversion to Chat Completions also dropped thinking blocks (Section 4.7).Beyond the conversion gateway, there are also hops where data is dropped. The Claude API drops thinking blocks that the model cannot read, as well as Claude Sonnet 5.5 and Claude Haiku 5.5 thinking blocks sent by accounts other than the account that created them (and linked accounts), while still reporting a successful request. For Claude Fable 5.1, Claude Opus 5.5, Claude Sonnet 5.5, and Claude Haiku 5.5, on accounts created on or after August 31, 2026, at 00:00 UTC, historical thinking with a changed prefix is rejected with a
400 error by default. Clients may also drop redacted_thinking by filtering blocks by type, or discard reasoning_signature on agentgateway's reverse route. Even pass-through gateways can lose thinking or prompt-cache hits by rewriting system or tools, or by stripping beta headers or cache_control.Some combinations are not supported in the target format. OpenAI's documentation indicates that GPT-6 Astra and GPT-6.1 Sol require Responses for tool calls. Bedrock's Converse does not support explicit cache breakpoints (
cachePoint) with GPT-6 Sol, GPT-6.1 Sol, and GPT-6 Luna.Finally, here are five things to verify in your environment:
- Does the path from the client to the model include a gateway that converts the format? If so, what version of that product runs, and which routes (source and target formats) does it traverse?
- Along the conversion route, are the thinking history, signatures,
redacted_thinking, cache breakpoints, tool call IDs, and citations being carried, converted, or dropped? Do the gateway's documentation and release notes contain version-specific information? - Are any gateways or libraries that don't change the format inadvertently rewriting or stripping out
system,tools, previousmessages,anthropic-beta, orcache_control? - Are you re-sending stored conversations from an account different from the one that created them?
- If using the Claude API, are you checking for dropped data using the
thinking-binding-controls-2026-08-01header andinput_transformations? If using OpenAI, are you checking the Prompt Cache Diagnostics? In either case, are you monitoring theusagecache values to see what was dropped?
12. References
- Messages - agentgateway
- Messages (1.5.x) - agentgateway
- Chat completions - agentgateway
- Custom providers - agentgateway
- Anthropic - agentgateway
- Amazon Bedrock - agentgateway
- Claude Code - agentgateway
- Release v1.6.0 - agentgateway/agentgateway - GitHub
- fix(llm): carry thinking blocks between Messages and Chat Completions - agentgateway/agentgateway Pull Request #3363
- Fix improper messages->responses id and prepare golden files - agentgateway/agentgateway Pull Request #3491
- messages to responses: various compatibility fixes - agentgateway/agentgateway Pull Request #3496
- llm: messages to responses better cache control - agentgateway/agentgateway Pull Request #3516
- completions.rs (v1.6.0) - agentgateway/agentgateway - GitHub
- responses.rs (v1.6.0) - agentgateway/agentgateway - GitHub
- bedrock.rs (v1.6.0) - agentgateway/agentgateway - GitHub
- mod.rs (v1.6.0) - agentgateway/agentgateway - GitHub
- Thinking - Claude Platform Docs
- Troubleshooting thinking - Claude Platform Docs
- Preserved thinking - Claude Platform Docs
- Prompt caching - Claude Platform Docs
- Citations - Claude Platform Docs
- Streaming messages - Claude Platform Docs
- Handle tool calls - Claude Platform Docs
- OpenAI SDK compatibility - Claude Platform Docs
- What's new in Claude Sonnet 5.5 - Claude Platform Docs
- What's new in Claude Haiku 5.5 - Claude Platform Docs
- Claude Haiku 5.5 migration guide - Claude Platform Docs
- Claude Platform release notes - Claude Platform Docs
- Other LLM gateways - Claude Code Docs
- Claude Code gateway compatibility guide - Claude Code Docs
- Claude Code CHANGELOG - anthropics/claude-code - GitHub
- Releases - anthropics/claude-code - GitHub
- Migrate to the Responses API - OpenAI API
- Reasoning models - OpenAI API
- Using GPT-6 - OpenAI API
- Prompt caching - OpenAI API
- Function calling - OpenAI API
- Streaming API responses - OpenAI API
- Changelog - OpenAI API
- Chat - OpenAI API Reference
- Create a model response - OpenAI API Reference
- Inference using Converse API - Amazon Bedrock
- Prompt caching for faster model inference - Amazon Bedrock
- Responses API - Amazon Bedrock
- Chat Completions API - Amazon Bedrock
- GPT-6 Sol - Amazon Bedrock
- GPT-6.1 Sol - Amazon Bedrock
- GPT-6 Luna - Amazon Bedrock
- ContentBlock - Amazon Bedrock API Reference
- ContentBlockDelta - Amazon Bedrock API Reference
- ReasoningTextBlock - Amazon Bedrock API Reference
- ToolUseBlock - Amazon Bedrock API Reference
- ToolResultBlock - Amazon Bedrock API Reference
References:
Tech Blog with curated related content
Written by Hidekazu Konishi