The result: Codex compaction fails because the compaction request replays a web_search_call from history while sending tools: []. Since around October 6, 2026, ChatGPT's Codex backend rejects that combination with response protection is unavailable. Declaring web_search (with tool_choice: "none") makes the same request complete. Official Codex doesn't have a fix yet, so the client-side workaround is web_search = "disabled" for new sessions. If you maintain your own gateway, it can add the declaration for you.
The engineering problem: a request body that was valid last week is now rejected, and nothing in it changed. The history still contains the search item. The compaction request still declares no tools, because a summary doesn't need any. What changed is that the upstream now checks whether a replayed web search comes with a declared web_search tool. One web search early in a long session is enough to make every later compaction fail, whether the request goes through a proxy or through the official Codex app on a direct ChatGPT login.
This post covers how to confirm you're hitting this failure, the failing and passing request shapes, a minimal repro, the gateway patch, and the client-side workaround.
Confirm it's this failure
"stream disconnected before completion" is a generic Codex message. Overloaded upstreams and WebSocket idle timeouts produce it too. This failure is specifically one of these:
stream disconnected before completion: response protection is unavailable
Error running remote compact task: stream disconnected before completion: stream closed before response.completed
The second line can be the same failure. Codex doesn't always surface the upstream SSE error and falls back to the generic message. To tell them apart, check three things.
1. Codex's local log. On affected machines, ~/.codex/logs_2.sqlite repeats remote compaction v2 stream failed and Failed to run pre-sampling compact. Codex retries about five times, then gives up. Resuming the session triggers the same compaction, so the session stays stuck.
2. What the upstream returned. If you run a gateway, look for one of two forms. The first is an HTTP 502 with this body, which a ChatGPT Plus user on gpt-6.1-sol also posted to OpenAI's developer forum:
{"message": "response protection is unavailable", "type": "internal_error"}
The second is an HTTP 200 whose SSE stream ends in response.failed with code: "upstream_error", or in an event: error that nests the text in error.message, and never reaches response.completed. In one gateway's logs, failed attempts took 2 to 13 seconds and recorded zero tokens.
3. Whether the history contains a search. Normal turns still work. Only requests that send the history without declaring tools fail. Codex keeps session history under ~/.codex/sessions/, so a quick check is:
grep -rl 'web_search_call' ~/.codex/sessions/
If the stuck session's file shows up, its history contains a hosted search item. Codex's web search defaults to "cached", so a search can be in there even if you never turned search on.
The request shapes
Here is a trimmed compaction request of the kind testers captured: POST /v1/responses, streaming, store: false, a trailing compaction_trigger (remote compaction v2), and an empty tools array. The values are placeholders.
{
"model": "gpt-6.1-sol",
"stream": true,
"store": false,
"tools": [],
"input": [
{"type": "message", "role": "user", "content": [{"type": "input_text", "text": "..."}]},
{"type": "web_search_call", "id": "ws_...", "status": "completed", "action": {"type": "search", "query": "..."}},
{"type": "message", "role": "assistant", "content": [{"type": "output_text", "text": "..."}]},
{"type": "compaction_trigger"}
]
}
For custom providers, Codex compacts locally instead (codex-rs/core/src/compact.rs). That request is a plain /responses call with the full history, tools: [] and no compaction_trigger. It fails the same way.
Normal turns don't fail because they declare {"type": "web_search"}. The passing version of the request above only changes two fields:
{
"tools": [{"type": "web_search", "external_web_access": false}],
"tool_choice": "none"
}
external_web_access: false limits the tool to the cached index; the declaration only has to satisfy the check. tool_choice: "none" means the model can't call it. In one test, the declared version completed in 2.63 seconds with zero new search calls.
Several groups ran A/B tests independently, posted on the openai/codex tracker and in the issue trackers of several open-source gateway projects. The combined results:
| History | Tools declared | Result |
|---|---|---|
| Messages and reasoning only | None | Completes |
| Function call and output | None | Completes |
| Includes a web_search_call | None | Fails |
| web_search_call, ID changed or removed | None | Fails |
| Includes a web_search_call | Function tool only | Fails |
| Includes a web_search_call | web_search | Completes |
| Search items removed | None | Completes |
Removing reasoning items didn't help in one tester's run. Request size was ruled out. The same split shows up with the official, unmodified codex-cli on a ChatGPT login: a history with three search items failed to compact, and the same history with only those items removed compacted fine. Developers who captured the raw requests sum up the rule this way: if the input replays a hosted web_search_call and no web_search* tool is declared, the upstream fails the stream.
Minimal repro
Size isn't the trigger, so you don't need a long session to see it. One tester used the same key and model for both runs:
- Start Codex with a tiny compaction threshold:
codex -c model_auto_compact_token_limit=2000. - In one session, chat without tools until compaction runs. It compacts with a 200.
- In a second session, ask something that makes the model use the built-in web search once, then keep going until compaction runs.
In that test, the second session failed compaction six times in a row with the protection error. Compaction isn't the only request type affected. One gateway also logged 66 failed thread_description requests on gpt-6-luna between October 3 and 9.
Fix it in your own gateway
The patches that work all follow the same rule. If input contains a web_search_call and no web_search* tool is declared, add a declaration. If the caller declared no tools, force tool_choice to "none". Leave input untouched.
def declare_replayed_web_search(body: dict) -> dict:
"""Let the upstream accept a replayed web_search_call in a tool-less request."""
items = body.get("input")
if not isinstance(items, list):
return body
if not any(isinstance(i, dict) and i.get("type") == "web_search_call" for i in items):
return body
tools = body.get("tools") or []
if any(isinstance(t, dict) and str(t.get("type", "")).startswith("web_search") for t in tools):
return body
caller_had_tools = bool(tools)
# Cached index only: the declaration exists to satisfy the check, not to search.
body["tools"] = tools + [{"type": "web_search", "external_web_access": False}]
if not caller_had_tools:
body["tool_choice"] = "none"
return body
Getting the function right is the easy part. The harder part is applying it in all the right places. Patches failed in these ways:
-
Detecting compaction by compaction_trigger only. Local compaction for custom providers has no
compaction_trigger. One early patch retried in plain text on the protection error, but custom-provider compactions never reached that code path. The patch that worked applies the declaration whenever the history contains a search and leaves every other request unchanged. One tester recommended running the check on every final request body, includingthread_descriptionrequests. - Patching HTTP but not WebSocket. The working fixes cover the WebSocket path as well as HTTP passthrough. WebSocket requests carry the same body, so changing the transport doesn't avoid the rule.
-
Catching only HTTP errors. When the error arrives as HTTP 200 followed by an SSE error event, a handler that checks only the status code sees a stream that ended early. One gateway's original parser also read only a top-level
message, missed the nestederror.message, and reported "compaction: the model failed" instead. -
Responses Lite. A top-level
web_searchdeclaration returns a 400 there. The working patches add the tool to the firstadditional_toolsinput item, or insert a developer-role item with it, and keep any trailingcompaction_triggerlast. -
Stripping the search items. This works, but the summary then loses whatever the search found. A softer variant turns each
web_search_callinto a plain note before compaction.
Work around it in Codex
Official Codex has no fix yet. As of October 10, no release up to Codex 0.162.1 or 0.163.0-alpha.5 mentions one, and the openai/codex issues have no maintainer response.
Until that changes, turn search off in ~/.codex/config.toml:
web_search = "disabled"
The Codex config reference lists disabled, cached (default), indexed and live. --yolo and other full-access sandbox settings default to live, so set the value explicitly. Use the setting for new sessions only. If you need search, keep it in a separate short session and pass the results along as a file.
To recover a stuck session, stop prompting it, because every prompt retries the failing compaction. Leave its file in ~/.codex/sessions/ and start a new session with search disabled. Hand it the old file, or a note with the goal, the changed files, the decisions made and what's left.
The API-key path
Every report comes from ChatGPT's subscription backend, the one Codex uses with a ChatGPT login. Codex can also run against the public Responses API with an API key. No one has reported the error on that path. To point Codex at AIHubMix, follow the Codex CLI tutorial:
model = "gpt-6.1-sol"
model_provider = "aihubmix"
[model_providers.aihubmix]
name = "AIHubMix"
base_url = "https://aihubmix.com/v1"
wire_api = "responses"
env_key = "AIHUBMIX_API_KEY"
Billing is per token, at the rates on the gpt-6.1-sol model page.
What we couldn't reproduce or verify
- The API-key path. We haven't run the repro against the public Responses API. "No reports there" is an observation, not a test result.
- The rule itself. It comes from community A/B tests. OpenAI hasn't confirmed it or replied in any of the threads, and hasn't said whether it's intended.
-
Disabling search mid-session. If a session already has
web_search_callin its history and search is then disabled, its normal turns stop declaring the tool and should start failing too. That follows from the rule, but nobody has reported testing it. -
The default "cached" mode. We assume it writes the same hosted
web_search_callitems, but we haven't checked a cached-mode session file ourselves.
The general lesson for anyone running a gateway: replayed history and declared tools now have to agree. Any rewrite that drops tools, whether for compaction, titles or summaries, needs to check what the history still references.
Keep reading: the GPT-6.1 Sol series
- If the failing requests are your first week on gpt-6.1-sol, the migration post lists the other request-level changes to check: Migrating to GPT-6.1 Sol: 9 Things That Can Go Wrong
- Effort settings decide how quickly a long agent session approaches its compaction threshold: Choosing a Reasoning Effort for GPT-6.1 Sol: low to max
Sources
- Windows Codex Desktop: context compaction always fails (openai/codex)
- Remote compaction fails after resuming hosted web_search_call history (openai/codex)
- Codex remote compaction fails with 502 "response protection is unavailable" (OpenAI Developer Community)
- Codex configuration reference (OpenAI)
- Codex CLI + AIHubMix integration tutorial (AIHubMix)
Top comments (0)