Preserved Thinking - Keeping Earlier Reasoning Valid Across a Conversation
If you arrived via
/claude-api preserved-thinking-migration(or opened this file directly): this is the right file. Execute the steps below in order rather than summarizing the guide back to the user - presenting the break profile, the ranked causes, and the measured result of each fix IS part of the execution. Start with Step 0 (scope, quality bar, baseline) and finish with Step 4's two deliverables: the break profile and the changes.
Preserved thinking is measured in units of conversations that keep their reasoning, not requests that pass. One edit to an earlier turn invalidates every thinking block after it, and the same stale block fails again on every later request that replays it - so a per-request count overstates the damage and a per-conversation count (did this conversation break, and at which turn) is the number that tells you whether a fix worked.
What the check is, in one paragraph. On models with preserved thinking (Claude Fable 5.1, Claude Opus 5.5, and Claude Sonnet 5.5 today; the Preserved thinking page lists the models and the enforced accounts, and says later models will enforce the check for all users), each thinking block's signature records the conversation that produced it - the top-level system prompt, the set of tools, and every message before the block - and the model that produced it. When the transcript comes back, the API recomputes that record from what you sent and requires a match (a separate model check decides whether the current model can read the block at all; see "Switching models mid-conversation" in shared/preserved-thinking-migration/causes.md). Integrations that keep the history append-only never notice. Integrations that rewrite earlier turns between requests - truncation, client-side compaction, a re-rendered system prompt, a tool list that grows when a plugin connects, a per-turn reminder that is injected and then stripped, old tool results trimmed after the fact, media dropped by a size cap, a lossy round trip through the app's own message types - lose the reasoning after the edit point (drop_block) or fail the request (error). The check compares the conversation as you sent it, before any server-side edit, so Anthropic's own server-side compaction and context editing never count as edits. The published explanation lives in shared/model-migration.md -> Migrating to Claude Fable 5.1 from Claude Fable 5 -> Breaking change 3 (the three-step check and the append-only form of each edit; read it first if the check is new to you) - this workflow restates the rules only as a lookup, the cause table and keep list in shared/preserved-thinking-migration/causes.md; it finds which edits this harness makes, proves them with the API's own response, fixes them one at a time, and proves the fix the same way.
Where this workflow sits: the migrate subcommand explains the check and lists the append-only form of each edit; this workflow is its executable form - scan, measure, fix, re-measure - for a harness that already exists. cost-optimize is a sibling, not a prerequisite: a history that invalidates its own thinking also restarts the prompt cache at the same point, so a fix here usually shows up in its cache-hit numbers too, and a prefix edit found there ("audit for mid-task cache-breakers") is the same finding as a break found here. Cache discipline and preserved-thinking discipline are very nearly the same discipline, so a harness that is already append-only for caching pays nothing extra here. Once the project has an eval, Step 3 runs as a hill-climb whose metric is the drop count - one change per round, measured, kept or reverted - and the hillclimb subcommand is the loop to use.
Two scripts ship with this guide, extracted beside it under shared/preserved-thinking-migration/, and are used by Steps 1 and 2. Both are dependency-free Python 3; neither needs credentials except the probe's live modes, and neither prints them. The commands below give their paths relative to this skill's base directory (the line at the top of the prompt); run them with that directory prefixed, from the user's project directory, so that relative capture paths resolve there. A reference file is extracted beside them, shared/preserved-thinking-migration/causes.md: the rules for a conversation that switches models, the "Cause -> detection -> fix" table, the keep list, and the failure modes to avoid. It is a separate file so that this guide fits in one Read; Read it when a step below sends you there (Step 1.4 at the latest), not before.
shared/preserved-thinking-migration/prefix_diff.py- diff consecutive request bodies in the parts the check compares, and name the difference in the API's own vocabulary;--scangreps a repository for the usual culprits.shared/preserved-thinking-migration/drop_block_probe.py- replay a captured conversation with the controls turned on and record what the API dropped and why, per turn and per conversation;--self-testis the three-request proof that the check is running (Step 0.5) and exits non-zero when it is not.
Severity tiers - one-off vs. recurring
Tier every cause before planning the work. Tiers 0-2 are one-time fixes: apply them once and the harness stops breaking. Tiers 3-4 keep costing - they recur on every conversation that reaches them, which is why they are the ones to measure before deciding.
| Tier | Meaning | Causes | What to do |
|---|---|---|---|
| 0 | Fine - not an edit | cache_control markers; reordering the tools array; server-side compaction and context editing; a retry; a regenerate, rewind, restored checkpoint or branch; the latest turn edited and resubmitted; a model switch (the model check is separate and is not a prefix edit) |
Nothing. None of these change the compared prefix - keep them out of the report |
| 1 | Accidental | The system prompt re-rendered with per-request content (a date, a counter, live state); drift from an SDK or domain-model round trip | Remove the edit. Nothing about the product needs it |
| 2 | Fixable | The tool set changed mid-conversation; a same-name tool's description or schema rebuilt; a per-turn reminder injected then stripped; tail state re-rendered every turn | Apply the append-only recipe (Step 3). Adding or withdrawing a tool has an append-only form (tool_addition / tool_removal, with the entry left in tools). A same-name description or schema change has one only under the inline-tools-2026-09-15 beta (Claude API), where a tool_addition carries the new definition (Step 3); without it, keep the first-sent bytes for the life of the conversation and accept the stale definition, or offer the changed text under a new name. Append tail state as new turns rather than rewriting it |
| 3 | Recurring | Tail-kept compaction; background compaction (a second request writes the summary while the session continues, then it is swapped in); rolling truncation; a pinned document rewritten every turn | Measure the drops and decide. Without the compact-2026-09-04 beta (on-demand compaction) no append-only client-side form exists for these: send drop_block from the swap onward, or strip the thinking from the kept turns; the recommended shape is simple compaction, done synchronously. With the beta, keep-tail and background compaction become append-only - see Step 3 |
| 4 | Stop | Prefix surgery - snipping, redacting, pruning old tool results, or removing content after a cache breakpoint to save cost; a missing predecessor; an unrecorded strip-and-retry | No workaround exists. The reasoning after the edit point is lost; the fix is to stop doing it |
A tier is a property of the cause, not of one conversation: rank by tier first, then by reasoning lost within a tier (Step 3).
Step 0: Establish scope, quality bar, and baseline
Does your harness change earlier turns, the system prompt, or the tool list? If not, stop - there is nothing to migrate, and saying so plainly is the finding.
When a replayed thinking block no longer matches, the default is a 400 error. Dropping the thinking instead is opt-in, and it is not a fix: it trades a visible failure for the silent loss of that reasoning, and it is not free - dropped blocks aren't billed, but the session's token usage might still increase because Claude can sometimes think more to re-create the dropped thinking ("Failure modes to avoid" in causes.md).
First, establish three things - from the request and the repository where they answer it, and from the user where they don't. This workflow is interactive by design: a capture of real request bodies, a test slice, and every live replay need the user's involvement or approval, and "which of these edits is deliberate" is a question only they can answer. State all three at the top of the report (the baseline may read "pending Step 2" at first).
- Scope. If an official Claude product or SDK (Claude Code, claude.ai, Claude Managed Agents, the Claude Agent SDK) manages the conversation history, there is nothing to migrate - say so and stop. Otherwise: if the request names files or directories, that is the scope. Otherwise it is every place the project builds the three parts of a request the check compares - the
systemprompt, thetoolsarray, and themessagesarray - across every path that touches them between two requests of one conversation: the request builder, compaction or truncation, reminder or context injection, media handling, persistence and resume (anything that re-reads the conversation from a store and re-renders it), process restart and deploy, plugin or MCP connection, sub-agent transcripts, and model switching. Note distinct traffic classes (an interactive chat path and a background agent loop are different harnesses even on one key): the profile, the fixes, and every validation later run per class. A capture taken on an older model may carry request shapes Claude Fable 5.1 rejects before any check runs -thinking.typeenabledordisabled, a forcedtool_choice(anyortool), an assistant prefill as the last message,temperature,top_portop_kwith a thinking configuration - so list them now as things to convert before measuring (Step 2.1 lists what the probe converts; it leaves a prefill alone, and that 400 shows in its "not evaluated" line). Also establish which platform the code targets (Claude API, Claude Platform on AWS, Amazon Bedrock, Google Cloud Vertex AI, or Microsoft Foundry) and which model it runs: the check applies only to models with preserved thinking. The Preserved thinking page states that beta names are the same on Amazon Bedrock and Google Cloud wherever the beta is available there; for any other platform, read that platform's own documentation andshared/platform-availability.mdrather than assuming. Finally establish whether the organization is enforced today. The rule, as the Preserved thinking page reads on 2026-09-23: on Claude Fable 5.1 and Claude Opus 5.5 the check is enforced by default for accounts created on or after August 31, 2026, 00:00 UTC, with the same definition on the Claude API and on the cloud platforms; a request that setsprefix_mismatch_behavioropts in regardless of account age; and later models will enforce the check for all users. Confirm it from the responses production already gets: an enforced account seesinput_transformationsentries (with the header) or 400s whose text says "bound to a different conversation"; an account that is not enforced yet sees no dropped blocks and no 400s until a request sets the field; with the beta header alone, its responses list each block that would fail as athinking_mismatch_allowedentry (Step 2.2). The page's own probe - send an edited history without the beta header: a 400 that names the header means the account is enforced by default, a 200 means it is not (the Message Batches API returns 200 either way). To confirm a 200, resend with the beta header and still no field: the response lists every thinking block after the edit asthinking_mismatch_allowed. Step 0.5 proves the enforced path - run it with a test key, never against production traffic. - Quality bar. Find the project's eval, test suite, or outcome checks for its model calls. The fixes in this workflow are behavior-preserving by construction (they change how the history is carried, not what the model is told), but two of them are not: replacing a compaction scheme, and choosing to drop thinking at a boundary. Those need the eval. If none exists, say so prominently in the report and do not stop: the drop count itself is a measurement, every fix that removes an edit is safe to propose, and the minimal eval recipe in Step 3 is the next step for the two that aren't.
- Baseline. Two numbers, both from Step 2's probe on the same test slice: the share of conversations with at least one prefix break, and the turn at which each one first breaks. Record them before any fix. The three-arm protocol in Step 2 also asks for the eval score with the current harness and preserved thinking off (that is the eval's existing number) - write it down now if it exists.
Step 0.5: Prove the mechanism is wired
Before trusting any measurement, run one deliberately broken conversation and confirm the API reports it. The probe's self-test sends three requests - one mint and two replays: it mints a thinking block with a one-turn conversation, replays it honestly as turn two (expect input_transformations: []), then replays it with the first user message edited (expect one {"type": "thinking_dropped", ...} entry with reason: "prefix_binding_mismatch", and - on the Claude API - the diagnosis header naming pattern=first_message_rewritten):
python3 shared/preserved-thinking-migration/drop_block_probe.py --self-test --model <the target model id> --yes
python3 shared/preserved-thinking-migration/drop_block_probe.py --self-test --model <the target model id> --yes --mode error # expect the 400 insteadThe probe always sends the thinking-binding-controls-2026-08-01 beta value; add --beta <value> for any other beta your production requests carry (the bare command is enough on the Claude API). The probe is first-party only: it authenticates against /v1/messages with ANTHROPIC_API_KEY (sent as x-api-key) or ANTHROPIC_AUTH_TOKEN (sent as a bearer token; when both variables are set only the token is sent, because the API rejects a request carrying both headers) and cannot replay captures taken on Amazon Bedrock or Google Cloud Vertex AI - there, run the account's own client and read input_transformations from its responses.
Three short requests (one mint, two replays), billed at the model's normal rates - state the cost and get the user's approval first, as for every run that exercises the model. If the edited replay comes back with no drop (the probe prints NOT WIRED and exits 1), stop: the check is not running for this request (the wrong model id, a platform without the controls, the header missing, the field misspelled, or a gateway or proxy between the harness and the API that drops the anthropic-beta header or the block_binding field - run the self-test through the same path production uses) and every later number would be meaningless. If the honest replay reports a drop, stop too: something in the probe's path is already editing the history - a proxy, an SDK middleware, a serializer - and that is finding number one. Record both responses in the report.
Step 1: Find the edits
Finding the edits has three sources, in order of evidence: the API's own diagnosis (Step 2), the diff of consecutive request bodies, and the code. Do the diff and the code read in this step; they tell you where to look before spending money on replays, and they are what localizes a break to a line of code after the API has named its shape.
1.1 Capture what the harness actually sends
Capture the exact request bodies of a few normal conversations - every request, in send order, JSON as it went over the wire, including the assistant turns with their thinking blocks and signature values exactly as the API returned them. Include one conversation that runs long enough to trigger compaction or truncation if the product has either, one that connects a plugin or tool mid-session if it can, and one that is resumed from storage or survives a process restart. Capture at the HTTP layer where possible (an SDK hook, a logging transport, a proxy) rather than from the application's own message objects - the second kind of capture hides exactly the re-serialization this check catches. If requests pass through a gateway, proxy, or model router on the way to the API, capture them as they leave that layer: a router can rewrite the system prompt, the tools, or the history after your code has built the request. Store one conversation per .jsonl file, one request body per line. Ten to thirty conversations are enough; they double as the test slice for Step 2.
Capture the request body and the anthropic-beta header only - never x-api-key or Authorization; an MCP server's authorization_token in the body is sent as the harness sent it (the connector's tools are part of what the check compares), so capture with a test-scoped token and rotate it afterwards. The probe reads nothing else and warns when a capture carries a credential.
Handling captures. A capture is the conversation as the end users had it, and it cannot be redacted without breaking the measurement (the check compares the bytes). Keep captures outside the repository (or ignored by version control), never commit them or an eval set derived from them, run the scripts from a machine that may hold that data, and delete the captures when the work is done. The probe's --json output is safe to share - it holds request ids, statuses, entries, headers, digests, token counts and file names, no message content; error text for the conversation check's own 400s is stored as a reconstruction of their fixed form (the block path, the fixed clause, and the first-changed-message diagnostic - never the server line itself), and every other error is reduced to its type and field path because API validation messages can echo request values (the full text still prints on the terminal); prefix_diff.py's output is not, because its attribution lines quote excerpts of the changed content.
If the application cannot capture bodies yet, adding that capture is itself the first diff of this workflow: it is the measurement channel for everything after it.
1.2 Diff consecutive pairs
python3 shared/preserved-thinking-migration/prefix_diff.py captures/conversation-0001.jsonlFor each pair of consecutive requests the script reports MATCH, MISMATCH with a verdict in the API's vocabulary - kind=system_changed; pattern=system_rerendered; sections=system; changed_validated=system.0 - and an attribution line that names the site and the first changed character:
system[0] changed at char 53: "... Be concise." -> "... Be concise. Current time: 2026-09-02T15:04:05Z."
tools: lookup_order description changed at char 23: "...order by id." -> "...order by id. Today is 2026-09-02."
messages[2] (user) content[1] (text) removed: {"text":"<reminder>Answer in one sentenc...
messages[1..2] removed (assistant, user)Two lines matter as much as the verdict. replayed thinking blocks in the later request: N - when N is 0 the pair proves nothing about preserved thinking (there was no block to check), which is common for the first pair of every conversation and for harnesses that strip thinking; and the ! chain line (printed as CHAIN-BREAK, counted in the exit status), which is a break, not a warning - it fires when the already-sent turns come back with their thinking blocks changed in a way the API rejects: the kept blocks must be a contiguous window of the original sequence (dropping from the front, from the back, or both is fine), so a block removed from the middle, or a reorder, fails the block after the gap even though the rest of the prefix is untouched.
The comparison ignores what the API ignores: cache_control markers, string content versus a single text block, leading and trailing whitespace of a text block, whitespace-only text blocks, key order, the order of tools in the array (they are compared as a name-keyed set), a defer_loading tool that no tool result, tool-search result, or tool_addition has named yet, request parameters outside system / tools / messages, everything before the last server-side compaction block (the check restarts there; the diff says when it compared from one), and the thinking blocks themselves. Interior whitespace, tool_use.input bytes, tool-result text, image bytes, and everything else count. Treat the script's pattern as a guess in the API's words - the API's own header in Step 2 is the authority when the two differ.
1.3 Read the code
python3 shared/preserved-thinking-migration/prefix_diff.py --scan path/to/repoThe scan prints file:line leads grouped by cause - timestamps and environment reads inside prompt builders, slicing of the messages array, tool lists mutated after session start, the opening message rebuilt from state, tool results trimmed after the fact, reminder tags stripped with a regex, thinking blocks filtered out, round trips through the app's own message model, media caps and URL re-signing. They are regex leads, not findings: read each one, and confirm it with the pair diff or the API's response before it goes in the report. The scan is optional; the checklist below is not. An application can always express an edit in words the patterns do not know, so a scan with no leads is not proof of compliance, and a scan with leads is a reading list - the pair diff is the instrument.
Whatever the scan finds, read these by hand - this checklist is the mandatory part of Step 1.3; they are where the edits hide:
- Prompt assembly: is anything in
systemor in the first user message computed per request - date or time, working directory, account or user line, git status, memory or instruction files re-read from disk, feature flags, model or client version strings, a token or turn counter? - Tool declaration: is
toolsbuilt from live state (connected MCP servers, plugins, permissions, a feature flag) so that it can differ between request 1 and request 2? Are descriptions or schemas rendered with anything dynamic? - History management: any path that shortens, summarizes, reorders, or rewrites messages already sent - sliding windows, keep-last-N, client-side summaries, "micro-compaction" of old tool results, media caps, context-length recovery after an error - and whether any turns are replayed verbatim after a summary (that decides which recipe applies - Step 3's, or the betas section of
shared/preserved-thinking-migration/causes.md). - Injection: anything appended to a user turn for one request only (reminders, status lines, token counts) and removed or rebuilt on the next.
- Modes: does entering a mode (plan, read-only) swap the tool list or the system prompt in place? Withdraw and re-offer tools with
tool_removal/tool_addition, and deliver the mode's instructions as an appended message. - Tools listed in the prompt: are the callable tools named in the system prompt, behind one generic dispatcher tool? Then adding or removing one is a system prompt edit: keep the prompt fixed and announce each change in an appended message ("You can now call X").
- Persistence and resume: does the conversation round-trip through a database or an ORM, and does the replay rebuild messages from those objects rather than from the stored wire JSON? Does a restart, a resume, a deploy, or a new template version re-render the system prompt or the opening message?
- Thinking handling: does any code filter
thinking/redacted_thinkingblocks (a serializer that skips a block with emptythinkingtext counts: the text is empty by default and thesignaturecarries the reasoning), reorder them, store their text truncated or re-wrapped (a modified thinking block is its own 400), or retry a 400 by stripping them without recording that it did? - Sessions: can a thinking block from one conversation be replayed under another (shared session keys, multiplexed users)?
1.4 Name each edit and decide whether it is deliberate
For every pair-diff verdict and every confirmed lead, record: the cause in the API's words (the pattern), the site in the code, which traffic class it is on, and whether the edit is deliberate (a compaction the product relies on; a user-invoked reset that starts a new conversation) or accidental (a timestamp nobody needed in the system prompt; a plugin landing inline on request 2). Accidental edits are removed outright in Step 3. Deliberate ones are replaced by their append-only form, or - where none exists yet - measured and decided (Step 2's caveats, Step 3's last section). The table under "Cause -> detection -> fix" in shared/preserved-thinking-migration/causes.md is the lookup for both: Read that file now if you have not yet, and check every candidate against its "Keep list" before it goes in the report.
Step 2: Measure with drop_block on a test slice
The API's response is the only ground truth. The client-side diff can miss what it cannot see (media bytes behind a URL, an edit in a part of the request the capture didn't include) and can flag what the API tolerates; the response cannot.
2.1 The request shape
Every request in the test slice carries the beta header and sets the behavior explicitly - this is what turns the check on for an organization that is not enforced by default, and it is what adds the report to the response:
POST /v1/messages
anthropic-beta: thinking-binding-controls-2026-08-01
{"model": "<the target model>", "max_tokens": 4096,
"thinking": {"type": "adaptive", "block_binding": {"prefix_mismatch_behavior": "drop_block"}},
"system": ..., "tools": [...],
"messages": [ ...the full history with thinking blocks replayed verbatim... ]}Rules that save a debugging hour:
- The field without the header is, today, a 400 ending in
block_binding: Extra inputs are not permitted. That is a different 400 from the one an enforced account gets when it replays an edited history without the header, whose text says the block is "bound to a different conversation" and ends by naming the beta value the setting requires; you will meet both, and only the second one means the check ran. The header without the field does not turn enforcement on for an organization that is not enforced yet: no block is dropped and nothing fails. The API records the check instead and lists each block that would have failed as athinking_mismatch_allowedentry (2.2) - the zero-risk way to find edits in production traffic. To measure what enforcement costs, set the field; it is the per-request opt-in. prefix_mismatch_behaviortakes"error"or"drop_block". Write exactly that field name; the probe removes any other key it finds underblock_bindingand says which.- The object is accepted, under the header, on every model that accepts
thinking, so one request body works before and after a model switch, provided the thinking configuration is one every model in the route accepts (enabledis a 400 on Claude Fable 5.1); on a model without preserved thinking it is a no-op that still reports model-check drops. - Keep everything else in the request exactly as production sends it. The probe touches only fields outside the compared prefix and prints each change: the header, this field,
max_tokens(capped, see 2.3),stream(off unless asked),tool_choice(set tonone, see 2.3), and the thinking configuration - a request with nothinkingconfiguration is given{"type": "adaptive"}(on Claude Fable 5.1 that is what a request without the key already runs with),enabledis rewritten toadaptive,temperaturebecomes 1 andtop_p/top_kare removed; none of those changes the verdict. Thinking disabled is skipped.system,toolsandmessagesgo out exactly as captured.
2.2 What the response tells you
Detect drops from the request diff. prefix_diff.py over consecutive requests (Step 1.2) is the detector: it reads only what your harness sent, so it works whatever shape the response takes. The surfaces below confirm and explain what it finds.
Three surfaces, in order of reliability:
input_transformations(response body, top-level, sibling ofusage) - the contract. With the header it is present on every response from a thinking-capable model:[]when nothing was dropped and nothing failed, otherwise one entry per block, of two types.{"type": "thinking_dropped", "path": "messages.7.content.0", "reason": "prefix_binding_mismatch"}: the block was removed.thinking_mismatch_allowed(samepath,reasonalwaysprefix_binding_mismatch): the block failed the prefix check on a request the API does not enforce (an older account with the field unset), so it reached the model unchanged and was billed; every block after the edit gets one, and a request that sets the field never does (the probe always sets the field, so it reads such an entry as the field stripped en route: turn not evaluated, run inconclusive).pathindexes themessagesarray as you sent it.reasonisprefix_binding_mismatch(your history changed - this workflow's subject) ormodel_binding_mismatch(the conversation switched to a model that cannot read the block - not a bug in your code; see "Switching models mid-conversation" inshared/preserved-thinking-migration/causes.md); the probe flags any reason it does not classify asWARN. Dropped blocks are not billed, whichever the reason. Ignore entries whosetypeorreasonyou don't recognize; later checks add values. When streaming, the array arrives on themessageobject inmessage_start(and again in the finalmessage_deltaonly after a mid-stream server-side model fallback). Without the header the field is absent and drops are silent.- A diagnosis header, if present. Some responses that report a drop (or a 400 in
errormode) also carry a response header namedanthropic-thinking-prefix-mismatch- which can be missed on streamed responses (the probe's--streammode may then print no "why:" line), so read it as a second check and detect from the request diff. Anthropic has not published this header; it may change or stop without notice. Use it if present; never depend on it. If it is there, itspatternandchanged_validatedfields are hints - the shape of the edit, in the same wordsprefix_diff.pyuses, and the first changed path in the request - and the probe prints them on its "why:" line. Treat its absence as "no diagnosis", not "no break": the detail (kind, pattern, changed path) is only given for blocks your own organization created - for other blocks the header carries only the bare fact - and a partner cloud's proxy is not guaranteed to forward it. - The 400 text in
errormode - the same diagnosis as one sentence, for code that never sees headers:messages.7.content.0: Invalidsignatureinthinkingblock. The block is bound to a different conversation. Remove the block, or setthinking.block_binding.prefix_mismatch_behaviorto "drop_block". Content that preceded this block when it was created is missing from this request, starting atmessages.2.It usually ends with one sentence naming the first changed path, as in that example (the sentence varies with the kind of edit and is sometimes absent). The request is rejected before any output; retrying the same body fails the same way (what production code does instead: "Failure modes to avoid" incauses.md).
The token-counting endpoint runs the same conversation check. /v1/messages/count_tokens applies it to the replayed blocks: in error mode it returns the same 400 (with the diagnosis header when that is present); in drop_block mode it returns 200 and leaves the dropped block out of the count. A harness that counts tokens before each request meets the 400 there first. The count endpoint costs nothing and samples nothing, so it is the cheapest first-break pass: replay the slice against it in error mode (drop_block_probe.py --count-tokens --mode error) before a paid /v1/messages replay. It returns no input_transformations, so it answers "is anything broken, and where", not "how many blocks".
None of the three says which line of your code made the edit. That is what Step 1's diff and scan are for: the header's pattern and changed_validated tell you where in the request to look; the pair diff tells you what changed there; the scan tells you who wrote it.
2.3 Run the slice
python3 shared/preserved-thinking-migration/drop_block_probe.py captures/ --dry-run # validates the capture, prints the plan, sends nothing
python3 shared/preserved-thinking-migration/drop_block_probe.py captures/ --yes --count-tokens --mode error # free first pass on the token-counting endpoint: 400s mark the breaks
python3 shared/preserved-thinking-migration/drop_block_probe.py captures/ --yes --json probe.json # drop_block replay; one conversation per file
python3 shared/preserved-thinking-migration/drop_block_probe.py captures/ --yes --mode error # the loud arm (prefix_mismatch_behavior "error"), if you want the 400 textNothing is sent without --yes: a bare invocation prints the plan and stops. For CI, a non-zero exit is the signal - the probe exits 0 with no break, 1 when a conversation was inconclusive, 2 when at least one conversation broke (not severity-ordered); the diff script exits 1 on any mismatch or chain break.
Replaying a request does not run the application's own tools - the capture already holds their results - but server-side tools in the request (web search, code execution, MCP connectors) run again on the API side if the model calls them, and every request is billed; the probe caps max_tokens at 16 by default because the verdict is decided before the first output token, which keeps the reply short and the cost to the input side. The probe sets tool_choice to none on every request that declares tools (tool_choice is outside the compared prefix), so no tool - server-side or the harness's own - can be called during a replay; the 16-token cap is for cost, not safety. Do not remove tools from the capture to the same end: the tool set is part of what the check compares and removing one would turn every request into a tool_set_changed break. This spends real money: estimate it from the slice (requests × their input size at current rates - fetch the rates, don't quote remembered ones) and get the user's approval once, for the whole measurement budget of this workflow, before the first run. --max-requests and --max-conversations cap a first run. Use a key from the same organization as the capture - a dedicated key or workspace in that organization is ideal; a capture replayed with another organization's key is not diagnosed, so the run teaches nothing about the harness.
What to log per request (the probe records all of it; production telemetry should too): the conversation id and turn index; the number of thinking blocks replayed in the request; every input_transformations entry; the anthropic-thinking-prefix-mismatch header if present; the HTTP status and, on a 400, the error text; the request id; and a client-side digest of the three bound parts - a hash of the canonical system, of the name-keyed tools, and of messages up to each replayed thinking block - so that the request where the digest changed can be found without the header.
The counting rule. Count new dropped blocks per conversation, not entries per request. A block that fails once fails again on every later request that replays it, so a ten-turn conversation with one edit at turn 3 shows entries on eight responses but has one break. Identify a dropped block by the thinking block itself - resolve each entry's path to the block in the body you sent and key it by its signature (a redacted_thinking block by its data) - not by the path: paths shift whenever the history is truncated or compacted, and the same block would be counted again at its new index. The probe does this and keeps the path in the record. Its per-conversation summary gives first_break_turn, distinct_dropped_blocks (the turns of reasoning the conversation lost) and the patterns seen; the share of conversations with a first break is the baseline from Step 0, and the first-break turn is what you compare against the code: the request where the digest changed is the request after the edit.
Suspect the test slice before the harness. A slice that never replays a thinking block produces a perfect score and proves nothing - the probe warns when a conversation's maximum replayed-thinking count is zero. Check that number first, on every run. The first request of every conversation replays nothing by definition; an adaptive-thinking model may answer a short turn with no thinking block at all (Claude Fable 5.1 at its default effort often does, so a slice of short exchanges can carry no thinking anywhere); a harness that strips thinking before sending has nothing to check. The slice must also be captured on the target model: a block minted by a model without preserved thinking carries no conversation record for the check to verify, so replaying such a capture can never show a prefix break (the probe prints the capture's model ids - check them). Before replaying, count the thinking blocks in the captured assistant turns (the probe's --dry-run prints the number it will replay). If the production traffic genuinely carries little thinking, say so in the report - that is a finding about exposure, not a pass - and, for a capture made specifically to test the harness, drive the conversations with tasks that need reasoning or capture them with output_config.effort raised ("max" on Claude Fable 5.1) so that the assistant turns hold thinking; keep everything else as production sends it. A zero on a slice whose conversations replay thinking on most turns is the result you want; a zero on any other slice is a broken measurement.
2.4 The three-arm protocol (validation against the eval)
When the project has an eval, treat the migration as an A/B/C experiment on one frozen set of inputs. Arm 1 is the harness as it is today, check not enforced: the eval's existing score, no new run. Arm 2 is the same harness with drop_block set in the eval runner's configuration (never in production code), with input_transformations recorded per response and joined to each conversation's score and token usage; the join answers whether the conversations that lost reasoning scored worse or used more tokens, and by how much. Arm 3 is the harness after Step 3's fixes, again with drop_block set; the target is no prefix_binding_mismatch entries on the slice (model-check entries are counted separately: "Switching models mid-conversation" in causes.md) and a score within noise of Arm 1. Report the three scores, the token usage and the two drop counts side by side, per traffic class.
2.5 Caveats that change what the measurement means
- Keep-tail and background compaction have no append-only client-side form without the
compact-2026-09-04beta. Summarizing older turns and keeping the newest ones verbatim fails on the kept turns (their thinking was produced with the full history present); compacting off the critical path and swapping the summary in later fails the same way for every turn produced above the swap point. The choices are: server-side compaction or context editing where the product can use them; simple compaction (summary plus the new turn, nothing older replayed); keep the scheme and strip thinking from the retained turns as a deterministic, recorded strip; or keep the scheme and senddrop_block(equivalent in effect: both lose the same reasoning; the strip is explicit,drop_blockis one field). Measure the scheme you keep - Arm 2 tells you what it costs - record the decision, and do not assume a compaction rewrite is required. With the beta (on-demand compaction; Step 3), its append-only form is one more scheme to measure the same way, with Arm 2 as its before number. - Same-name tool definition changes and tools that re-list. Under
mid-conversation-tool-changes-2026-07-01,tool_additionis by reference, so it cannot express "the same tool, now with a different description", and a connector whose tool list changes on reconnect has no natural append-only form. Strictly, one exists - declare the revised tool under a new name withdefer_loading: true, announce it withtool_addition, and withdraw the old name withtool_removal- but whereinline-tools-2026-09-15is not available the simpler answer, and the recommendation, is to freeze each tool's text for the life of the conversation and replay it: stale but valid. On the Claude API, send the changed definition by value under that beta instead (Step 3). - Partner clouds run the check too (Amazon Bedrock and Google Cloud Vertex AI). The Preserved thinking page states that beta names are the same on Amazon Bedrock and Google Cloud wherever the beta is available there; each beta's own page lists its platforms, and for any other platform, read that platform's own documentation and
shared/platform-availability.mdbefore assuming the controls are offered. On every platform the response is the only place the result is visible, and whether the diagnosis header is forwarded on a 200 is up to the platform. Where the controls are not offered the opt-in test does not apply, and recovery from a 400 is the same as Step 3's last resort: strip thinking from the rejected block onward, recorded and persisted. - The Message Batches API drops failing blocks under its unset default instead of failing the item, and sends no header; set
"error"explicitly if batch items should fail. Whether batch results carryinput_transformationsis unverified - test one batch before relying on it. - Organizations that are not enforced yet get no dropped blocks and no 400 until a request sets the field. With the header alone they get
thinking_mismatch_allowedentries instead, so sending the header in production and logging those entries finds the edits without changing what the model receives - a record-only pass worth running before Step 2's replay. The field opts a replayed request in, one request at a time, with the rest of production untouched. The same fact is the production hazard: copying the field into production code lifts the exemption on every request that carries it, and every conversation that breaks today starts losing its reasoning (or failing) at once - do not do that before Step 3's fixes have landed. - Two checks share the response. A conversation that moves to a model that cannot read the block gets
model_binding_mismatchentries; those are expected, unbilled, not a 400 in any test to date, and not a harness bug - but they are reasoning lost, so the probe counts them apart from prefix breaks and the report states them (see "Switching models mid-conversation" inshared/preserved-thinking-migration/causes.md). Onlyprefix_binding_mismatchis this workflow's metric.
Step 3: Fix one cause per diff, re-measure, keep or revert
Work the causes in order of turns of reasoning lost: for each cause, sum distinct_dropped_blocks over the conversations whose first-break diagnosis carried that pattern (the probe's per-conversation summary gives both; a conversation with an early break loses more blocks than one that breaks late) - the cause that breaks every conversation at turn 2 comes before the one that breaks a tenth of them at turn 30.
When several causes hit the same request - the usual case in a harness that grew over time - the unit of work is the attribution line, not the pattern. Run prefix_diff.py on the first-break pair of each conversation; every line it prints (system[0] changed at char 78, messages[0] (user) content[0] (text) changed, messages[2] (user) content[1] (text) removed, ...) is one edit with one site in the code, and the API reports only the earliest of them (the header names the first failing block's cause; the kind says multiple). Rank the lines, fix each as its own diff, and measure each fix with the pair diff: the fix is right when that line disappears from the first-break pair. Use the probe for the end-to-end re-measure only after the whole set of lines on that pair is gone - the drop count cannot move while any edit on the first-break pair remains, so a probe run after a single correct fix will show the same breaks. Do not read that as "the fix did nothing" and revert it; read the pair diff.
Each cause that earns a place becomes its own diff (one cause per diff, so a revert is clean and the effect attributes). Diffs are proposed by default - presented to the user with the attribution line they clear and the measurement that will prove it - and applied only when the user asks; then measured: the pair diff first, then - once the first-break pair is clean - re-run the probe on the same slice, compare the share of conversations with a break and the first-break turn against the previous kept state, and, when the eval exists and the change is one of the two behavior-affecting kinds, re-run Arm 3. A diff whose attribution line goes to MATCH is kept; a diff that changes nothing in the pair diff is either a miss (the slice didn't exercise that path - extend the slice, not the claim) or a wrong diagnosis; a diff that clears its line but moves the eval is reverted and recorded. Never keep or revert on one conversation's swing; the slice is the unit.
Accidental edits are removed. The recipes below are the append-only forms for the deliberate ones, the same shapes Anthropic's own agent products use, since they face every one of these problems (the "Cause -> detection -> fix" table in shared/preserved-thinking-migration/causes.md maps each pattern to its recipe here, and its model-switch section covers a harness that routes between models). They share one principle: the transcript is the source of truth, and everything the model needs to know later is added at the tail, never written into the head.
Freeze the rendered prompt; deliver changes as appended messages. Render the system prompt once, at conversation start, and store the rendered bytes with the conversation record; every later request of that conversation sends the stored bytes - across process restarts, deploys, template updates, and client versions - not whatever this turn would render. Move every per-session or per-request fact out of system and out of the opening message: date and time, the user or account line, working directory, environment, instruction or memory files, feature flags, model and client version. Announce them once in the first user turn (or in a role: "system" message appended after it), and afterwards send only deltas, as a new appended message that says what changed ("Primary working directory: /repo/worktrees/x (was /repo)"; "Instruction files were re-read; these differ from their earlier copies: ..."). Mid-conversation role: "system" messages carry system-prompt authority and become part of the conversation record later blocks are checked against, so a change delivered this way is as strong as a re-rendered prompt and invalidates nothing; a plain one needs no beta header on Claude Fable 5.1 (only clear_at and the tool-change blocks below do). The only times the prompt is rendered again are deliberate boundaries - a new conversation, a user-invoked reset, the request after a full compaction - and a deliberate boundary is declared in the logs so a diff at that point is not mistaken for a bug.
Declare the initial tool set at the start; never edit an entry; surface late tools and withdrawals by reference. For a small fixed set, build the full tools array before the first request, with defer_loading: true on tools that may not be ready; store the array as sent and replay it. For a large catalogue the model will mostly never use, leave not-yet-enabled tools out of tools and, on the turn one becomes available - or a tool connects later (an MCP server, a plugin, a permission granted mid-session) - append it to tools with defer_loading: true (a deferred tool nothing has referenced yet is outside the compared prefix, so appending it is safe; the Preserved thinking page documents this form) and announce it with a tool_addition block in an appended role: "system" message (beta mid-conversation-tool-changes-2026-07-01; it must follow a user message, such as the tool_result turn; right after a paused assistant turn that ends in a server-tool result a text-only system message is accepted but a tool change is a 400, so resume that turn first); never append a regular tool. A tool whose provider goes away is never removed from the tools array: to withdraw it, announce a tool_removal block in an appended role: "system" message (same beta) and leave the definition in place, returning an ordinary "not available" error if the model still calls it. Freeze each tool's description and schema text for the life of the conversation - a refreshed token, a date, a live listing, or a version string inside a description is a re-render.
With the tool search tool in the request, keep a tool out of tools until it is available. Search can find and call a deferred tool (the Tool search page), so append the tool with defer_loading: true and a tool_addition block on the turn it becomes available.
Keep per-turn reminders in the history. A reminder that should apply to one turn goes out as a turn-scoped system message - {"role": "system", "clear_at": "next_user_message", "content": "..."} appended after the tool_result message it applies to (beta mid-conversation-system-clear-at-2026-08-21; a role: "system" message must follow a user message, or an assistant message that ends in a server-tool result - anywhere else it is a 400, not a binding failure) - and every earlier copy stays where it is, byte for byte: a cleared message renders nothing, costs no input tokens, and is still part of the conversation record the thinking is checked against. Three details from the platform page: a turn-scoped message carries text content only; it takes no cache_control marker, so put the cache breakpoint on the user turn before it; and a user message that holds only tool_result blocks counts as the "next user message" that clears it. Without that beta, append the reminder as a text block after the tool_result blocks in the same user message and leave earlier copies in place; the model acts on the newest one. Rewording, rebuilding from current state, or deleting a copy already sent is an edit like any other. The same rule covers any mid-conversation role: "system" message: persist it with the transcript and replay it, including in sub-agent transcripts.
Size old messages before they are sent, never after. Tool results, documents, and images are sized at ingestion - truncate the output, downscale the image, count the tokens - before the first request that carries them, and never touched again. Later trimming goes through server-side context editing (tool-result clearing, thinking clearing) or server-side compaction, which do not count as edits. For images specifically, the options in order of preference: (a) downscale at ingestion so that keeping every image in the history is affordable - the only option that loses nothing; (b) return generated or fetched images inside the tool_result of the tool that produced them, because server-side context editing can clear old tool results without a client edit, whereas there is no server-side way to prune an image that sits in a plain user turn; (c) if a client-side cap over user-turn images is unavoidable, make what it strips a deterministic function of the append-only history, stripping down to the cap minus a headroom so that a crossing happens once every N images rather than on every request - and say plainly that each crossing is still an edit that costs the reasoning after it; (d) the Files API (file_id) is for content whose bytes would otherwise drift between turns (a re-fetched URL, a re-encoded upload) - it does not reduce the tokens an image costs, so it is not a cap.
Store and echo wire bytes; never rebuild history from a domain model. Persist the messages array exactly as sent and the assistant content exactly as received - in particular tool_use.input as the API produced it (keep a normalized copy for your own execution if you need one, but echo the original) and text blocks untrimmed. Replay those bytes. A round trip through ORM objects, dataclasses, or a "normalize" pass is where interior whitespace, number formatting, key coercion, and string-versus-block shapes drift; the check tolerates leading and trailing whitespace and the string-versus-single-text-block shape, and nothing else.
Compact in a shape the check honours. Prefer server-side compaction (its instructions parameter takes your own summarization prompt; on Claude 5.1 and later models threshold compaction with custom instructions summarizes without the earlier thinking, while on-demand compaction's summarizer always reads it) or context editing - the checked prefix restarts at the compaction block. Client-side, the recommended shape is simple compaction: when the conversation grows too long, summarize it into one message and start the next request with that summary and the new user turn, replaying nothing older - no earlier turns, no earlier thinking. The summary is a plain user message, so there is no thinking left to fail the check. Threshold compaction writes its summary with the model named in the request; an on-demand compaction request can name another supported model, but kept turns' thinking stays valid only if every compaction request since it was produced ran on a model with preserved thinking. Never compact in the middle of a tool round (an assistant turn whose tool_use is still waiting on its tool_result goes back with its thinking intact). "Summarize the last N turns and drop the rest" is simple compaction too, as long as nothing older than the summary is replayed verbatim. If the product keeps a verbatim tail, strip the thinking and redacted_thinking blocks from the retained turns - text and tool calls stay - and make the strip a deterministic function of the stored compacted transcript (every assistant turn older than the compaction boundary goes out without thinking), so the same bytes go out on every later request and after a restart; a harness that re-derives the transcript from a store that still holds the thinking needs a recorded marker to get the same result. A rolling keep-last-N scheme pays this at every compaction (Step 2.5 compares the strip with drop_block). Name compaction in the logs as the one sanctioned boundary where the prefix legitimately changes.
Where the product can use one of two newer betas, its append-only form is one more scheme to measure: compact-2026-09-04 (on-demand compaction; not on Amazon Bedrock - its page's Compatibility list names the models and platforms) for background and keep-tail compaction, inline-tools-2026-09-15 (Claude API) for tools learned mid-session, same-name tool changes and connectors that re-list. Recipes and rules: "Append-only forms under newer betas" in shared/preserved-thinking-migration/causes.md; check the platform pages for availability first.
Remove thinking only as a contiguous run, and record every strip. The check accepts any contiguous window of the original thinking blocks - a run dropped from the front (the oldest first; after a compaction block, the oldest after it), a run dropped from the back, or both - and nothing else: a block removed from the middle, or a reorder, fails the block after the gap and every one after that. Two shapes follow from it. When a 400 forces a strip-and-retry, strip from the rejected block onward (a trailing run) and persist that the strip happened, so later requests send the same stripped history, not the refused blocks. And never thin the middle. Once a block is removed, leave it out: putting it back invalidates the thinking produced while it was gone.
Keep conversations apart. A block from another conversation fails as kind=unrelated / pattern=foreign_prefix. That is a session-keying bug, not a prefix edit; fix the key. Reasoning cannot be carried into a new conversation: a branch that replays the history unchanged up to the fork keeps it; anything else starts from a summary.
When no natural append-only form exists for a shape - keep-tail or background compaction without the compact-2026-09-04 beta, a same-name tool definition change or a connector that re-lists without the inline-tools-2026-09-15 beta (the rename-and-withdraw form above exists but is rarely worth it) - the diff is the decision, not a code change: measure the cost with Arm 2, choose "error" (a mismatch can only mean a bug, fail loudly) or "drop_block" (degrade, keep serving) for production, set it explicitly under the header, and log input_transformations or the 400s either way. Record the cause as measured and decided in the report, with the setting chosen and what it costs, so the next person doesn't re-litigate it.
Minimal eval recipe, for the two behavior-affecting fixes when no eval exists: a frozen set of 20-30 real conversations from the capture; a per-conversation judgment that is cheapest for the workload (golden outputs to diff against, a short rubric, or an automated checker); a runner that replays one configuration and reports pass rate beside the probe's break share. Three arms, approved as one budget.
Finally, the production setting is its own diff, and the last one. Under the header, choose "error" or "drop_block" and set it explicitly; do not leave the field unset, because the defaults differ by surface and by account age, and an unset field on an account that is not yet enforced means the check is recorded, not applied. Apply that diff only with the user's explicit approval, only after the slice shows zero new drops for every cause that was fixed, and never "error" in production while a measured-but-unfixed cause remains - "error" belongs in CI, where one multi-turn capture per traffic class replays with it so that a new prefix edit fails the build. In production, alert on the first prefix_binding_mismatch entry per conversation, not on the count.
Before the report, check the run against "Failure modes to avoid" in shared/preserved-thinking-migration/causes.md.
Step 4: Deliverables
- The break profile and plan: the Step 0 assumptions (scope, traffic classes, platform and model, enforcement status, quality bar), the Step 2 baseline (share of conversations with a break, first-break turn distribution, replayed-thinking coverage of the slice), and the causes found - each with its
pattern, its site in the code, whether it is deliberate, and the turns of reasoning it costs - ranked by reasoning lost. Causes with no append-only form are listed as measured and decided, with the setting chosen. - The changes: one diff per cause, in the order proposed (and applied, when the user asked for that), each tagged applied and measured (break share and first-break turn before and after; Arm 3 score where the eval exists), proposed (with the expected effect), or needs an eval (the two behavior-affecting kinds without one). Plus the production setting chosen (
"error"or"drop_block") as its own, last, explicitly approved diff - applied only once the slice is clean for the fixed causes, and never"error"while an unfixed cause remains - where it is set, the CI replay, and the alert. "No changes recommended" - the slice replayed thinking on most turns and nothing was dropped - is a successful outcome; say it plainly.
Report skeleton (section order and required columns - keep the rest flexible):
- Scope / traffic classes / platform and model / enforcement status / quality bar (Step 0)
- Baseline - conversations with a break (share), first-break turn (median and range), replayed-thinking coverage of the slice (Step 2)
- Causes found - table columns:
Pattern | Site (file:line) | Traffic class | Deliberate? | Conversations affected | First-break turn | Reasoning lost (turns)- captioned "ranked by reasoning lost" - Changes - one diff per cause, numbered in application order, each tagged applied and measured / proposed / needs an eval, with before-and-after break share
- Causes measured and decided - shapes with no append-only form: what was measured, the cost of keeping it, the setting chosen
- Reasoning lost by routing - conversations routed to a model that cannot read their thinking (
model_binding_mismatch), the turns affected, and the routing decision recommended - Production setting, CI replay, alert - what was set where
- Next step / approvals needed - measurement budget, eval prerequisite, or "no changes recommended"
Sources and live references
Facts about the check, the request fields, and the response surfaces above are snapshots; the pages win where they differ. Fetch them when the user needs the full write-ups or current availability:
- Preserved thinking - the platform guide (
https://platform.claude.com/docs/en/build-with-claude/preserved-thinking): the check,prefix_mismatch_behavior,input_transformations, enforcement by account age, the beta names on Amazon Bedrock and Google Cloud, and an FAQ that covers model switches, non-Claude turns, and resuming after a restart. Anchors:#tool-changes(Step 3's tool recipe),#custom-compaction-on-the-clientand#keep-tail-compaction(Tier 3). - The Claude Fable 5.1 migration guide (
https://platform.claude.com/docs/en/models/fable-5-1/migration-guide, breaking-changes item 3, "Editing earlier turns invalidates thinking blocks"): the three-step check and the append-only form of each edit; the same material is inshared/model-migration.md§ "Breaking change 3", which this guide links to. - Extended thinking (
https://platform.claude.com/docs/en/build-with-claude/thinking#preserved-for-model): the section "Only for the model that produced it, or a newer one" - which models read which models' thinking blocks. - Tool search (
https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-search-tool):defer_loadingand search. - Mid-conversation system messages and tool changes (
https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages):role: "system"messages,clear_at,tool_addition/tool_removal, and defining a tool inside a message (betainline-tools-2026-09-15). - Compaction and context editing - the platform's server-side compaction and context-management docs (
https://platform.claude.com/docs/en/build-with-claude/compaction, with its on-demand compaction and "Keep thinking blocks valid" pages): the shapes the check honours, including on-demand compaction (betacompact-2026-09-04). - Per-platform availability of the opt-in controls:
shared/platform-availability.md, and the Preserved thinking page.