All skills
anthropics avatar

/claude-api

@8a1541c official
by Anthropicanthropics/skills179k stars
21,194

Reference for the Claude API / Anthropic SDK — model ids, pricing, params, streaming, tool use, MCP, agents, caching, token counting, model migration. TRIGGER — read BEFORE opening the target file; don't skip because it "looks like a one-liner" — whenever: the prompt names Claude/Anthropic in any form (Claude, Anthropic, Fable, Opus, Sonnet, Haiku, `anthropic`, `@anthropic-ai`, `claude-*`, `us.anthropic.*`, `[1m]`); the user asks about an LLM (pricing/model choice/limits/caching) — never answer from memory; OR the task is LLM-shaped with provider unstated (agent/MCP/tool-definition/multi-agent/RAG/LLM-judge/computer-use; generate/summarize/extract/classify/rewrite/converse over NL; debugging refusals/cutoffs/streaming/tool-calls/tokens). SKIP only when another provider is being worked on (overrides all triggers): OpenAI/GPT/Gemini/Llama/Mistral/Cohere/Ollama named in the query; OR `grep -rE 'openai|langchain_openai|google.generativeai|genai|mistralai|cohere|ollama'` over the project hits (run this grep FIRST if no provider named — don't Read the file).

Use this Skill: https://skilld.dev/gh/anthropics/skills/claude-api

This session only. Nothing lands on disk.

sharedpreserved-thinking-migrationcauses.md

≈9.5k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Preserved Thinking - Causes, Fixes, and the Keep List

Reference for shared/preserved-thinking-migration.md (the preserved-thinking-migration workflow). This file is not a workflow: it holds four lookups that the guide uses - the rules for a conversation that switches models, the "Cause -> detection -> fix" table, the keep list, and the failure modes to avoid - kept apart from the guide so that the guide fits in one Read. Read this file when the guide sends you here (Step 1.4, Step 2, or Step 3), and use the step numbers below as references into that guide.

Switching models mid-conversation

A harness that routes one conversation to more than one model - a cheaper model for easy turns, a fallback when the primary is unavailable, an upgrade from the model the conversation started on - meets a second check that shares the response with the prefix check. The rules below are what the API does today, as observed on the Claude API with claude-opus-5, claude-opus-5-5, claude-fable-5, and claude-fable-5-1 in both prefix_mismatch_behavior modes; the Preserved thinking page (under "Sources and live references" in shared/preserved-thinking-migration.md) states the same rules, and the Extended thinking page (section "Only for the model that produced it, or a newer one") lists which models read which blocks.

Is the model part of what the signature records? Yes. A thinking block's signature records the model that produced it alongside the conversation record and the previous thinking block. When the block is replayed, the API first asks whether the model now reading it reads blocks from the model that produced it (the page's rule: the model that produced it, or a newer one) - the model check - and only then whether the conversation before the block is unchanged - the prefix check. The model field of the request is not part of the conversation record, so changing it is not an edit: a model switch by itself never fails the prefix check.

Does a switch break the check? As of 2026-09-03 the API behaves as follows: a model switch does not fail a request - upgrade, downgrade, a round trip, and a downgrade combined with an edit all return 200, in "error" mode and in "drop_block" mode. "error" makes edits loud; it does not make a model switch loud, because the model check has no error setting: a block the current model cannot read is dropped, not rejected. Read the Preserved thinking page for the current rule. What differs by direction is whether the earlier reasoning is used:

  • Upgrade (an older model's blocks replayed to Claude Fable 5.1): nothing to do. Claude Fable 5.1 reads thinking produced by Claude Opus 5 and Claude Fable 5, among other earlier models (the Extended thinking page has the full list). The blocks are kept, fed to the model, counted in input tokens, and input_transformations is []. An older model's block carries no conversation record for Claude Fable 5.1 to compare, so an edit made before such a block does not invalidate it; a conversation migrated from Claude Opus 5 can only break on the turns Claude Fable 5.1 produces from then on.
  • Downgrade (Claude Fable 5.1's blocks replayed to Claude Opus 5 or Claude Fable 5): the reasoning is lost for that request, not the request. The older model cannot read them, so the API leaves them out of that call, reports each one as {"type": "thinking_dropped", "reason": "model_binding_mismatch", "path": ...} when the beta header is on, and does not bill the dropped tokens. As of 2026-09-03 this is not a 400 in either mode, and there is no diagnosis header (that header belongs to the prefix check). Your messages array is never edited: the blocks stay in your history.
  • Round trip (Claude Fable 5.1, then an older model, then Claude Fable 5.1 again): nothing is lost, provided the client keeps the history intact. Back on Claude Fable 5.1 every block is read again - the ones from before the switch, the older model's own thinking block, and the ones minted after the return. Observed through four turns: a Claude Fable 5.1 block minted after the older model's turn is valid, and the chain of Claude Fable 5.1 blocks does not care that a turn between them came from another model. The same four-turn history sent back to the older model drops exactly the Claude Fable 5.1 blocks and keeps its own.
  • Claude Opus 5.5 runs the same prefix check, and sits beside Claude Fable 5.1, not under it. Everything in this file that names Claude Fable 5.1 as the model that runs the prefix check holds for Claude Opus 5.5 too (observed 2026-09-23: an edited history is a 400 in "error" mode, a prefix_binding_mismatch drop in "drop_block" mode, and a thinking_mismatch_allowed entry with the field unset). What differs is who reads whose blocks, per the Preserved thinking page: Claude Opus 5.5 reads thinking from Claude Opus 5 and earlier Opus, Sonnet and Haiku models, but not from Claude Fable or Claude Mythos models; on the Claude API, Claude Fable 5.1 and Claude Mythos 5.1 read Claude Opus 5.5's blocks, and no other model does. So Claude Opus 5 to Claude Opus 5.5, and Claude Opus 5.5 up to Claude Fable 5.1 on the Claude API, keep the reasoning ([]); Claude Fable 5.1 to Claude Opus 5.5, or Claude Opus 5.5 to any other model, is a downgrade in the sense above - model_binding_mismatch drops, 200 in both modes.
  • A downgrade can hide an edit for one request. The model check runs first, so a block the older model cannot read is never prefix-judged: a request that switches down and edits earlier content reports only model_binding_mismatch, even in "error" mode. The edit is still there; it surfaces on the next Claude Fable 5.1 turn as an ordinary prefix break (a 400 in "error" mode, prefix_binding_mismatch in "drop_block" mode). Scan and measure every turn of a mixed-model conversation, not only the turn where the switch happened.
  • Turns from a non-Claude model do not invalidate earlier thinking, provided they are appended after the existing history as ordinary assistant messages (text and tool_use content, no thinking blocks) and nothing earlier changes. Claude Fable 5.1 thinking minted after such a turn stays valid.
  • A twin model that reads the blocks but runs no conversation check (the Extended thinking page lists which models read which): its turns report [] even on an edited history, so an edit made on such a turn surfaces only on the next Claude Fable 5.1 turn - the same one-request blind spot as a downgrade, with [] instead of a model drop to notice it by.

Does changing a tool description break the check? Yes. Tools are compared as their full definitions - name, description, and input_schema - so changing even one tool's description invalidates every thinking block minted before the change, reported as prefix_binding_mismatch with pattern=tool_schema_changed (the API reports a description edit and a schema edit with the same word). Adding or removing a plain tool has the same effect, reported as tool_set_changed. Reordering tools is fine, and a defer_loading: true tool is outside the comparison until something references it. The fix is the tool_schema_changed and tool_set_changed rows: freeze each tool's text for the life of the conversation, and add or withdraw tools by reference - or, under inline-tools-2026-09-15 (Claude API), append a tool_addition carrying the new definition ("Append-only forms under newer betas", below).

Cause -> detection -> fix. A model switch produces findings in two shapes, and only the second is a harness bug:

  1. The conversation is routed to a model that cannot read its thinking. Detection: model_binding_mismatch entries on the older model's turns; the probe prints them as model_drops=N per turn and model_drop_turns per conversation, apart from the prefix-break count, and prefix_diff.py prints model_switch=A->B on the pair where the model field changed (never as a MISMATCH). Fix: a routing decision, not a history edit. Send the same messages - thinking blocks included - to every model, and leave the beta header and block_binding field in place across the switch (the object is accepted on every model that accepts thinking). If the product needs the reasoning on those turns, keep the conversation on the model that produced it and switch models at conversation boundaries; if it needs the older model on those turns, accept that they run without the newer reasoning. Report it as reasoning lost by routing - conversations and turns affected - beside the prefix breaks, so the owner can decide.
  2. The switch triggers an edit. Three kinds, each an ordinary prefix break with the switch as its trigger: (a) the harness strips thinking blocks on a switch, or rebuilds the history from what each model was shown - the diff shows the blocks removed (predecessor_missing if from the middle), and the reasoning is gone for good when the conversation returns; fix: stop stripping, the API already leaves out what the current model cannot read. (b) A re-rendered system prompt or tool list that persists past the switch - a fallback banner that stays in system from then on, a prompt that depends on which model answered last - reaches Claude Fable 5.1 together with blocks minted under the old text: system_rerendered, tool_set_changed, or tool_schema_changed on the first Claude Fable 5.1 turn after it (on a downgrade the older model's turn in between reports only the model drop). A prompt or tool set that is a pure function of the model being called is different: every Claude Fable 5.1 request then carries the same text the blocks were minted under, so the check finds nothing (verified: the original prompt restored on the return to Claude Fable 5.1 gave 200 and [] with the block fed, in both modes), even though the pair diff flags the switch pairs and the older model's turns still lose the reasoning through the model check. Fix for the persisting case: the rows for those patterns - keep each model's prompt and tool text stable across that model's own turns, and deliver anything that must change mid-conversation as an appended role: "system" message. (c) A request shape the target model does not accept - a thinking configuration it rejects (a 400 from request validation whose text names the field, not a binding failure), or mid-conversation role: "system" messages, and with them tool_addition, tool_removal and clear_at, on a model the Mid-conversation system messages page does not list (it names Claude Sonnet 5 as unsupported: "Use the top-level system field there instead"); fix: one request body that every model in the route accepts.

Cause -> detection -> fix

The pattern column is the word the API's diagnosis header and the diff script both use. "Diff shows" is the attribution line from prefix_diff.py; "scan lead" is the heuristic --scan reports; the fix is the append-only form from Step 3 of shared/preserved-thinking-migration.md. One caution on reading the pattern word: for a summary-plus-tail compaction the header's word depends on how much was removed - the same code path reports tail_kept, compaction_summary, or unknown with kind=blocks_replaced on different conversations - so identify that cause by the kind (blocks_removed or blocks_replaced) together with the diff's attribution lines, not by one pattern word.

Tier pattern (and kind) What the harness did Diff shows Scan lead Fix
1 system_rerendered (system_changed) The system prompt was rebuilt with per-request content: time, cwd, account line, memory or instruction files, flags, version strings system[i] changed at char N time and environment reads near prompt builders; templates rendered per request Render once, store the bytes with the conversation, replay; per-session facts go into the first turn; changes go out as appended role: "system" messages
2 tool_set_changed (tools_changed) A tool was added or removed after the first request: a plugin or MCP server connected late, a provider disconnected, a permission changed tools: X added / removed tools mutated after session start; a tool listing fetched per request Declare the full set at start; append a late tool with defer_loading: true (safe while unreferenced) and announce it with tool_addition in an appended system message - never append a regular tool; never remove one from the array - withdraw it with a tool_removal block and leave the definition in place, returning an ordinary "not available" error if the model still calls it
2 tool_schema_changed (tools_changed) Same tool names, different description or schema text: a date, a refreshed token, a live listing, a version inside a description tools: X description changed at char N descriptions or schemas built from templates or state Freeze each tool's text for the conversation; store and replay the definitions as sent. Without the inline-tools-2026-09-15 beta no append-only form expresses a same-name change - a new name is the only way to offer changed text; under it, append a tool_addition carrying the new definition instead (the betas section below)
1-2 system_and_tools_changed (multiple) Both re-rendered, messages untouched: a connector landing on request 2, or a restart or resume re-deriving both both of the above startup, resume, reconnect paths Replay the stored prompt and tool text across restarts, and across model switches (a switch is not a boundary: a re-render that persists into later requests on the model that minted the blocks is this break, with the switch as its trigger; a prompt that is a pure function of the model called is stable on each model's own turns and is not); the only declared boundaries are a new conversation, a user-invoked reset, and the request after a full compaction
1 first_message_rewritten (blocks_modified / blocks_removed, often with system in sections) The opening user message carried context rebuilt from live state: environment, instructions, a session date messages[0] (user) content[j] changed at char N messages[0] assigned after creation; a context header rendered per request Announce context once and freeze it; send later changes as an appended message describing the delta
3 rolling_truncation (blocks_removed) The oldest turns were dropped whole - a sliding window messages[0..k] removed messages[-N:], keep-last, window size Without the compact-2026-09-04 beta no client-side form keeps the thinking (with it, on-demand compaction does, but the window must summarize rather than only drop - the betas section below). The choices: server-side compaction or context editing; simple compaction (summary plus new turn, nothing older); or keep the window and strip the retained turns' thinking as a deterministic, recorded strip, or send drop_block (equivalent in effect) - measured
3 tail_kept (blocks_removed) A run of older turns removed (or replaced by a summary the record can't see) with the first message kept and the newest turns verbatim - keep-tail compaction or keep-first truncation messages[i..j] removed, messages[0] intact summarize(messages[:-k]) plus messages[-k:] Same as above; the retained turns' thinking cannot verify without the compact-2026-09-04 beta (it does behind an on-demand compaction block, the betas section below) - send drop_block from the compaction onward or strip that thinking as a recorded decision, never mid tool-round; measure, decide, and record the decision
3 compaction_summary (blocks_replaced) Older turns replaced in place by a shorter summary, the tail intact messages[i..j] replaced by 1 message(s) same Same; or move the summary to simple compaction (replay nothing older than the summary); or, under compact-2026-09-04, on-demand compaction (the betas section below)
4 tool_results_rewritten (blocks_modified) Old tool_result content trimmed or cleared after it was sent messages[i] (user) content[j] (tool_result) changed at char N tool-result truncation applied to earlier turns Bound outputs before the first send; later clearing through server-side context editing (clear_tool_uses_20250919, beta context-management-2025-06-27); a client-side prune only at a declared boundary, as a pure function of the growing history
1 tool_use_rewritten (blocks_modified) Old tool_use.input re-encoded or normalized on replay content[j] (tool_use) changed input normalizers, to_dict on tool calls Echo tool_use.input exactly as received; normalize a copy for execution only
1 reserialized (blocks_modified) Many blocks differ slightly across types: a lossy round trip through the app's own message model (interior whitespace, number formatting, coerced keys, trimmed text) many changed at char N lines across messages from_dict/to_dict, JSON re-encoding of history, .strip() on content Persist and replay the wire JSON; never rebuild messages from domain objects
2 reminder_stripped / history_block_stripped / block_inserted (blocks_removed / blocks_inserted) A per-turn text block injected into a user turn and removed on the next request (or added after the fact) messages[i] (user) content[j] (text) removed / inserted regex strips of reminder tags; inject-then-strip helpers Turn-scoped system message (clear_at: "next_user_message") appended after the tool results, every earlier copy left in place; without the beta, a text block after the tool_result blocks, left in place
2 system_block_rerendered / system_blocks_stripped A mid-conversation role: "system" message re-rendered in place, or several dropped (a sub-agent transcript replayed without them) messages[i] (system) changed / removed transcript stores that don't keep system messages Persist them with the transcript and replay verbatim
3 media_stripped (blocks_removed / blocks_modified) Images or documents in earlier turns dropped, downsized, or replaced by a placeholder - a client media cap content[j] (image) removed image caps, resizing of stored turns Downscale at ingestion; return images inside the producing tool's tool_result so server-side context editing can clear them; if a user-turn cap is unavoidable, strip deterministically to cap-minus-headroom and accept that each crossing is an edit; file_id only for bytes that would drift
1 image_url_resigned (pattern) / media_content_changed (kind) A URL-sourced image or document whose block changed, or whose bytes differ from the first fetch content[j] (image) changed URL re-signing, re-uploads The check compares the bytes, not the URL string: a rotated URL to the same bytes is fine; for content referenced across turns use a file_id or base64
4 predecessor_missing / predecessor_reordered (kinds of the chain check; the header reads kind=predecessor_missing; pattern=not_applicable) A thinking block removed from the middle, or re-ordered, with the rest of the prefix intact ! messages[i] re-sent with a different set of thinking blocks filters on type == "thinking"; serializers that drop empty fields or unknown block types (a thinking block with empty text is still a block); a hand-rolled stream parser that loses the signature_delta; strip-and-retry without a record Keep the replayed thinking blocks a contiguous window of the original (drop from the front or the back, never the middle); make any forced strip deterministic and recorded so it replays identically
3 unknown with kind blocks_removed or blocks_replaced A shortening of the history that the API does not name more specifically, or several edits at once messages[i..j] removed / replaced by 1 message(s) the same leads as the truncation and compaction rows Read the attribution lines; the fix is the truncation or compaction one above
- foreign_prefix (unrelated) A block from another conversation replayed (on a long conversation; a short one reports an ordinary multiple / system_and_tools_changed) no pair diff (it is a different conversation) session keys, multiplexed stores Fix the session keying
- (no pattern) A drop or 400 on a pair where nothing you sent differs - the diff shows no change and the digests match no diff - Not a harness bug; report the request id to Anthropic
4 thinking_modified (a separate 400, "cannot be modified") A replayed thinking block's text differs from what the API returned - truncated, summarized, re-wrapped ! the thinking text of the block in messages[i] differs stores that trim or reformat thinking text Store and replay thinking blocks byte for byte
0 model_binding_mismatch (model check; no pattern, no header) The conversation was routed to a model that cannot read its earlier thinking - a downgrade, a cheaper-model route, a fallback model_switch=A->B on the pair, verdict unchanged; the probe's model_drops model ids chosen per turn; fallback or router code Not a history edit: send the same messages to every model, keep the field and header in place, and decide the routing - pin the conversation to the producing model, or accept that the older model's turns run without the newer reasoning; report it as reasoning lost by routing
1-2 system_rerendered / tool_set_changed / tool_schema_changed / predecessor_missing, triggered by a switch A re-rendered prompt or tool list that persists past the switch (a fallback banner, a prompt keyed to the last responder), or thinking stripped on the switch; on a downgrade the edit is reported only on the next Claude Fable 5.1 turn. A prompt that is a pure function of the model called is not this break the row for that pattern, on the pair after the switch (the diff also flags a per-model prompt that the API accepts - confirm with the probe) prompts or tool lists that change with the route; thinking filtered on a switch The row for that pattern: each model's prompt and tool text stable across its own turns, changes as an appended role: "system" message, and never strip thinking on a switch

Append-only forms under newer betas

Two newer betas add an append-only form for shapes the table above marks as having none without them. On-demand compaction (compact-2026-09-04) is on the Claude API, Claude Platform on AWS, Google Cloud and Microsoft Foundry, not on Amazon Bedrock; the Compatibility list on its page names the models and platforms, and the Models API reports each model's capabilities.compaction with the beta header. Defining a tool inside a message (inline-tools-2026-09-15) is Claude API only; by-reference tool changes under the older mid-conversation-tool-changes-2026-07-01 header also work on Amazon Bedrock and Google Cloud. Where a beta is not available, treat those shapes as the table says: measure and decide, freeze and replay.

Background and keep-tail compaction, the append-only way: on-demand compaction (beta compact-2026-09-04, "compaction": {"type": "summarize"}; the on-demand compaction page, https://platform.claude.com/docs/en/build-with-claude/compaction-on-demand). Send a compaction request that carries exactly the messages of a request you already sent, on the conversation's model and under its system, tools and thinking settings, with the compaction field and a max_tokens large enough for a summary; send the beta header on it and on every request that carries the block. Exactly those messages, because the kept turns must directly follow the summarized messages and the first kept message must not be one the API would merge into the last summarized one (the same role, or a role: "system" message). A compaction request whose last assistant turn is still waiting on a tool result is rejected, so send the results first. Leave out output_config.format, stop_sequences and a tool_choice of type any or tool (the API rejects a compaction request that carries them), and never send output_config.task_budget.remaining on the compaction request or on any request that carries the block (a 400). The response holds one signed compaction block and nothing else (stop_reason: "compaction"). When no summary could be written there is no block: the response is still a 200 with empty content, and stop_reason is the summarization call's own - max_tokens (cut off), model_context_window_exceeded (no room for the summarization prompt), refusal, tool_use, or end_turn (no text) - so give max_tokens a few thousand tokens at least, resend with more room or fewer messages as the reason suggests, or continue without one. Custom instructions replace the server's summarization prompt whole (a blank value counts as absent), so ask for text only and no tool call yourself; the summarizer reads the whole conversation either way, earlier thinking included (unlike threshold compaction with custom instructions on Claude 5.1 and later models, which leaves earlier thinking out). Keep taking turns against the full history while it runs, and do not edit anything already sent.

On the first request after the block arrives, drop exactly the messages you sent to the compaction request from the front of the history and put the block first, as an assistant message of its own (the request that does this adopts the block; the API also accepts it as the first content block of the first kept message, and prefix_diff.py checks the kept turns only in the message-of-its-own form); everything appended since stays, thinking included, and keeps verifying. Keep nothing else from the dropped messages: summarized messages re-sent after the block are not rejected - the model sees them twice, summary then verbatim - so drop them yourself. Keep the block first on every later request; to compact again, send compaction on a request that starts with the current block, and from then on send only the newest block (a request that carries more than one compaction block is a 400). Do not send compaction and context_management in the same request. Text instructions in role: "system" messages and tool_addition / tool_removal blocks that sat inside the summarized messages stop applying at adoption (tool changes excepted when the block carries tool_changes - next sentence): re-declare them in a role: "system" message on the first request that adopts the block, directly after that request's new user turn (which comes after the kept turns - a system message between the block and the kept turns breaks their thinking), and leave it there afterward. Where the compaction request also carried inline-tools-2026-09-15, the block records the summarized messages' net tool changes in a tool_changes field: send the block back unmodified and they carry over by themselves, so re-declare no tool change; a block without that field carries none, so re-declare as above. The block is accepted on any model that supports the beta, with any later system or tools, but the kept turns' thinking verifies only on a model that can read it, only if every compaction request since that thinking was produced ran on a model with preserved thinking, and only while system and the tools other than defer_loading: true ones stay what the compaction request had: changing either can invalidate the kept turns' thinking and has no other effect, so to change them without losing any, compact the whole conversation first (keeping no turns), then change them on the next request. Nothing before the block is sent, but the kept turns are still checked against the summarized messages as they stood when you sent the compaction request, so do not touch them in between; prefix_diff.py compares them against the earlier request by aligning on a kept thinking block both carry, and when the compaction request itself is in the capture it checks that request's system and tools against the conversation's, compares the adopting request's system and tools against that request, and checks that the dropped range is the messages it carried.

Tool changes by value (beta inline-tools-2026-09-15, Claude API; the Mid-conversation system messages and tool changes page, "Define tools in a message"). Keep tools exactly as the first request sent it, on every request, and make every later change by appending one role: "system" message: a tool_addition whose tool is {"type": "tool_definition", "definition": {...}} with the full tools entry (name, description, input_schema) inside definition, for a new tool or for a same-name tool whose description or schema changed - a different definition replaces the tool from that message on, an identical one changes nothing (safe to resend on a retry) - and a tool_removal by reference to withdraw one. Nothing already sent moves, so earlier thinking keeps verifying; rewording or deleting a definition message already sent is an edit like any other. The header also covers changes by reference, so it replaces mid-conversation-tool-changes-2026-07-01; the placement rules are the same, including no tool change directly after a paused assistant turn. A tool added this way may itself be defer_loading: true (inside definition); a tool already known at the first request belongs in tools with defer_loading: true, shown later by reference. Keep at least one non-deferred tool in tools (a tool search tool counts): otherwise the first tool defined by value costs one full cache miss. cache_control goes on the block or in the definition, not both, and never on a deferred definition. Reusing a name for a different type of tool is a 400, and during the beta some tool types (computer use among them) cannot be defined in a message - declare those in tools and add them by reference. A definition stays in the history after the tool is replaced or withdrawn, so a beta header that a dated tool type needs goes on every later request of the conversation. For an MCP connector server, add mcp-client-2026-09-15 (in place of mcp-client-2025-11-20, which it includes): the definition can then be an mcp_toolset (connection details stay in mcp_servers), and a response for which the API fetched a server's tool list starts with one mcp_tool_listing block per server fetched (code that reads content[0] must skip them) - send the assistant message back as it came, those blocks included, keep the header on every request that carries one, and later requests reuse that list instead of asking the server again.

Keep list - what never to flag

The scan and the diff will tempt you to report things the check does not care about. These stay out of the report (or go in a "checked, fine" line):

  • Reordering tools in the tools array without changing them - compared as a name-keyed set.
  • Adding a defer_loading: true tool that nothing has referenced yet.
  • Adding, moving, or removing cache_control markers anywhere.
  • Changing model, max_tokens, temperature and other sampling parameters, tool_choice, metadata, stop_sequences, stream, service_tier, output_config (including effort), or the thinking configuration itself - none is part of the compared prefix today (a model change triggers the separate model check, reported as model_binding_mismatch, not this one). Request headers are not part of it either, except that a beta which makes the API inject a tool server-side (code execution, the web-search fallback) changes the tool set with an identical body.
  • String content versus a single text block of the same text; leading or trailing whitespace of a text block; whitespace-only text blocks; JSON key order; 1 versus 1.0.
  • A rotated or re-signed URL for an image or document that serves the same bytes.
  • Removing thinking blocks from the start of the history (oldest first), from the end, or all of them - allowed, and the Preserved thinking page documents all three. The kept blocks must stay an unbroken run of the original sequence; it costs the reasoning, not the validity of the kept blocks. An assistant turn whose tool_use still awaits its tool_result should keep its thinking (the page asks for that). Two things are not on this list: removing a block from the middle, which invalidates every block after it, and putting a removed block back, which invalidates the blocks produced while it was gone.
  • Anything the API itself adds or rewrites server-side (server-side compaction, context editing, thinking clearing, its own injected text) - the check runs on the request as you sent it.
  • Mid-conversation role: "system" messages and cleared turn-scoped messages that are left in place and replayed verbatim.

Failure modes to avoid

  • Measuring on a slice that never replays thinking. The first request of every conversation replays nothing; short turns may return no thinking block; a harness that strips thinking has nothing to check. Read the replayed-thinking count before the drop count, every run.
  • Counting entries instead of breaks. A stale block re-fails on every later request. Report conversations with a first break and the turn it happened at; an entries-per-request number only ever goes up with conversation length.
  • Trusting the header over the body. The diagnosis header is best-effort and unpublished; input_transformations is the contract. A drop with no header is still a drop; a header-only pipeline misses every drop where the header did not arrive.
  • Reading a header-only run as an enforcement test. On an organization that is not enforced yet the header alone records failures as thinking_mismatch_allowed and drops nothing: good for finding edits in production, no measure of what enforcement costs. Set prefix_mismatch_behavior for that.
  • Misspelling the field. Only block_binding.prefix_mismatch_behavior is the documented name; the probe removes any other key it finds under block_binding and says which; production code should write the documented name.
  • Reading the capture from the app's own objects. A capture rebuilt from ORM or domain objects hides the re-serialization the check catches. Capture at the HTTP layer.
  • Replaying someone else's conversation. A capture from another organization still gets its blocks dropped, but it is not diagnosed, and the drop teaches nothing about the harness. Replay with a key from the organization that produced the capture.
  • Bundling fixes. Two causes fixed in one diff cannot be attributed or reverted separately. One cause per diff, re-measured each time.
  • Prescribing a compaction rewrite as if it were required. Keep-tail and background compaction have no append-only client-side form without the compact-2026-09-04 beta (on-demand compaction); measure the scheme the product has, choose error or drop_block, and record the decision.
  • Setting drop_block and calling it done. "drop_block" hides the error but doesn't fix the edit that caused it. Dropped blocks aren't billed, but a session's token usage might still increase because Claude can sometimes think more to re-create the dropped thinking. The increase tends to be larger when more thinking blocks are dropped, or when blocks are dropped on more turns of a long session. Use drop_block to measure (Step 2) and as a recorded stopgap, count the responses in each session whose input_transformations has a prefix_binding_mismatch entry, and fix the edit.
  • Resending the refused body, or fixing a broken session per request. The Preserved thinking page's "Handle the error in code": retry once with the beta header and prefix_mismatch_behavior: "drop_block", and store that choice with the session so every later request sends it too, including after a restart; where the beta header cannot be sent, remove every thinking and redacted_thinking block from the history once and leave them out. A saved session that now fails on every request has the edit stored in it: the same remedy applies, thinking produced from then on stays valid as long as nothing before it changes again, and the edit still has to be found so new sessions do not hit it.
  • A library, proxy or gateway that rewrites what it forwards. Its rewrites are edits its users cannot see or fix. Forward the caller's anthropic-beta values and thinking.block_binding unchanged and return input_transformations to them (an options schema that rejects unknown keys stops a caller from choosing "drop_block"); leave a role: "system" message where the caller put it - moving it into the top-level system field invalidates every thinking block in the conversation; turn tool use off with tool_choice: {"type": "none"}, never by removing tools; and do not hide the 400 - code that catches it, strips thinking and retries on the caller's behalf logs that it did.
  • An unrecorded strip. Stripping thinking after a 400 without making the strip deterministic and recorded re-sends the refused blocks on the next turn and fails again, every turn, for the rest of the conversation.
  • Leaving the production value unset. Defaults differ by surface and by account age; an unset field on a not-yet-enforced account means the check is only recorded (thinking_mismatch_allowed), and that default changes the day the account or the model is enforced. Set it, and monitor the entries or the 400s.
  • Mistaking the model check for this one. model_binding_mismatch entries after a model switch are expected and unbilled; only prefix_binding_mismatch is a harness finding.

Source: SKILL.md on GitHub

1 warning2d5 checks · Risk SAFE
  • Gen Agent Trust Hub2d

    This skill is a developer reference for the Claude API and Anthropic SDKs. It includes some security considerations related to building agents with powerful capabilities like shell command execution and web fetching. While these present a potential surface for indirect prompt injection, the skill provides extensive security guidance, emphasizing sandboxing and input validation as mitigation strategies. All external resources and packages originate from trusted official sources.

  • Socket2d

    No alerts

  • Snyk2d

    Risk: LOW · No issues

  • Runlayer7mo

    12/26 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 8a1541c. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 3 days ago.

Activeupdated 3 days ago

README badge

README badge for anthropics/skills/claude-api

Reference for the Claude API and official Anthropic SDKs — model IDs, pricing, parameters, streaming, tool use, MCP, managed agents, caching, token counting, and model migration. Read this skill before opening a file that involves Claude, an Anthropic model, agent workflows, or LLM-shaped tasks with no specified provider.

Generated from the current SKILL.md.

Which Claude model should I use by default?
Use Claude Opus 4.8 (model ID: `claude-opus-4-8`) as the default. Also default to adaptive thinking (`thinking: {type: "adaptive"}`) for anything complex, and streaming for requests with long input, output, or high max_tokens.
What should I do if the project uses OpenAI or another non-Anthropic provider?
Stop and ask the user whether they want to switch the file to Claude or want a non-Claude implementation. Do not edit a non-Anthropic file with Anthropic SDK calls.
Should I use the official SDK or raw HTTP?
Use the official Anthropic SDK for your language whenever one exists (Python, TypeScript, Java, Go, Ruby, C#, PHP). Only use raw HTTP (curl, requests, fetch) if the user explicitly asks for it, the project is shell/cURL, or the language has no official SDK.
When should I use Managed Agents versus Claude API with tool use?
Use Managed Agents when you want Anthropic to run the agent loop and host a per-session container for tool execution (file ops, bash, code). Use Claude API with tool use for multi-step workflows where you control the orchestration and host the compute yourself.
Does this skill work with Amazon Bedrock, Google Vertex AI, or Microsoft Foundry?
Managed Agents is not available on those platforms. Use Claude API with tool use instead. Claude Platform on AWS (Anthropic-operated) has full feature parity with the first-party API.

Generated from the current SKILL.md. These answers refresh after source changes.