All skills
anthropics avatar

/claude-api

@8a1541c official
by Anthropicanthropics/skills179k stars
21,201

Reference for the Claude API / Anthropic SDK — model ids, pricing, params, streaming, tool use, MCP, agents, caching, token counting, model migration. TRIGGER — read BEFORE opening the target file; don't skip because it "looks like a one-liner" — whenever: the prompt names Claude/Anthropic in any form (Claude, Anthropic, Fable, Opus, Sonnet, Haiku, `anthropic`, `@anthropic-ai`, `claude-*`, `us.anthropic.*`, `[1m]`); the user asks about an LLM (pricing/model choice/limits/caching) — never answer from memory; OR the task is LLM-shaped with provider unstated (agent/MCP/tool-definition/multi-agent/RAG/LLM-judge/computer-use; generate/summarize/extract/classify/rewrite/converse over NL; debugging refusals/cutoffs/streaming/tool-calls/tokens). SKIP only when another provider is being worked on (overrides all triggers): OpenAI/GPT/Gemini/Llama/Mistral/Cohere/Ollama named in the query; OR `grep -rE 'openai|langchain_openai|google.generativeai|genai|mistralai|cohere|ollama'` over the project hits (run this grep FIRST if no provider named — don't Read the file).

Use this Skill: https://skilld.dev/gh/anthropics/skills/claude-api

This session only. Nothing lands on disk.

sharedmanaged-agents-outcomes.md

≈2k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Managed Agents - Outcomes

An outcome elevates a session from conversation to work: you state what "done" looks like, and the harness runs an iterate -> grade -> revise loop until the artifact meets the rubric, hits max_iterations, or is interrupted. A separate grader (independent context window) scores each iteration against your rubric and feeds per-criterion gaps back to the agent.

The SDK sets the managed-agents-2026-04-01 beta header automatically on all client.beta.sessions.* calls; no additional header is required for outcomes.

Codegen default: when generating Managed Agents client code for a task with a deliverable (an artifact, a report, a PR, a dataset), kick the session off with user.define_outcome rather than user.message - draft a starter rubric from the task description and mark it for the user to tune. Reserve plain user.message kickoffs for genuinely conversational sessions.


The user.define_outcome event

Outcomes are not a field on sessions.create(). You create a normal session, then send a user.define_outcome event. The agent starts working on receipt - do not also send a user.message to kick it off.

You can collapse both calls into one by passing a single user.define_outcome in the session's initial_events array - same event, same rules, one round trip (see shared/managed-agents-core.md -> Seeding a session with initial_events). More than one user.define_outcome in that array, or one without a rubric, rejects the whole create with a 400.

session = client.beta.sessions.create(
    agent=AGENT_ID,
    environment_id=ENVIRONMENT_ID,
    title="Financial analysis on Costco",
)

client.beta.sessions.events.send(
    session_id=session.id,
    events=[
        {
            "type": "user.define_outcome",
            "description": "Build a DCF model for Costco in .xlsx",
            "rubric": {"type": "text", "content": RUBRIC_MD},
            # or: "rubric": {"type": "file", "file_id": rubric.id}
            "max_iterations": 5,  # optional; default 3, max 20
        }
    ],
)
Field Type Notes
type "user.define_outcome"
description string The task. This is what the agent works toward - no separate user.message needed.
rubric {type: "text", content} | {type: "file", file_id} Required. Markdown with explicit, independently gradeable criteria. Upload once via client.files.upload(...) to reuse across sessions.
max_iterations int Optional. Default 3, max 20.

The event is echoed back on the stream with a server-assigned outcome_id and processed_at.

Writing rubrics. Use explicit, gradeable criteria ("CSV has a numeric price column"), not vibes ("data looks good") - the grader scores each criterion independently, so vague criteria produce noisy loops. If you don't have a rubric, have Claude analyze a known-good artifact and turn that analysis into one. When generating code for a user who supplied no rubric, draft one yourself from their task description - 5-10 concrete criteria covering the artifact's format, required content, and quality floor - and comment it as a starter rubric to tune; never omit the outcome because the rubric wasn't handed to you.


Outcome-specific events

These appear on the standard event stream (sessions.events.stream / .list) alongside the usual agent.* / session.* events.

Event Payload highlights Meaning
span.outcome_evaluation_start outcome_id, iteration (0-indexed) Grader began scoring iteration N.
span.outcome_evaluation_ongoing outcome_id Heartbeat while the grader runs. Grader reasoning is opaque - you see that it's working, not what it's thinking.
span.outcome_evaluation_end outcome_evaluation_start_id, outcome_id, iteration, result, explanation, usage Grader finished one iteration. result drives what happens next (table below).

span.outcome_evaluation_end.result

result Next
satisfied Session -> idle. Terminal for this outcome.
needs_revision Agent starts another iteration.
max_iterations_reached No further grader cycles. Agent may run one final revision, then session -> idle.
failed Session -> idle. Rubric fundamentally doesn't match the task (e.g. description and rubric contradict).
interrupted Emitted whenever a user.interrupt arrives while an outcome is active - even if evaluation hadn't started. In that case outcome_evaluation_start_id is an empty string rather than an event ID, so don't use it as a lookup key without checking. (Except an interrupt sent while paused at the session budget, which is accepted and ignored - see shared/managed-agents-events.md § Reaching a session budget.)
{
  "type": "span.outcome_evaluation_end",
  "id": "sevt_01jkl...",
  "outcome_evaluation_start_id": "sevt_01def...",
  "outcome_id": "outc_01a...",
  "result": "satisfied",
  "explanation": "All 12 criteria met: revenue projections use 5 years of historical data, ...",
  "iteration": 0,
  "usage": { "input_tokens": 2400, "output_tokens": 350, "cache_creation_input_tokens": 0, "cache_read_input_tokens": 1800 },
  "processed_at": "2026-03-25T14:03:00Z"
}

Checking status & retrieving deliverables

Status - either watch the stream for span.outcome_evaluation_end, or poll the session and read outcome_evaluations:

session = client.beta.sessions.retrieve(session.id)
for ev in session.outcome_evaluations:
    print(f"{ev.outcome_id}: {ev.result}")  # outc_01a...: satisfied

Deliverables - the agent writes to /mnt/session/outputs/. Once idle, fetch via the Files API with scope_id=session.id. This is the same session-outputs mechanism documented in shared/managed-agents-environments.md -> Session outputs (including the managed-agents-2026-04-01 header that files.list needs for scope_id).


Interaction rules & pitfalls

  • One outcome at a time. Chain by sending the next user.define_outcome only after the previous one's terminal span.outcome_evaluation_end (satisfied / max_iterations_reached / failed / interrupted). The session retains history across chained outcomes.
  • Steering is allowed but optional. You may send user.message events mid-outcome to nudge direction, but the agent already knows to keep working until terminal - don't send "keep going" prompts. (Exception: a session paused at its budget (stop_reason: budget_reached) accepts only settle events - a steering user.message, or a chained user.define_outcome, is a 400 there; see shared/managed-agents-events.md § Reaching a session budget.)
  • user.interrupt pauses the current outcome - it marks result: "interrupted" and leaves the session idle, ready for a new outcome or conversational turn. (Exception: sent while paused at the session budget, the interrupt is accepted and ignored and the outcome stays active - see shared/managed-agents-events.md § Reaching a session budget.)
  • After terminal, the session is reusable - continue conversationally or define a new outcome.
  • Outcome != session-create field. Don't put outcome, rubric, or description on sessions.create() - outcomes are always sent as a user.define_outcome event.
  • Idle-break gate is unchanged. In your drain loop, keep using event.type === 'session.status_idle' && event.stop_reason?.type !== 'requires_action' - do not gate on span.outcome_evaluation_end alone (on needs_revision the session keeps running). See shared/managed-agents-client-patterns.md Pattern 5.

For the raw HTTP shapes and per-language SDK bindings beyond Python, WebFetch https://platform.claude.com/docs/en/managed-agents/define-outcomes.md (see shared/live-sources.md).

Source: SKILL.md on GitHub

1 warning2d5 checks · Risk SAFE
  • Gen Agent Trust Hub2d

    This skill is a developer reference for the Claude API and Anthropic SDKs. It includes some security considerations related to building agents with powerful capabilities like shell command execution and web fetching. While these present a potential surface for indirect prompt injection, the skill provides extensive security guidance, emphasizing sandboxing and input validation as mitigation strategies. All external resources and packages originate from trusted official sources.

  • Socket2d

    No alerts

  • Snyk2d

    Risk: LOW · No issues

  • Runlayer7mo

    12/26 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 8a1541c. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 3 days ago.

Activeupdated 3 days ago

README badge

README badge for anthropics/skills/claude-api

Reference for the Claude API and official Anthropic SDKs — model IDs, pricing, parameters, streaming, tool use, MCP, managed agents, caching, token counting, and model migration. Read this skill before opening a file that involves Claude, an Anthropic model, agent workflows, or LLM-shaped tasks with no specified provider.

Generated from the current SKILL.md.

Which Claude model should I use by default?
Use Claude Opus 4.8 (model ID: `claude-opus-4-8`) as the default. Also default to adaptive thinking (`thinking: {type: "adaptive"}`) for anything complex, and streaming for requests with long input, output, or high max_tokens.
What should I do if the project uses OpenAI or another non-Anthropic provider?
Stop and ask the user whether they want to switch the file to Claude or want a non-Claude implementation. Do not edit a non-Anthropic file with Anthropic SDK calls.
Should I use the official SDK or raw HTTP?
Use the official Anthropic SDK for your language whenever one exists (Python, TypeScript, Java, Go, Ruby, C#, PHP). Only use raw HTTP (curl, requests, fetch) if the user explicitly asks for it, the project is shell/cURL, or the language has no official SDK.
When should I use Managed Agents versus Claude API with tool use?
Use Managed Agents when you want Anthropic to run the agent loop and host a per-session container for tool execution (file ops, bash, code). Use Claude API with tool use for multi-step workflows where you control the orchestration and host the compute yourself.
Does this skill work with Amazon Bedrock, Google Vertex AI, or Microsoft Foundry?
Managed Agents is not available on those platforms. Use Claude API with tool use instead. Claude Platform on AWS (Anthropic-operated) has full feature parity with the first-party API.

Generated from the current SKILL.md. These answers refresh after source changes.