AI/LLM API Design Patterns
Design patterns for APIs that expose AI/LLM capabilities — streaming, tool use, structured output, and safety.
2026-08 baseline:
- OpenAPI 3.2 (2025-09-23) gives streaming endpoints a first-class contract via
text/event-stream+itemSchema,application/jsonl, andapplication/json-seq(spec). Document SSE response events withitemSchemainstead of prose.- OpenAI Structured Outputs with
strict: trueconstrains supported models to a JSON Schema (OpenAI API reference). This is a syntax guarantee only: a fully conformant object can still carry invented field values. Strict-mode constraints includeadditionalProperties: falseand all properties inrequired(use anullunion for optional fields). Function calling supports the samestrictflag.- Anthropic Prompt Caching (GA) cuts long-prompt input cost up to 90% and latency up to 85%; cache hits are 0.1× input price, 5-min TTL default (1-hour optional). Anthropic now automatically identifies cached segments — manual
cache_controlmarkers are still supported but no longer required for many cases (Anthropic docs).- OWASP Top 10 for Agentic Applications 2026 (released 2025-12, genai.owasp.org) — ASI01 Agent Goal Hijacking is the #1 risk for agent-facing APIs. Apply principle-of-least-agency on every tool exposed via function-calling.
Example model IDs are placeholders: resolve a supported, authorized stable ID through
_common/CLI_COMPATIBILITY.mdbefore executing.
Streaming Response Pattern (SSE)
Server-Sent Events (SSE) is the standard for streaming LLM token output to clients.
POST /v1/responses
Content-Type: application/json
Accept: text/event-stream
{
"model": "<authorized-model-id>",
"stream": true,
"input": [{"role": "user", "content": "Hello"}]
}HTTP/1.1 200 OK
Content-Type: text/event-stream
Cache-Control: no-cache
X-Accel-Buffering: no
data: {"type":"response.output_text.delta","item_id":"msg_01","output_index":0,"content_index":0,"delta":"Hello","sequence_number":1}
data: {"type":"response.output_text.delta","item_id":"msg_01","output_index":0,"content_index":0,"delta":"!","sequence_number":2}
data: {"type":"response.completed","response":{"id":"resp_01","status":"completed"},"sequence_number":3}Streaming Design Rules
| Rule | Description |
|---|---|
text/event-stream content type |
Always set Content-Type: text/event-stream and Cache-Control: no-cache |
| Terminate explicitly | End with the provider's documented completion event (for example response.completed); do not infer completion from a closed socket |
| Include event IDs | id: field on each event enables client-side reconnect with Last-Event-ID |
| Heartbeat events | Send : keep-alive comment lines every 15s to prevent proxy/load balancer timeouts |
| Error mid-stream | Send data: {"type":"error","error":{"type":"server_error","message":"..."}} then close stream |
| OpenAPI spec | Document streaming endpoints with text/event-stream response type and link to SSE event schema |
Tool Use / Function Calling
Tool use allows the model to request structured data from external systems during generation.
Schema Design
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get current weather for a location. Use when the user asks about weather conditions.",
"parameters": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "City name or 'city, country' format. Example: 'Tokyo, JP'"
},
"unit": {
"type": "string",
"enum": ["celsius", "fahrenheit"],
"description": "Temperature unit. Default is celsius."
}
},
"required": ["location"]
}
}
}Best Practices
- Descriptions are prompts — write
descriptionfields as instructions to the model, not documentation for humans. Specify when to call the tool and how to interpret the result. - Narrow parameter schemas — use
enum,minimum,maximum,patternto constrain valid inputs. Broad schemas lead to hallucinated parameter values. - Idempotent tools first — prefer read-only tools; flag state-mutating tools with
"confirm_required": true(custom extension) to prompt user confirmation before execution. - Return structured data — tool results should be JSON, not prose. The model parses the result; unstructured text increases hallucination risk.
- Design for parallel calls — models may invoke multiple tools simultaneously; tool implementations must be safe to run concurrently.
- Namespace names by service (
github_,slack_) and put user-facing keywords in descriptions — once a catalog is searched rather than fully loaded, discoverability is the naming and description quality. - Design for a searched catalog past ~10 tools — tool-selection accuracy degrades past 30-50 available tools. Anthropic's tool search (
tool_search_tool_regex_20251119/_bm25_20251119) plusdefer_loading: trueloads only the tools a request needs. Note thatdefer_loadingcontrols context entry, not the wire: you still transmit every definition each request. Full contract, versioning semantics, and per-tool model support →oracle/reference/advanced-tool-use.md.
Structured Output
Forces the model to produce JSON that conforms to a specified schema. On supported OpenAI models, strict: true constrains decoding to the supported JSON Schema subset.
Schema conformance is not correctness. Decode-time constraints guarantee shape, never meaning — a schema-valid record can name a price or a brand that appears nowhere in the input. Value correctness, evidence-span existence, and execution authorization are separate validators that run after parsing, not properties the schema buys. See oracle/reference/evaluation-observability.md ("schema validity says nothing about answer quality").
Strict-Mode Structured Output (OpenAI Responses API)
{
"model": "<authorized-model-id>",
"input": [
{ "role": "user", "content": "Extract product fields from: ..." }
],
"text": {
"format": {
"type": "json_schema",
"name": "extract_product",
"strict": true,
"schema": {
"type": "object",
"additionalProperties": false,
"required": ["name", "price", "sku"],
"properties": {
"name": { "type": "string" },
"price": { "type": "number" },
"sku": { "type": ["string", "null"] }
}
}
}
}
}Strict-mode constraints (enforced at request validation):
additionalProperties: falseon every object.- All properties listed in
required; mark optional fields by addingnullto the type array. - Subset of JSON Schema 2020-12 (no
oneOfwith conflicting types, no recursive$refcycles).
Legacy JSON Mode (still useful for non-OpenAI / older models)
{
"model": "claude-opus-5",
"response_format": { "type": "json_object" },
"messages": [
{ "role": "user", "content": "Extract product name, price, SKU from: ..." }
]
}Structured Output Rules
| Rule | Description |
|---|---|
Prefer strict: true when available |
Use the provider's current structured-output mechanism and verify model support; Anthropic tool use provides analogous schema constraints. |
| Provide schema in prompt | Even with strict mode, include the target schema in the system prompt — models still hallucinate field semantics without context |
| Validate server-side | Never trust model output as schema-valid — always parse and validate with Zod/Pydantic before passing downstream (defense in depth, even with strict mode) |
| Handle partial JSON | Streaming structured output may arrive as partial JSON; buffer and parse only after the provider's documented completion event |
| Version your schemas | Include a schema_version field in output schemas; models may produce output with old field names after schema changes |
Rate Limiting for AI APIs
AI APIs have unique cost dimensions that require multi-axis rate limiting.
| Dimension | Unit | Why It Matters |
|---|---|---|
| Requests per minute (RPM) | Count | Prevents API abuse and thundering herd |
| Input tokens per minute (TPM) | Tokens | Direct cost driver — long prompts consume quota fast |
| Output tokens per minute (TPM) | Tokens | Streaming output is billed per output token; unconstrained streaming can exhaust quota |
| Concurrent streams | Count | SSE connections hold server resources; limit per user/org |
HTTP/1.1 429 Too Many Requests
Content-Type: application/json
Retry-After: 30
X-RateLimit-Limit-Requests: 60
X-RateLimit-Remaining-Requests: 0
X-RateLimit-Limit-Tokens: 100000
X-RateLimit-Remaining-Tokens: 0
X-RateLimit-Reset-Requests: 2026-04-01T12:00:30Z
X-RateLimit-Reset-Tokens: 2026-04-01T12:00:05Z
{
"error": {
"type": "rate_limit_error",
"message": "Token quota exceeded. Retry after 30 seconds."
}
}Error Handling
AI API errors require distinct handling from standard REST errors due to partial streaming and model-specific failure modes.
Error Response Format
{
"error": {
"type": "invalid_request_error",
"code": "context_length_exceeded",
"message": "Input tokens (128500) exceed model maximum (128000). Reduce prompt length.",
"param": "messages"
}
}Error Handling Rules
- Distinguish error types:
invalid_request_error(4xx, fix the request),authentication_error(401, check API key),rate_limit_error(429, backoff),api_error(5xx, retry with exponential backoff). - Handle stream interruption: If a stream stops without the documented completion event, treat it as
api_error— do not present partial output as complete to the user. - Content filter errors:
content_policy_violationerrors should be surfaced to users with a user-friendly message; do not silently retry — log for safety review. - Token budget errors: Return
context_length_exceededwith the actual and maximum token counts to help clients truncate inputs correctly.