All skills
aws avatar

/amazon-bedrock

@3b23681

Builds generative AI applications on Amazon Bedrock. Covers model invocation (Converse API, InvokeModel), RAG with Knowledge Bases, Bedrock Agents, Guardrails, and AgentCore (including the Harness managed agent loop). Applies when invoking models, setting up Knowledge Bases, creating agents, applying guardrails, deploying to AgentCore, migrating/porting/converting a Bedrock Agent (including inline agents) to an AgentCore Harness, troubleshooting Bedrock errors (ThrottlingException, AccessDeniedException), or choosing models (Claude, Llama, Nova, Titan). Also for prompt caching, quota and throttling diagnosis, cost tracking, migrating between Claude model generations (4.5 to 4.6 to 4.7), chunking strategies, API selection (Converse vs InvokeModel), guardrail capabilities, and model selection. Also covers AgentCore Payments (x402, microtransactions, Payment Manager, Connector, Instrument, Coinbase CDP, Stripe Privy, paid endpoints, agent payments). NOT for custom model training, Rekognition, or Comprehend.

Use this Skill: https://skilld.dev/gh/aws/agent-toolkit-for-aws/amazon-bedrock

This session only. Nothing lands on disk.

referencesquota-health.md

≈1k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Bedrock Quota Health Check

Monitor and manage Bedrock model quotas to prevent throttling. Bedrock enforces two quota types per model per region: requests per minute (RPM) and tokens per minute (TPM).

Table of Contents

How Quota Reservation Works

Bedrock reserves TPM quota at request start based on: InputTokens + CacheWriteInputTokens + CacheReadInputTokens + maxTokens. If maxTokens is unset, it defaults to the model's maximum (up to 64K–128K), reserving far more quota than needed.

Example (Claude Sonnet, 2M TPM quota):

  • maxTokens=1000, 500 input tokens: reserves 1,500 → ~1,333 concurrent requests
  • maxTokens unset (defaults to 64K): reserves ~64,500 → ~31 concurrent requests

This is the most common cause of unexpected ThrottlingException. Always set maxTokens explicitly.

Cache read tokens are included in the initial reservation but released at settlement — prompt caching effectively increases your usable TPM capacity.

Audit Workflow

1. Check Current Quotas

aws service-quotas list-service-quotas --service-code bedrock --region <REGION> --profile <PROFILE> --query "Quotas[?starts_with(QuotaName, 'Invoke')].{Name:QuotaName, Value:Value}" --output table

2. Check Recent Usage vs Limits

Run the quota health script:

python3 scripts/check-quota-health.py --region <REGION> --profile <PROFILE>

The script compares current quota limits against peak CloudWatch metrics over the last 24 hours and flags models approaching their limits.

3. Assess maxTokens Impact

Review application code for Bedrock calls without explicit maxTokens. Each unset call wastes quota proportional to the model's max output tokens.

CloudWatch Metrics

Key metrics in the AWS/Bedrock namespace (dimension: ModelId):

Metric What It Tells You
InvocationCount RPM usage — compare against RPM quota
InvocationThrottles Throttled requests — any value > 0 needs attention
InputTokenCount Input token consumption per request
OutputTokenCount Actual output tokens — use to right-size maxTokens
InvocationLatency Latency distribution — spikes may correlate with throttling

Sample CloudWatch Logs Insights query (requires model invocation logging enabled):

fields @timestamp, @message
| filter modelId like /claude/
| stats count() as requests, sum(inputTokenCount) as totalInput, sum(outputTokenCount) as totalOutput by bin(1m)
| sort @timestamp desc

When You're Being Throttled

Decision table for resolving ThrottlingException:

Situation Action
maxTokens not explicitly set Set it to expected output length — biggest single impact
Traffic is bursty Use cross-region inference profiles (us., eu., global. prefix) to distribute across regions
Steady-state traffic exceeds quota Request a quota increase (see below)
Latency-sensitive workload Use priority service tier for preferential processing
Non-time-critical workload Use flex service tier (may queue during peak, lower cost)
Consistent high-volume Request quota increase + use cross-region inference for headroom

Quota Increase Requests

aws service-quotas request-service-quota-increase --service-code bedrock --quota-code <QUOTA_CODE> --desired-value <VALUE> --region <REGION> --profile <PROFILE>

To find the quota code for a specific model:

aws service-quotas list-service-quotas --service-code bedrock --region <REGION> --profile <PROFILE> --query "Quotas[?contains(QuotaName, '<MODEL_NAME>')].{Code:QuotaCode, Name:QuotaName, Value:Value}"

Quota increases are reviewed by AWS — plan 1–3 business days. For urgent production needs, open an AWS Support case.

Source: SKILL.md on GitHub

1 warning2d3 checks · Risk SAFE
  • Gen Agent Trust Hub2d

    This skill provides a comprehensive and secure framework for building generative AI applications on Amazon Bedrock. It incorporates industry-standard security practices, including IAM least-privilege guidance, SSRF protections, and robust encryption recommendations for sensitive data.

  • Socket2d

    1 alert: gptSecurity

  • Snyk2d

    Risk: LOW · No issues

Signed by skilld at 3b23681. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 3 days ago
metadata
{
  "version": "6"
}

README badge

README badge for aws/agent-toolkit-for-aws/amazon-bedrock