All skills
aws avatar

/amazon-bedrock

@3b23681

Builds generative AI applications on Amazon Bedrock. Covers model invocation (Converse API, InvokeModel), RAG with Knowledge Bases, Bedrock Agents, Guardrails, and AgentCore (including the Harness managed agent loop). Applies when invoking models, setting up Knowledge Bases, creating agents, applying guardrails, deploying to AgentCore, migrating/porting/converting a Bedrock Agent (including inline agents) to an AgentCore Harness, troubleshooting Bedrock errors (ThrottlingException, AccessDeniedException), or choosing models (Claude, Llama, Nova, Titan). Also for prompt caching, quota and throttling diagnosis, cost tracking, migrating between Claude model generations (4.5 to 4.6 to 4.7), chunking strategies, API selection (Converse vs InvokeModel), guardrail capabilities, and model selection. Also covers AgentCore Payments (x402, microtransactions, Payment Manager, Connector, Instrument, Coinbase CDP, Stripe Privy, paid endpoints, agent payments). NOT for custom model training, Rekognition, or Comprehend.

Use this Skill: https://skilld.dev/gh/aws/agent-toolkit-for-aws/amazon-bedrock

This session only. Nothing lands on disk.

referencescost-tracking.md

≈1.1k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Bedrock Cost Attribution and Tracking

Track, allocate, and manage Bedrock inference costs across teams, products, and models. Bedrock charges per input/output token with model-specific rates.

Table of Contents

Cost Attribution Approaches

Approach Best For Setup Effort
Application inference profiles + cost allocation tags Per-product or per-team cost tracking in Cost Explorer Medium — create profiles, tag, activate in Billing
IAM principal-based (CUR 2.0) Per-developer or per-role attribution Low — automatic in CUR 2.0, no Bedrock config needed
Model invocation logging + custom analytics Fine-grained per-request analysis (token counts, latency, model) High — enable logging, build queries

For most teams, application inference profiles with cost allocation tags is the recommended approach. It provides clean cost breakdowns in Cost Explorer without custom analytics.

Application Inference Profiles

Setup Workflow

1. Create an Application Inference Profile
aws bedrock create-inference-profile \
  --inference-profile-name "<TEAM_OR_PRODUCT_NAME>" \
  --model-source "copyFrom=arn:aws:bedrock:<REGION>::foundation-model/<MODEL_ID>" \
  --region <REGION> --profile <PROFILE>

Note the returned inferenceProfileArn.

2. Tag the Profile
aws bedrock tag-resource \
  --resource-arn <INFERENCE_PROFILE_ARN> \
  --tags key=CostCenter,value=<COST_CENTER> key=Project,value=<PROJECT> \
  --region <REGION> --profile <PROFILE>
3. Activate Cost Allocation Tags

In the AWS Billing console (or via API), activate the tags as cost allocation tags. Tags take ~24 hours to appear in Cost Explorer after activation.

4. Use the Profile for Inference

Replace the base model ID with the inference profile ARN in application code:

response = bedrock_runtime.converse(
    modelId="<INFERENCE_PROFILE_ARN>",
    messages=[...],
    inferenceConfig={"maxTokens": 1024}
)
5. Verify in Cost Explorer

After 24–48 hours, filter Cost Explorer by the tag keys. Bedrock costs appear under Amazon Bedrock service, grouped by tag values.

IAM Principal-Based Attribution

CUR 2.0 automatically records the IAM caller identity for every Bedrock API call. No Bedrock-specific setup required.

To use: tag IAM roles/users with keys like department, costCenter, or project, then filter CUR 2.0 data by those tags. Works for per-developer tracking when each developer assumes a distinct IAM role.

Limitation: only tracks who made the call, not which product or feature triggered it. Use inference profiles for product-level attribution.

CloudWatch Usage Monitoring

Key metrics for cost monitoring (namespace AWS/Bedrock, dimension ModelId):

Metric Cost Signal
InputTokenCount Input token spend (charged per token)
OutputTokenCount Output token spend (higher per-token rate)
InvocationCount Request volume
CacheReadInputTokens Tokens served from cache (90% cheaper than standard input)
CacheWriteInputTokens Cache write tokens (25% surcharge over standard input)

Cost Analysis Script

python3 scripts/analyze-bedrock-costs.py --days <DAYS> --region <REGION> --profile <PROFILE>

The script queries Cost Explorer for Bedrock spend grouped by usage type (model + token direction) over the specified period.

Budget Alerts

Set up AWS Budgets to alert when Bedrock spend approaches a threshold:

aws budgets create-budget --account-id <ACCOUNT_ID> \
  --budget '{"BudgetName":"bedrock-monthly","BudgetLimit":{"Amount":"<AMOUNT>","Unit":"USD"},"TimeUnit":"MONTHLY","BudgetType":"COST","CostFilters":{"Service":["Amazon Bedrock"]}}' \
  --notifications-with-subscribers '[{"Notification":{"NotificationType":"ACTUAL","ComparisonOperator":"GREATER_THAN","Threshold":80},"Subscribers":[{"SubscriptionType":"EMAIL","Address":"<EMAIL>"}]}]' \
  --profile <PROFILE>

This alerts at 80% of the monthly budget. Adjust threshold and notification targets as needed.

Source: SKILL.md on GitHub

1 warning2d3 checks · Risk SAFE
  • Gen Agent Trust Hub2d

    This skill provides a comprehensive and secure framework for building generative AI applications on Amazon Bedrock. It incorporates industry-standard security practices, including IAM least-privilege guidance, SSRF protections, and robust encryption recommendations for sensitive data.

  • Socket2d

    1 alert: gptSecurity

  • Snyk2d

    Risk: LOW · No issues

Signed by skilld at 3b23681. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 3 days ago
metadata
{
  "version": "6"
}

README badge

README badge for aws/agent-toolkit-for-aws/amazon-bedrock