All skills
zxkane avatar

/aws-agentic-ai

@e4ef2e2
by Mengxin Zhuzxkane/aws-skills365 stars
40

AWS Bedrock AgentCore comprehensive expert for deploying and managing AI agents at scale. Use when working with any AgentCore service including Gateway, Runtime, Memory, Identity, Code Interpreter, Browser, Observability, Agent Registry, or Evaluations. Covers agent deployment, MCP tool integration, credential management, agent discovery, governance workflows, and automated quality assessment. Essential when user mentions AgentCore, agent runtime, agent registry, agent evaluation, MCP gateway, deploy agent, register MCP server, discover agents, evaluate agent quality, agent credentials, or wants to build, deploy, catalog, or monitor AI agents on AWS.

Use this Skill: https://skilld.dev/gh/zxkane/aws-skills/aws-agentic-ai

This session only. Nothing lands on disk.

referencesagentcore-runtime-deploy.md

≈9.2k tokens on demand. Your agent reads this file only when SKILL.md points to it.

AgentCore Runtime Deployment & Operations: CDK, Security, Observability, Framework Comparison

Covers CDK deployment (L1/L2), multi-Runtime architecture, runtime data flow, security model (OAuth/IAM), observability (OTel/CloudWatch), and BedrockAgentCoreApp vs FastAPI comparison.

Companion document: Core Mechanisms (container contract, Session model, Agent lifecycle, tool integration, startup flow)


Table of Contents


1. CDK Deployment

Use CDK for all production deployments. AgentCore Runtime depends on peripheral resources such as Cognito, ECR, IAM Roles, and Secrets Manager, all of which need to be orchestrated through CDK. AgentCore provides the following deployment paths:

  • CDK L2 Construct (Recommended): @aws-cdk/aws-bedrock-agentcore-alpha, supports four artifact sources — fromAsset (points to a local directory, CDK automatically builds the image), fromEcrRepository (existing ECR image), fromS3 (zip package, no Docker required), fromImageUri (pre-built image URI).
  • CDK L1 + Custom Container: Write your own Dockerfile, push to ECR, and deploy via CfnRuntime. Suitable when you need full control over the build process or integration into an existing IaC pipeline.
  • Starter Toolkit CLI: agentcore configure + agentcore deploy, uses Direct Code Deploy by default, no Dockerfile required. Only suitable for demos and prototyping — production projects cannot deploy just a Runtime without managing peripheral resources.

1.1 CfnRuntime Definition

AgentCore Runtime uses the L1 CDK construct CfnRuntime:

const cfnRuntime = new agentcore_cfn.CfnRuntime(this, 'Runtime', {
  // Runtime name
  agentRuntimeName: 'my-agent',

  // Container image (ECR ARM64)
  agentRuntimeArtifact: {
    containerConfiguration: {
      containerUri: ecrImage.imageUri,
    },
  },

  // Network: PUBLIC mode
  networkConfiguration: { networkMode: 'PUBLIC' },

  // Protocol: HTTP
  protocolConfiguration: 'HTTP',

  // IAM role
  roleArn: agentCoreRole.roleArn,

  // Authentication: Cognito JWT
  authorizerConfiguration: {
    customJwtAuthorizer: {
      discoveryUrl: `${cognitoIssuer}/.well-known/openid-configuration`,
      allowedClients: [cognitoClientId],
    },
  },

  // Environment variables (customize as needed)
  environmentVariables: {
    MY_API_URL: apiUrl,                        // Application-specific
    AGENTCORE_MEMORY_ID: memoryId,             // AgentCore Memory integration
    AGENT_OBSERVABILITY_ENABLED: 'true',       // AgentCore Observability
    // ...
  },
});

1.2 CDK L2 Construct (Recommended)

@aws-cdk/aws-bedrock-agentcore-alpha provides a more concise L2 API with four artifact sources:

import * as agentcore from '@aws-cdk/aws-bedrock-agentcore-alpha';
import * as path from 'path';

// Option A: fromAsset — points to a local directory, CDK automatically builds and pushes the Docker image to ECR
const runtime = new agentcore.Runtime(this, 'MyAgent', {
  runtimeName: 'my-agent',
  agentRuntimeArtifact: agentcore.AgentRuntimeArtifact.fromAsset(
    path.join(__dirname, '../agent')   // Directory must contain a Dockerfile
  ),
});

// Option B: fromS3 — zip package deployment, no Docker required
const runtime = new agentcore.Runtime(this, 'MyAgent', {
  runtimeName: 'my-agent',
  agentRuntimeArtifact: agentcore.AgentRuntimeArtifact.fromS3(
    { bucketName: 'my-code-bucket', objectKey: 'agent.zip' },
    agentcore.AgentCoreRuntime.PYTHON_3_12,
    ['opentelemetry-instrument', 'main.py'],  // Startup command
  ),
});

// Option C: fromEcrRepository — pre-built image from CI/CD pipeline
const runtime = new agentcore.Runtime(this, 'MyAgent', {
  runtimeName: 'my-agent',
  agentRuntimeArtifact: agentcore.AgentRuntimeArtifact.fromEcrRepository(
    repository, 'v1.0.0'
  ),
});

// Common configuration (applies to all options)
runtime.addEndpoint('production', { version: '1' });  // Fixed-version endpoint
model.grantInvoke(runtime);                            // Grant permission to invoke Bedrock model
runtime.grantInvoke(invokerFunction);                  // Grant Lambda permission to invoke this Runtime

The L2 Construct automatically handles IAM role creation, version management, and endpoint configuration. fromAsset is ideal for local development (CDK auto-builds and pushes), fromS3 suits quick deployments without Docker, and fromEcrRepository is best for CI/CD pipelines.

1.3 Docker Image Build

# Multi-stage build, ARM64
FROM --platform=linux/arm64 ghcr.io/astral-sh/uv:python3.12-bookworm-slim AS builder
WORKDIR /app
COPY pyproject.toml uv.lock ./

# UV_PROJECT_ENVIRONMENT specifies the venv location
ENV UV_PROJECT_ENVIRONMENT=/app/.venv
RUN uv venv /app/.venv && \
    uv sync --frozen --no-dev --no-cache && \
    test -f /app/.venv/bin/uvicorn || (echo "ERROR: uvicorn not found!" && exit 1)

FROM --platform=linux/arm64 ghcr.io/astral-sh/uv:python3.12-bookworm-slim

# Security patches
RUN apt-get update && apt-get upgrade -y && rm -rf /var/lib/apt/lists/*

WORKDIR /app
COPY --from=builder /app/.venv ./.venv

ENV PATH="/app/.venv/bin:$PATH"
ENV PYTHONPATH=/app/.venv/lib/python3.12/site-packages

COPY . ./

RUN useradd -m -u 1000 bedrock_agentcore
USER bedrock_agentcore
EXPOSE 8080
CMD python3 -m uvicorn main:app --host 0.0.0.0 --port 8080

Key points:

  • Non-root user: bedrock_agentcore (uid=1000), required by AgentCore security policy
  • ARM64 platform: AgentCore Runtime requires ARM64-compatible images
  • Build verification: test -f uvicorn ensures dependencies are installed correctly, avoiding post-deployment failures
  • uv package manager: UV_PROJECT_ENVIRONMENT specifies the venv location, --frozen ensures reproducible builds

1.4 Deployment Architecture Overview

CDK Deploy
    |
    +-- Cognito UserPool
    |     +-- UserPool + Domain
    |     +-- App Client (Authorization Code + PKCE)
    |     +-- Discovery URL -> AgentCore JWT Authorizer
    |
    +-- ECR Repository
    |     +-- Docker Image (ARM64)
    |
    +-- Secrets Manager (Optional)
    |     +-- API Key / Secrets
    |
    +-- AgentCore CfnRuntime
    |     +-- Container: ECR image
    |     +-- Auth: Cognito JWT Authorizer (discoveryUrl + allowedClients)
    |     +-- Network: PUBLIC / VPC
    |     +-- Env: Environment variables
    |
    +-- Backend Server (API Gateway + Lambda / ECS / EC2)
    |     +-- Business APIs (data queries, writes, etc.)
    |     +-- Auth: API Key / Cognito JWT / IAM
    |     +-- Agent calls via HTTP tools
    |
    +-- (Optional) MCP Gateway
    |     +-- Auth: Cognito M2M (client_credentials)
    |     +-- Gateway Lambda
    |     +-- Target Lambda
    |
    +-- CfnOutput
          +-- backendApiUrl -> Frontend
          +-- agentRuntimeArn -> Frontend

The frontend has two call paths:

  +-------------------------------------------------------------+
  |  Frontend (SPA)                                             |
  |  Cognito login -> access_token                              |
  |                                                             |
  |  Path 1: General business logic (CRUD, lists, config, etc.) |
  |  ------> API Gateway / Backend Server                       |
  |          Auth: Cognito JWT (API Gateway Authorizer)          |
  |                                                             |
  |  Path 2: Agent capabilities (Chat, intelligent Q&A, etc.)   |
  |  ------> AgentCore Runtime endpoint (direct HTTPS)           |
  |          Auth: Cognito JWT (AgentCore JWT Authorizer)         |
  |          Header: Bearer Token + Session-Id                   |
  |          Response: SSE streaming                             |
  +-------------------------------------------------------------+

  Both paths share the same Cognito access_token; the frontend
  does not need to manage two separate authentication flows.

2. Multi-Runtime Architecture Patterns

In complex systems, multiple AgentCore Runtimes are typically deployed, each focusing on different responsibilities. Common division patterns:

Dimension Tool-Intensive Agent Knowledge Query Agent Process-Driven Agent (Skills MD)
Core Capability Execute complex operations with many Python tools RAG retrieval + knowledge base Q&A Markdown-defined Skills drive processes
Tool Definition Python @tool functions Retrieval tools + a few helper tools Markdown files (process instructions + API specs)
MCP Gateway Required (multi-tool routing) Optional Optional
Adding Capabilities Write Python + deploy Update knowledge base documents Add/modify .md files, no redeployment needed
Data Access Direct database / Gateway Knowledge Base / vector database Via Backend API (HTTP tools)
Auth Model Gateway OAuth IAM (Bedrock KB) API Key / service-to-service auth
Tool Count Many (10+) Few (2-3) A few fixed tools + dynamic Skills
Use Cases Complex business workflows (claims, approvals) FAQ, document Q&A, policy consultation Configurable query/operation workflows

2.1 How Multiple Runtimes Communicate

There are several ways for multiple Runtimes to collaborate:

Option A: A2A Protocol (Agent-to-Agent)

AgentCore natively supports the A2A protocol (created by Google, hosted by the Linux Foundation), designed specifically for peer-to-peer Agent collaboration. Each Agent remains a black box, interacting through JSON-RPC 2.0 messages:

Agent A (Runtime 1, HTTP protocol)          Agent B (Runtime 2, A2A protocol)
  |                                           |
  |  1. Via MCPClient or direct HTTP          |
  |     call Agent B's A2A endpoint           |
  |  -------- JSON-RPC message/send --------> |
  |                                           |  2. Agent B processes request
  | <-------- JSON-RPC result --------------- |  3. Returns artifacts
  |                                           |
  |  Agent Card discovery:                    |
  |  GET /.well-known/agent-card.json         |

The core advantage of A2A is Agent opacity — Agents do not need to share internal memory, tool implementations, or private logic; they collaborate entirely through message passing. Suitable for cross-framework (Strands + LangGraph) and cross-organization Agent collaboration.

Option B: Shared Data Layer (Indirect Collaboration)

Multiple Runtimes collaborate indirectly through shared data stores (DynamoDB, S3) without direct communication:

Agent A (Runtime 1)     Agent B (Runtime 2)
  |                       |
  +---> DynamoDB <--------+
        (shared state table)

Suitable for pipeline-style collaboration — Agent A processes and writes to the database, Agent B reads and continues processing.

Option C: Via Orchestration Layer

The frontend or backend API acts as an orchestration layer, calling different Runtimes sequentially based on business logic:

Orchestration Layer (Backend API / Frontend)
  |
  +-- 1. POST Runtime-A/invocations -> Agent A processes
  |      <- Returns intermediate result
  |
  +-- 2. POST Runtime-B/invocations -> Agent B processes
         <- Returns final result

The simplest approach — no direct communication between Agents is required, but orchestration logic resides outside the Agents.

2.2 How to Choose a Communication Pattern

Scenario Recommended Approach Rationale
Agents need real-time conversational collaboration A2A Standard protocol, supports streaming and async
Cross-framework/cross-organization Agent interop A2A Agent Card dynamic discovery, framework-agnostic
Pipeline processing (A completes, then B continues) Shared Data Layer Decoupled, each Agent runs independently
Simple sequential calls Orchestration Layer Simplest, logic centralized in orchestration layer
Agent needs to call another Agent's tools MCP Expose Agent as MCP Server

The three protocols are complementary: MCP is the downward connection for Agents to acquire tools and data, A2A is the horizontal collaboration between Agents, and AG-UI is the upward interaction between Agents and frontend UIs. See Protocol Reference for details.


3. Runtime Data Flow

3.1 Typical Request Flow (Local Tools + HTTP Mode)

User: "Help me check my recent orders"
         |
    Frontend (Web App)
         | POST /invocations
         v
    AgentCore Runtime
    +--------------------------------------------+
    |  Agent Container                           |
    |                                            |
    |  1. Agent selects search_orders tool       |
    |                                            |
    |  2. @tool function calls Backend API       |
    |     -> Auth headers auto-injected          |
    |                                            |
    |  3. Backend API returns data               |
    |                                            |
    |  4. Agent aggregates results, streams out  |
    +--------------------------------------------+
         | SSE Stream
         v
    Frontend renders chat bubbles

3.2 Typical Request Flow (MCP Gateway Mode)

User: "Help me create a Jira Issue"
         |
    Frontend (Web App)
         | POST /invocations
         v
    AgentCore Runtime (MicroVM)
    +------------------------------------------+
    |  Agent Container                         |
    |                                          |
    |  1. Strands Agent selects create_issue   |
    |     tool                                 |
    |                                          |
    |  2. MCPClient calls Gateway              |
    |     -> Gateway routes to Lambda Target   |
    |     -> Gateway auto-injects OAuth token  |
    |                                          |
    |  3. Lambda calls Jira API                |
    |     -> Issue created successfully        |
    |                                          |
    |  4. Result returned to Agent context     |
    |                                          |
    |  5. Agent continues reasoning, streams   |
    |     output                               |
    +------------------------------------------+
         | SSE Stream
         v
    Frontend renders

4. Security Model

4.1 Authentication Layers Overview

Layer 1: AgentCore Identity — Inbound JWT Authorizer
  +- User/Caller -> Runtime / Gateway OAuth 2.0 authentication
  +- OIDC discovery + JWT validation, supports any OAuth 2.0 IdP

Layer 2: AgentCore Identity — Outbound Credential Provider
  +- Agent -> Third-party service OAuth 2.0 / API Key authentication
  +- Token secure storage, automatic refresh; Agent code never touches plaintext credentials

Layer 3: M2M Service API Key (Secrets Manager)
  +- Agent -> Backend API service-to-service authentication
  +- secrets.compare_digest prevents timing attacks
  +- Environment variables store only the Secret ARN; secret values are read from Secrets Manager at runtime

Layer 4: IAM Role
  +- AgentCore Runtime Role restricts AWS resource access scope
  +- Least privilege principle: only Bedrock InvokeModel, S3, DynamoDB, etc.

4.2 AgentCore Identity — OAuth Authentication

AgentCore Identity provides comprehensive OAuth 2.0 integration capabilities, divided into Inbound (verifying callers) and Outbound (accessing third-party APIs on behalf of users). Gateway builds on this to achieve transparent credential injection for MCP tools.

For complete code examples, configuration references, and Cognito setup scripts, see the dedicated document: agentcore-oauth-integration.md

Inbound JWT Authorizer (Quick Overview)

Configured via authorizerConfiguration when creating a Runtime. AgentCore automatically validates the JWT before the request reaches the container:

response = client.create_agent_runtime(
    agentRuntimeName='my-agent',
    authorizerConfiguration={
        "customJWTAuthorizer": {
            "discoveryUrl": "https://cognito-idp.us-east-1.amazonaws.com/POOL_ID/.well-known/openid-configuration",
            "allowedClients": ["your-client-id"],
        }
    },
    ...
)

The container can obtain user identity in two ways:

# Option A: From the request body; the container never touches the JWT
user_id = request.user_id   # Passed by the caller; already verified at the AgentCore platform layer

# Option B (when additional claims are needed): Decode the JWT (requires --request-header-allowlist "Authorization")
claims = jwt.decode(auth_header[7:], options={"verify_signature": False})
user_id, scopes = claims.get("sub"), claims.get("scope")

Once OAuth is enabled, callers must send HTTPS + Bearer Token directly; they cannot use boto3.invoke_agent_runtime() (SigV4).

Outbound Credential Provider (Quick Overview)

When an Agent calls a third-party OAuth API on behalf of a user, use the @requires_access_token decorator:

from bedrock_agentcore.identity.auth import requires_access_token

@requires_access_token(
    provider_name="google-provider",
    scopes=["https://www.googleapis.com/auth/drive.metadata.readonly"],
    auth_flow="USER_FEDERATION",
    on_auth_url=lambda url: print("Please authorize:", url),
)
async def list_drive_files(*, access_token: str):
    # access_token is auto-injected; Agent never touches refresh_token / client_secret
    return requests.get("https://www.googleapis.com/drive/v3/files",
                        headers={"Authorization": f"Bearer {access_token}"}).json()
Gateway OAuth (Quick Overview)

Gateway transparently injects OAuth credentials at the MCP tool layer; Agent code and LLM are completely unaware:

Agent calls MCP tool -> Gateway validates JWT -> Token Vault retrieves access_token -> Injects into downstream header -> Executes API

Pre-integrated services (1-Click): Salesforce, Slack, Jira, Asana, Zendesk.

4.3 Authentication Scheme Selection

Scenario Recommended Scheme
Agent only calls your own backend APIs Inbound JWT + API Key (Secrets Manager)
Agent calls third-party OAuth APIs on behalf of users Inbound JWT + Outbound Credential Provider
Agent calls third-party services via MCP tools Inbound JWT + Gateway OAuth (zero Agent code changes)

Evolution path: Start with API Key + Secrets Manager for simple service-to-service authentication. When you need to call Slack, Jira, or other third-party services on behalf of users, introduce Outbound Credential Provider + Gateway OAuth without managing OAuth tokens in Agent code.

4.4 LLM Security Isolation

  • Service API Keys are not injected into the system prompt and are invisible to the LLM
  • The tool wrapper layer injects auth headers at the Python code level
  • AgentCore Identity: In Gateway mode, credentials are entirely outside Agent code; in @requires_access_token mode, Agent code accesses the access_token but never the refresh_token / client_secret
  • Even if the LLM is subject to prompt injection, it cannot leak secrets

5. Observability

AgentCore Runtime provides out-of-the-box observability based on the OpenTelemetry (OTel) standard, sending Traces, Metrics, and Logs to CloudWatch in a unified manner.

5.1 Overall Architecture

Strands Agent (auto-generates OTel Spans + Metrics)
    |
    v
ADOT Sidecar (ADOT Collector auto-injected by AgentCore)
    |
    +---> AWS X-Ray (distributed tracing)
    +---> CloudWatch Logs (aws/spans log group)
    +---> CloudWatch Metrics (bedrock-agentcore namespace)

AgentCore Runtime automatically injects an ADOT (AWS Distro for OpenTelemetry) Sidecar as the telemetry collector; applications do not need to configure an OTel Collector themselves. Spans generated by the Strands SDK are sent by default via UDP to 127.0.0.1:2000 (the ADOT Sidecar address), which forwards them to X-Ray and CloudWatch.

5.2 Environment Variable Configuration

Observability-related environment variables configured in CDK:

// AgentCore Runtime environment variables
AGENT_OBSERVABILITY_ENABLED: 'true',
OTEL_RESOURCE_ATTRIBUTES: `service.name=${props.agentName}`,
OTEL_LOG_LEVEL: 'info',
Environment Variable Purpose Example Value
AGENT_OBSERVABILITY_ENABLED Enable AgentCore observability true
OTEL_RESOURCE_ATTRIBUTES OTel resource attributes (identifies service name) service.name=my-agent
OTEL_LOG_LEVEL OTel SDK log level info
OTEL_EXPORTER_OTLP_ENDPOINT OTLP export endpoint (auto-configured by ADOT Sidecar) http://127.0.0.1:4318
OTEL_TRACES_SAMPLER Sampling strategy (AgentCore defaults to X-Ray sampling) xray

5.3 Spans Auto-Generated by Strands SDK

Strands Agents SDK has built-in OpenTelemetry instrumentation that generates complete Agent execution traces without manual instrumentation. Span hierarchy:

invoke_agent {agent_name}              <- Root Span: entire Agent invocation
  +- execute_event_loop_cycle           <- Each iteration of the Agent event loop
       +- chat                          <- Model call (Bedrock API)
       +- execute_tool {tool_name}      <- Tool execution

Key attributes carried by each Span type:

The attribute names below come from the Strands SDK source code. Some attributes such as gen_ai.agent.tools and gen_ai.server.time_to_first_token are not explicitly listed in the official Strands documentation and may change across SDK versions.

Span Key Attributes
invoke_agent gen_ai.system=strands-agents, gen_ai.agent.name, gen_ai.request.model, gen_ai.agent.tools (tool list), gen_ai.usage.* (cumulative tokens)
chat gen_ai.request.model, gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, gen_ai.server.time_to_first_token, gen_ai.server.request.duration
execute_tool gen_ai.tool.name, gen_ai.tool.call.id, gen_ai.tool.status

Spans also carry Events that record the complete conversation content:

  • gen_ai.user.message — User input
  • gen_ai.assistant.message — Assistant response
  • gen_ai.tool.message — Tool input/output
  • gen_ai.choice — Model choice (includes finish_reason)

5.4 Metrics Auto-Generated by Strands SDK

Strands SDK also generates OTel Metrics, which can be exported to CloudWatch Metrics via ADOT:

Counters:

Metric Name Description
strands.event_loop.cycle_count Total Agent event loop iterations
strands.tool.call_count Total tool invocations
strands.tool.success_count Tool success count
strands.tool.error_count Tool failure count

Histograms:

Metric Name Unit Description
strands.event_loop.latency ms Agent invocation end-to-end latency
strands.event_loop.cycle_duration s Duration of a single loop iteration
strands.tool.duration s Tool execution duration
strands.event_loop.input.tokens count Input token count
strands.event_loop.output.tokens count Output token count
strands.model.time_to_first_token ms TTFT (Time to First Token latency)

5.5 Viewing Data in CloudWatch

AgentCore writes telemetry data to the following CloudWatch locations:

Data Type CloudWatch Location Purpose
Traces X-Ray -> CloudWatch ServiceMap Distributed tracing, visualize call chains
Spans CloudWatch Logs aws/spans ADOT-formatted Span documents
Runtime Logs CloudWatch Logs /aws/bedrock-agentcore/runtimes/{agent-id} Application logs (stdout/stderr)
Metrics CloudWatch Metrics bedrock-agentcore namespace Performance metrics

In the X-Ray console, you can see the complete Agent call chain: the duration, status, and token consumption of each chat (model call) and execute_tool (tool execution) round.

5.6 Evaluation: Quality Assessment Based on Spans

The bedrock-agentcore SDK provides StrandsToADOTConverter, which converts OTel Spans into ADOT format for the AgentCore Evaluation API:

from bedrock_agentcore.evaluation.span_to_adot_serializer import convert_strands_to_adot

# Retrieve raw OTel Spans from CloudWatch or in-memory
raw_spans = telemetry.in_memory_exporter.get_finished_spans()

# Convert to ADOT documents (includes Span documents + conversation logs + tool logs)
adot_docs = convert_strands_to_adot(raw_spans)

# Submit to AgentCore Evaluation API for quality assessment

The converted ADOT documents contain three types:

  • Span Documents: trace_id, span_id, duration, status
  • Conversation Log: Complete conversation record of user input -> assistant response
  • Tool Log: Execution record of tool input -> tool output

5.7 Application-Level Supplementary Logging

In addition to telemetry data auto-generated by the SDK, the application layer can log additional performance metrics via logging, written to CloudWatch Logs:

# TTFT (Time to First Token)
logger.info(f"[PERF] TTFT: {first_token_time:.3f}s")

# Tool execution time
logger.info(f"[PERF] Tool completed: {tool_name} took {execution_time_ms}ms")

# Total duration
logger.info(f"[PERF] Total: {total_time:.3f}s, events: {event_count}")

These logs complement OTel Spans: OTel provides structured tracing data (visualizable in X-Ray), while application logs provide human-readable diagnostic information.


6. BedrockAgentCoreApp vs FastAPI: Build Approach Comparison

In addition to building all endpoints yourself with FastAPI, the bedrock-agentcore SDK also provides the BedrockAgentCoreApp wrapper class, which auto-generates /invocations, /ping, and other endpoints. Both approaches can run on AgentCore Runtime, but they differ significantly in developer experience, built-in capabilities, and flexibility.

6.1 Core Differences

Dimension BedrockAgentCoreApp FastAPI Custom Build
Code Volume ~20 lines to deploy ~200+ lines (endpoints, middleware, SSE formatting)
/invocations Auto-generated via @app.entrypoint Manually define @app.post("/invocations")
/ping Auto-generated with built-in Healthy/HealthyBusy state machine Manually defined, returns fixed {"status": "healthy"}
/ws Auto-generated via @app.websocket Requires manual Starlette WebSocket integration
SSE Streaming Just yield; SDK auto-formats as SSE Manually construct data: {...}\n\n format
Async Tasks Built-in add_async_task / @app.ping / Worker Loop Requires custom task tracking + thread safety
Session Management context.session_id auto-injected Manually extract request.id from request body
Middleware Supports Starlette Middleware Native FastAPI middleware
Deployment CDK (fromAsset/fromS3) or Starter Toolkit (demo only) CDK (fromAsset/fromEcrRepository) + Dockerfile
Local Development python my_agent.py (app.run() auto-starts uvicorn) uvicorn main:app --reload
Framework Lock-in Tied to bedrock-agentcore SDK Standard FastAPI, portable to any platform

6.2 Minimal Code Comparison

BedrockAgentCoreApp (~15 lines):

# [BedrockAgentCoreApp]
from bedrock_agentcore import BedrockAgentCoreApp
from strands import Agent

app = BedrockAgentCoreApp()
agent = Agent(system_prompt="You are an intelligent assistant")

@app.entrypoint
async def main(payload):
    async for event in agent.stream_async(payload.get("prompt", "")):
        if "data" in event:
            yield event["data"]

if __name__ == "__main__":
    app.run()  # Auto-listens on 8080, auto /ping, /invocations, /ws

FastAPI Custom Build (simplified, ~60 lines):

# [FastAPI Custom Build]
from fastapi import FastAPI
from fastapi.responses import StreamingResponse
from strands import Agent

app = FastAPI()
agent_system_prompt = "You are an intelligent assistant"

@app.post("/invocations")
async def stream_agent(request: ChatRequest):
    async def event_generator():
        session_manager = _create_session_manager(request.id, request.user_id)
        agent = Agent(system_prompt=agent_system_prompt, session_manager=session_manager, ...)

        yield f"data: {json.dumps({'type': 'start', 'session_id': request.id})}\n\n"
        async for event in agent.stream_async(user_message):
            if "delta" in event and "text" in event["delta"]:
                yield f"data: {json.dumps({'type': 'text-delta', 'delta': event['delta']['text']})}\n\n"
        yield f"data: {json.dumps({'type': 'finish'})}\n\n"
        yield "data: [DONE]\n\n"

    return StreamingResponse(event_generator(), media_type="text/event-stream")

@app.get("/ping")
async def ping():
    return {"status": "healthy"}

@app.get("/health")
async def health():
    return {"status": "healthy", "agent_name": "my-agent"}

6.3 Architecture Differences

BedrockAgentCoreApp Architecture:

app.run()
  +-- Starlette ASGI App (created internally by SDK)
        +-- GET  /ping        <- Auto-generated with built-in state machine
        +-- POST /invocations <- @app.entrypoint route
        +-- WS   /ws          <- @app.websocket route (optional)
        +-- Worker Loop thread <- Isolates handler execution, prevents blocking /ping

FastAPI Custom Build Architecture:

uvicorn main:app
  +-- FastAPI ASGI App (created by developer)
        +-- GET  /ping        <- Manually defined
        +-- GET  /health      <- Manually defined
        +-- POST /invocations <- Manually defined + SSE formatting
        +-- CORS / Timing middleware <- Manually registered

6.4 When to Choose Which

Choose BedrockAgentCoreApp:

  • New project, rapid prototyping
  • Need async background tasks (HealthyBusy state management)
  • Need WebSocket bidirectional communication
  • No need for complex custom middleware or endpoints

Choose FastAPI Custom Build:

  • Need full control over request/response format (custom SSE event types)
  • Need additional endpoints (/health, /, custom APIs)
  • Need complex middleware chains (authentication, rate limiting, logging, timing)
  • Need portability (may migrate to ECS/EKS/Lambda in the future)
  • Team is familiar with the FastAPI ecosystem

6.5 Migration Path

If you later need to migrate from FastAPI to BedrockAgentCoreApp (e.g., for async task management), the core changes are:

# Before [FastAPI Custom Build]
app = FastAPI()

@app.post("/invocations")
async def stream_agent(request: ChatRequest):
    ...

@app.get("/ping")
async def ping():
    return {"status": "healthy"}

# After [BedrockAgentCoreApp]
from bedrock_agentcore import BedrockAgentCoreApp

app = BedrockAgentCoreApp()

@app.entrypoint
async def main(payload, context):
    # payload = original request body
    # context.session_id = original request.id
    ...
    yield chunk  # SDK auto-formats as SSE

@app.ping
def custom_ping():
    if has_background_tasks():
        return PingStatus.HEALTHY_BUSY
    return PingStatus.HEALTHY

Key changes:

  • @app.post("/invocations") -> @app.entrypoint
  • ChatRequest parsing -> payload dict + context.session_id
  • Manual SSE formatting -> yield raw data
  • Manual /ping -> @app.ping or automatic state machine
  • StreamingResponse -> SDK handles automatically

Appendix: Typical Project File Structure

Use CDK for all production deployments. AgentCore Runtime is not an isolated service — it depends on peripheral resources such as Cognito, ECR, IAM Roles, Secrets Manager, and DynamoDB, all of which need to be orchestrated through CDK. The Starter Toolkit CLI (agentcore configure + agentcore deploy) is suitable for quick demos and prototyping, but production projects cannot deploy just a Runtime alone.

BedrockAgentCoreApp (CDK Deployment):

my-agent/
+-- main.py                    # BedrockAgentCoreApp entry point (@app.entrypoint + app.run())
+-- tools/
|   +-- custom_tools.py        # Custom @tool functions
+-- requirements.txt           # Dependency list (used with CDK L2 fromS3)
+-- Dockerfile                 # ARM64 container build (used with CDK L2 fromAsset)
+-- pyproject.toml             # uv project config (optional, replaces requirements.txt)
+-- uv.lock                    # Lock file (optional)

CDK deployment for BedrockAgentCoreApp is exactly the same as for FastAPI — CDK does not care which framework the application uses internally; it only sees a container or zip package. fromS3 does not require a Dockerfile (startup command ['python', 'main.py']), while fromAsset requires a Dockerfile (CMD python3 main.py).

FastAPI Custom Build (CDK Deployment):

my-agent/
+-- main.py                    # FastAPI main entry (/invocations, /ping)
+-- config.py                  # pydantic-settings configuration
+-- schema.py                  # Request/response model definitions
+-- config_manager.py          # S3 + local config loading
+-- tools/
|   +-- http_request.py        # HTTP tool wrapper (auto auth)
|   +-- custom_tools.py        # Custom @tool functions
+-- configs/
|   +-- agent_config.yaml      # Agent configuration (system prompt, tools)
+-- Dockerfile                 # ARM64 container build
+-- pyproject.toml             # uv project config
+-- uv.lock                    # Lock file

CDK Side (shared by both build approaches):

infrastructure/
+-- lib/agentcore/
|   +-- runtime-construct.ts   # CfnRuntime (L1) or Runtime (L2) definition
|   +-- gateway-construct.ts   # MCP Gateway (optional)
+-- lib/auth/
|   +-- cognito-construct.ts   # Cognito UserPool + App Client
+-- lib/backend/
|   +-- api-construct.ts       # Backend API (API Gateway + Lambda / ECS)
+-- lib/main-stack.ts          # Overall deployment orchestration

Source: SKILL.md on GitHub

1 alert16d4 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    The skill provides a comprehensive AWS Bedrock AgentCore orchestration guide with documentation, templates, and reference materials. No security issues, prompt injections, malicious dependencies, or obfuscation layers were detected.

  • Socket16d

    1 alert: gptAnomaly

  • Snyk16d

    Risk: MEDIUM · 1 issue

  • Runlayer6mo

    10/13 files flagged

Signed by skilld at e4ef2e2. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 months ago.

Steadyupdated 6 months ago
What it can do
Network Runs commands
MCP servers
aws-mcpawsdocsacdocs
Modelsonnet
aliases
[
  "bedrock-agentcore"
]
context
fork
model
sonnet
All 12 allowed tools
mcp__aws-mcp__*mcp__awsdocs__*mcp__acdocs__search_agentcore_docsmcp__acdocs__fetch_agentcore_docBash(aws bedrock-agentcore *)Bash(aws bedrock-agentcore-control *)Bash(aws bedrock-agentcore-runtime *)Bash(aws bedrock *)Bash(aws s3 cp *)Bash(aws s3 ls *)Bash(aws secretsmanager *)Bash(aws sts get-caller-identity)
Other metadata
skills
[
  "aws-mcp-setup"
]
hooks
{
  "PreToolUse": [
    {
      "matcher": "Bash(aws bedrock-agentcore-control create-*)",
      "command": "aws sts get-caller-identity --query Account --output text",
      "once": true
    }
  ]
}

README badge

README badge for zxkane/aws-skills/aws-agentic-ai