All skills
datadog-labs avatar

/agent-observability-experiment-bootstrap

@47d0cf5 official

Bootstrap a reproducible LLM Observability experiment through the Python ddtrace SDK or the Node dd-trace SDK. Use for experiment, dataset, evaluator, benchmark, regression, or LLM-as-a-judge scaffolding. The legacy Python invocation remains supported.

Use this Skill: https://skilld.dev/gh/datadog-labs/agent-skills/agent-observability-experiment-bootstrap

This session only. Nothing lands on disk.

referencespythonprovidersbedrock.md

≈590 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Provider: AWS Bedrock

Triggered by introspection (Workflow step 2.5) when the call-site function imports boto3 and calls:

  • boto3.client("bedrock-runtime").invoke_model(...)
  • boto3.client("bedrock-runtime").converse(...) (newer API)
  • boto3.Session(...).client("bedrock-runtime")...

{{PROVIDER_ASSERTS}} substitution

assert os.getenv("AWS_ACCESS_KEY_ID"), "AWS_ACCESS_KEY_ID is required for the wired task_fn (Bedrock)."
assert os.getenv("AWS_SECRET_ACCESS_KEY"), "AWS_SECRET_ACCESS_KEY is required for the wired task_fn (Bedrock)."

Optional env vars (do NOT assert; document in # TODO comments)

  • AWS_SESSION_TOKEN — required when using short-lived credentials (SSO, IAM Identity Center, etc.). If set, must be present at runtime.

  • AWS_REGION (or AWS_DEFAULT_REGION) — Bedrock is region-scoped. Defaults to us-east-1 if unset; emit a comment:

    # AWS Bedrock is region-scoped. Defaults to us-east-1; set AWS_REGION if your
    # Bedrock-enabled region differs (us-west-2, eu-central-1, etc.).
  • AWS_PROFILE — alternative to access-key/secret pair when using ~/.aws/credentials. If the user's function uses boto3.Session(profile_name=...), key-pair asserts may not apply — emit a comment instead.

Adapter notes

  • client.invoke_model(modelId=..., body=...) (older API) — body is a JSON string with provider-specific shape (Anthropic Claude, Amazon Titan, AI21, Cohere, Meta Llama — each has different body schema). Extract response via json.loads(response["body"].read()).
  • client.converse(modelId=..., messages=[...]) (newer Converse API) — standardized request/response across providers. Extract via response["output"]["message"]["content"][0]["text"].
  • If the user's function uses invoke_model, leave their body construction intact in task_fn — Anthropic-on-Bedrock vs Llama-on-Bedrock have different request shapes.

Common gotchas

  • Model IDs differ from upstream provider IDs (e.g., anthropic.claude-3-5-sonnet-20240620-v1:0 rather than claude-3-5-sonnet-20240620). Don't rewrite the model ID.
  • Bedrock charges per call; rate limits apply. Consider --jobs 1 for the first experiment run to gauge cost.
  • Cross-region inference profiles use a different model ID prefix (us.anthropic..., eu.anthropic...). Trust the user's setup.

Source: SKILL.md on GitHub

No alerts1mo3 checks · Risk SAFE
  • Gen Agent Trust Hub1mo

    This skill provides scaffolding for LLM observability experiments using Datadog's official SDKs. It adheres to security best practices by explicitly forbidding the hardcoding of credentials, implementing PII scrubbing for datasets, and utilizing official vendor libraries and documentation.

  • Socket1mo

    No alerts

  • Snyk1mo

    Risk: LOW · No issues

Signed by skilld at 47d0cf5. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated last month

README badge

README badge for datadog-labs/agent-skills/agent-observability-experiment-bootstrap