All skills
launchdarkly avatar

/built-in-metrics

@add614f official

Instrument an existing codebase with LaunchDarkly config tracking. Walks the four-tier ladder (managed runner → provider package → custom extractor + trackMetricsOf → raw manual) and picks the lowest-ceremony option that still captures duration, tokens, and success/error.

Use this Skill: https://skilld.dev/gh/launchdarkly/agent-skills/built-in-metrics

This session only. Nothing lands on disk.

referencesbedrock-tracking.md

≈1.4k tokens on demand. Your agent reads this file only when SKILL.md points to it.

AWS Bedrock Metrics Tracking

There is no LaunchDarkly provider package for Bedrock today (neither Python nor Node). Two practical paths:

  1. Route Bedrock through LangChain (ChatBedrockConverse / langchain-aws). If you're open to LangChain, this is the closest thing to Tier 2 — you use the LangChain provider package's getAIMetricsFromResponse and inherit the whole trackMetricsOf pattern for free.
  2. Custom extractor on boto3 (this file's primary pattern). Bedrock Converse returns a stable response shape with usage.inputTokens / usage.outputTokens / usage.totalTokens, so the extractor is three lines.

Tier 1 is not available

ManagedModel does not ship a Bedrock provider today (Python or Node). If you want Tier 1 for a Bedrock chat app, route via LangChain — ManagedModel can wrap a ChatBedrockConverse through the LangChain provider package.

Tier 3 — Custom extractor + trackMetricsOf (primary)

Converse API (recommended)

Python:

import boto3
from ldai.providers.types import LDAIMetrics, TokenUsage

bedrock = boto3.client("bedrock-runtime")

def bedrock_converse_extractor(response) -> LDAIMetrics:
    usage = response.get("usage", {})
    return LDAIMetrics(
        success=True,
        tokens=TokenUsage(
            total=usage.get("totalTokens", 0),
            input=usage.get("inputTokens", 0),
            output=usage.get("outputTokens", 0),
        ),
    )

def call_with_tracking(ai_config, user_prompt: str) -> str | None:
    if not ai_config.enabled:
        return None

    system_content = ai_config.messages[0].content if ai_config.messages else ""

    def call_bedrock():
        kwargs = {
            "modelId": ai_config.model.name,
            "messages": [{"role": "user", "content": [{"text": user_prompt}]}],
        }
        if system_content:
            kwargs["system"] = [{"text": system_content}]
        return bedrock.converse(**kwargs)

    tracker = ai_config.create_tracker()
    # Exceptions are tracked automatically — track_metrics_of catches
    # exceptions, records tracker.track_error(), and re-raises.
    response = tracker.track_metrics_of(bedrock_converse_extractor, call_bedrock)
    return response["output"]["message"]["content"][0]["text"]

Node:

import { BedrockRuntimeClient, ConverseCommand, type ConverseCommandOutput } from '@aws-sdk/client-bedrock-runtime';
import type { LDAIMetrics } from '@launchdarkly/server-sdk-ai';

const bedrock = new BedrockRuntimeClient({});

const bedrockConverseExtractor = (response: ConverseCommandOutput): LDAIMetrics => ({
  success: true,
  tokens: {
    total: response.usage?.totalTokens ?? 0,
    input: response.usage?.inputTokens ?? 0,
    output: response.usage?.outputTokens ?? 0,
  },
});

async function callWithTracking(
  aiConfig: LDAICompletionConfig,
  userPrompt: string,
): Promise<string | null> {
  if (!aiConfig.enabled) return null;

  const systemContent = aiConfig.messages?.[0]?.content;

  const tracker = aiConfig.createTracker();
  // Exceptions are tracked automatically — trackMetricsOf catches
  // exceptions, records tracker.trackError(), and re-throws.
  const response = await tracker.trackMetricsOf(
    bedrockConverseExtractor,
    () => bedrock.send(new ConverseCommand({
      modelId: aiConfig.model!.name,
      messages: [{ role: 'user', content: [{ text: userPrompt }] }],
      ...(systemContent ? { system: [{ text: systemContent }] } : {}),
    })),
  );
  return response.output?.message?.content?.[0]?.text ?? null;
}

Legacy InvokeModel API

InvokeModel returns per-model shapes (Anthropic on Bedrock returns Anthropic's shape, Llama on Bedrock returns Meta's shape, etc.), so the extractor has to branch. Prefer Converse unless you're locked into InvokeModel by an older model that Converse doesn't support. If you must use InvokeModel, switch the extractor based on the model family:

def invoke_model_extractor(response) -> LDAIMetrics:
    body = json.loads(response["body"].read())
    # Claude on InvokeModel
    if "usage" in body:
        return LDAIMetrics(
            success=True,
            tokens=TokenUsage(
                total=body["usage"]["input_tokens"] + body["usage"]["output_tokens"],
                input=body["usage"]["input_tokens"],
                output=body["usage"]["output_tokens"],
            ),
        )
    # Llama / Titan — use the fields on the specific body shape
    # ...
    return LDAIMetrics(success=True, tokens=TokenUsage(total=0, input=0, output=0))

This is a good reason to migrate to Converse if you can.

Tier 2 option — route via LangChain

If the app uses LangChain, the LangChain provider package's ChatBedrockConverse support gives you the Tier-2 experience:

from ldai_langchain import create_langchain_model, get_ai_metrics_from_response

ai_config = ai_client.completion_config("my-config-key", context, default_config)
llm = create_langchain_model(ai_config)  # ChatBedrockConverse when provider=bedrock

tracker = ai_config.create_tracker()
response = tracker.track_metrics_of(
    get_ai_metrics_from_response,
    lambda: llm.invoke(messages),
)

LangChain normalizes the Converse response shape into AIMessage.usage_metadata, which get_ai_metrics_from_response reads — so you don't need a Bedrock-specific extractor.

Tier 4 — Manual (streaming only)

Bedrock Converse streaming (ConverseStream) needs manual TTFT tracking. The pattern is identical to OpenAI streaming. See streaming-tracking.md.

Source: SKILL.md on GitHub

No alerts2d3 checks · Risk SAFE
  • Gen Agent Trust Hub2d

    The skill provides patterns and best practices for instrumenting AI applications with LaunchDarkly's monitoring capabilities. It correctly suggests using environment variables for secrets and interacts with vendor-owned domains. However, it establishes an attack surface for indirect prompt injection by demonstrating how to wrap LLM provider calls that process untrusted user input without providing examples of input sanitization or boundary enforcement.

  • Socket2d

    No alerts

  • Snyk2d

    Risk: LOW · No issues

Signed by skilld at add614f. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 days ago.

Activeupdated 4 months ago
metadata
{
  "author": "launchdarkly",
  "version": "1.0.0-experimental"
}
Other metadata
compatibility
Requires the LaunchDarkly server-side AI SDK (`launchdarkly-server-sdk-ai>=0.20.0` for Python or `@launchdarkly/server-sdk-ai>=0.20.0` for Node) and an existing config.
  • AI/ML
  • launchdarkly
  • metrics
  • instrumentation
  • tracking
  • monitoring
  • openai
  • langchain
  • anthropic

README badge

README badge for launchdarkly/agent-skills/built-in-metrics

Instruments an existing codebase with LaunchDarkly config tracking by walking a four-tier ladder from managed runner down to raw manual calls, selecting the lowest-ceremony option that still captures duration, tokens, and success/error. Targets Python and Node codebases using OpenAI, LangChain, Vercel AI SDK, Anthropic, Gemini, Bedrock, or custom HTTP providers.

Generated from the current SKILL.md.

What LaunchDarkly SDK version does this skill require?
The skill requires launchdarkly-server-sdk-ai version 0.20.0 or later (Python: `launchdarkly-server-sdk-ai>=0.20.0`, Node: `@launchdarkly/server-sdk-ai>=0.20.0`) and an existing LaunchDarkly config.
Does this skill work with streaming responses?
Yes, but streaming with time-to-first-token (TTFT) tracking requires Tier 4 (raw manual tracking). Node offers `trackStreamMetricsOf` for the streaming wrapper, but TTFT must be tracked explicitly via `trackTimeToFirstToken`.
Which AI providers are supported?
The skill supports OpenAI, LangChain, Vercel AI SDK, AWS Bedrock, Anthropic, Gemini, Google GenAI, Strands Agents, and custom HTTP providers. Provider package availability and tracking tier options differ by framework and language—see the included reference matrix.
Can I use this with chat loops or only one-shot completions?
The skill supports both. Chat loops use Tier 1 (managed runner, highest-priority tier with zero tracker calls), while one-shot completions, agent steps, and other non-chat patterns use Tiers 2–4 depending on available provider packages.
What metrics does this capture?
The skill captures duration, input/output token counts, success/error status, and time-to-first-token for streaming—the four core metrics the LaunchDarkly Monitoring tab displays. The exact tracking method depends on which tier you implement (Tier 1 captures all automatically, Tier 2–3 require minimal code, Tier 4 is fully manual).

Generated from the current SKILL.md. These answers refresh after source changes.