All skills
launchdarkly avatar

/built-in-metrics

@add614f official

Instrument an existing codebase with LaunchDarkly config tracking. Walks the four-tier ladder (managed runner → provider package → custom extractor + trackMetricsOf → raw manual) and picks the lowest-ceremony option that still captures duration, tokens, and success/error.

Use this Skill: https://skilld.dev/gh/launchdarkly/agent-skills/built-in-metrics

This session only. Nothing lands on disk.

referencesanthropic-tracking.md

≈1.6k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Anthropic Metrics Tracking

There is no LaunchDarkly provider package for Anthropic direct API today. The canonical path is the generic trackMetricsOf wrapper (Tier 3) with a small custom extractor that reads response.usage.input_tokens and response.usage.output_tokens. The instinct to "default to generic" is correct here — there is no Tier-2 shortcut to take.

Three viable paths, in order of preference:

  1. Route Anthropic through LangChain. If the app already uses LangChain (or can adopt it cheaply), install the LangChain provider package and use it as Tier 2. LangChain's ChatAnthropic wrapper exposes the standardized usage_metadata that getAIMetricsFromResponse reads.
  2. Route Anthropic through Bedrock Converse. If the app can switch to Bedrock Converse (Claude is available on Bedrock), you inherit Bedrock's Converse response shape and a custom-extractor pattern that's slightly cleaner. See bedrock-tracking.md.
  3. Custom extractor on the direct SDK (this file's primary pattern).

Tier 1 is not available

ManagedModel does not currently ship an Anthropic provider. If you need Tier 1 for a chat app, use option 1 or 2 above — the LangChain provider package lets ManagedModel wrap a ChatAnthropic under the hood, which restores the zero-tracker-call experience.

Tier 3 — Custom extractor + trackMetricsOf (primary)

Python — direct Anthropic SDK:

import anthropic
from ldai.providers.types import LDAIMetrics, TokenUsage

client = anthropic.Anthropic()

def anthropic_extractor(response) -> LDAIMetrics:
    return LDAIMetrics(
        success=True,
        tokens=TokenUsage(
            total=response.usage.input_tokens + response.usage.output_tokens,
            input=response.usage.input_tokens,
            output=response.usage.output_tokens,
        ),
    )

def call_with_tracking(ai_config, user_prompt: str) -> str | None:
    if not ai_config.enabled:
        return None

    system_content = ai_config.messages[0].content if ai_config.messages else ""

    def call_anthropic():
        return client.messages.create(
            model=ai_config.model.name,
            max_tokens=1024,
            system=system_content,
            messages=[{"role": "user", "content": user_prompt}],
        )

    tracker = ai_config.create_tracker()
    # Exceptions are tracked automatically here: track_metrics_of catches
    # exceptions, records tracker.track_error(), and re-raises. Do NOT add
    # except: tracker.track_error() on top — it's a noop that trips the
    # at-most-once guard. Wrap in your own try/except only if you need
    # local handling (logging, fallback, alert); the error is already tracked.
    response = tracker.track_metrics_of(anthropic_extractor, call_anthropic)
    return response.content[0].text

Node — direct Anthropic SDK:

import Anthropic from '@anthropic-ai/sdk';
import type { LDAIMetrics } from '@launchdarkly/server-sdk-ai';

const client = new Anthropic();

const anthropicExtractor = (response: Anthropic.Message): LDAIMetrics => ({
  success: true,
  tokens: {
    total: response.usage.input_tokens + response.usage.output_tokens,
    input: response.usage.input_tokens,
    output: response.usage.output_tokens,
  },
});

async function callWithTracking(
  aiConfig: LDAICompletionConfig,
  userPrompt: string,
): Promise<string | null> {
  if (!aiConfig.enabled) return null;

  const systemContent = aiConfig.messages?.[0]?.content ?? '';

  const tracker = aiConfig.createTracker();
  // Exceptions are tracked automatically: trackMetricsOf catches exceptions,
  // records tracker.trackError(), and re-throws. Do NOT add
  // catch (err) { tracker.trackError(); throw err } on top — it's a noop
  // that trips the at-most-once guard. Wrap in your own try/catch only if
  // you need local handling (logging, fallback); the error is already tracked.
  const response = await tracker.trackMetricsOf(
    anthropicExtractor,
    () => client.messages.create({
      model: aiConfig.model!.name,
      max_tokens: 1024,
      system: systemContent,
      messages: [{ role: 'user', content: userPrompt }],
    }),
  );
  return response.content[0].type === 'text' ? response.content[0].text : null;
}

Notes on the extractor shape:

  • Anthropic returns input_tokens / output_tokens on response.usage. Compute total yourself; Anthropic does not provide it.
  • LDAIMetrics is a typed surface — Python has it at ldai.providers.types, Node exports it from @launchdarkly/server-sdk-ai. Keep the extractor pure: no side effects, no network calls.
  • success: true in the extractor is not a lie — trackMetricsOf only calls the extractor on the success path. On the error path, trackMetricsOf records trackError() internally and re-throws; no caller-side catch block is required.

Tier 2 option — route via LangChain

If the app can adopt LangChain, the LangChain provider package handles Anthropic (via @langchain/anthropic) through the same trackMetricsOf(getAIMetricsFromResponse, ...) pattern used for any other LangChain model. This is often the cleanest answer if the app already uses or is open to LangChain, because the extractor is built in and shared with every other LangChain-wrapped model.

from ldai_langchain import create_langchain_model, get_ai_metrics_from_response

ai_config = ai_client.completion_config("my-config-key", context, default_config)
llm = create_langchain_model(ai_config)  # ChatAnthropic under the hood

tracker = ai_config.create_tracker()
response = tracker.track_metrics_of(
    get_ai_metrics_from_response,
    lambda: llm.invoke(messages),
)

Tier 4 — Manual (streaming only)

Streaming Anthropic needs manual TTFT tracking; the pattern is identical to OpenAI streaming. See streaming-tracking.md.

What NOT to do

  • Do not hand-wire track_duration_of + track_tokens + track_success as three separate calls unless you're on the streaming path. That's Tier 4, and trackMetricsOf gives you the same three metrics in one call with half the drift surface.
  • Do not look for a track_anthropic_metrics helper — it doesn't exist, never has, and won't be added. Anthropic direct support lives in the extractor you write above.
  • Do not invent a provider package like @launchdarkly/server-sdk-ai-anthropic. It doesn't exist as of this writing. Check js-core ai-providers before recommending one.

Source: SKILL.md on GitHub

No alerts2d3 checks · Risk SAFE
  • Gen Agent Trust Hub2d

    The skill provides patterns and best practices for instrumenting AI applications with LaunchDarkly's monitoring capabilities. It correctly suggests using environment variables for secrets and interacts with vendor-owned domains. However, it establishes an attack surface for indirect prompt injection by demonstrating how to wrap LLM provider calls that process untrusted user input without providing examples of input sanitization or boundary enforcement.

  • Socket2d

    No alerts

  • Snyk2d

    Risk: LOW · No issues

Signed by skilld at add614f. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 days ago.

Activeupdated 4 months ago
metadata
{
  "author": "launchdarkly",
  "version": "1.0.0-experimental"
}
Other metadata
compatibility
Requires the LaunchDarkly server-side AI SDK (`launchdarkly-server-sdk-ai>=0.20.0` for Python or `@launchdarkly/server-sdk-ai>=0.20.0` for Node) and an existing config.
  • AI/ML
  • launchdarkly
  • metrics
  • instrumentation
  • tracking
  • monitoring
  • openai
  • langchain
  • anthropic

README badge

README badge for launchdarkly/agent-skills/built-in-metrics

Instruments an existing codebase with LaunchDarkly config tracking by walking a four-tier ladder from managed runner down to raw manual calls, selecting the lowest-ceremony option that still captures duration, tokens, and success/error. Targets Python and Node codebases using OpenAI, LangChain, Vercel AI SDK, Anthropic, Gemini, Bedrock, or custom HTTP providers.

Generated from the current SKILL.md.

What LaunchDarkly SDK version does this skill require?
The skill requires launchdarkly-server-sdk-ai version 0.20.0 or later (Python: `launchdarkly-server-sdk-ai>=0.20.0`, Node: `@launchdarkly/server-sdk-ai>=0.20.0`) and an existing LaunchDarkly config.
Does this skill work with streaming responses?
Yes, but streaming with time-to-first-token (TTFT) tracking requires Tier 4 (raw manual tracking). Node offers `trackStreamMetricsOf` for the streaming wrapper, but TTFT must be tracked explicitly via `trackTimeToFirstToken`.
Which AI providers are supported?
The skill supports OpenAI, LangChain, Vercel AI SDK, AWS Bedrock, Anthropic, Gemini, Google GenAI, Strands Agents, and custom HTTP providers. Provider package availability and tracking tier options differ by framework and language—see the included reference matrix.
Can I use this with chat loops or only one-shot completions?
The skill supports both. Chat loops use Tier 1 (managed runner, highest-priority tier with zero tracker calls), while one-shot completions, agent steps, and other non-chat patterns use Tiers 2–4 depending on available provider packages.
What metrics does this capture?
The skill captures duration, input/output token counts, success/error status, and time-to-first-token for streaming—the four core metrics the LaunchDarkly Monitoring tab displays. The exact tracking method depends on which tier you implement (Tier 1 captures all automatically, Tier 2–3 require minimal code, Tier 4 is fully manual).

Generated from the current SKILL.md. These answers refresh after source changes.