All skills
launchdarkly avatar

/built-in-metrics

@add614f official

Instrument an existing codebase with LaunchDarkly config tracking. Walks the four-tier ladder (managed runner → provider package → custom extractor + trackMetricsOf → raw manual) and picks the lowest-ceremony option that still captures duration, tokens, and success/error.

Use this Skill: https://skilld.dev/gh/launchdarkly/agent-skills/built-in-metrics

This session only. Nothing lands on disk.

referencesstrands-tracking.md

≈1.3k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Strands Agents Metrics Tracking

There is no LaunchDarkly provider package for Strands. Strands is a provider-agnostic agent SDK — the same Agent class runs against Anthropic, OpenAI, and Bedrock by swapping the model argument — so the tracking pattern plugs in at the agent layer, not the provider layer. Tier 3 (custom extractor + trackMetricsOf) is the canonical path.

The Strands AgentResult object exposes a metrics.accumulated_usage dict (Python) / metrics.accumulatedUsage object (Node) that already aggregates token counts across every provider call the agent made in a single invoke_async turn — including any tool-calling round trips. That means one extractor call covers the whole turn, unlike the per-response shape from Anthropic or OpenAI direct.

The key names inside accumulated_usage are camelCase even in Python: inputTokens, outputTokens, totalTokens.

Tier 1 is not available

ManagedModel does not currently ship a Strands runner. Strands owns its own agent loop and short-term memory (SlidingWindowConversationManager), so wrapping it in a LaunchDarkly managed runner would fight against the framework. Stay on Tier 3.

Tier 3 — Explicit track_duration_of + manual track_tokens (primary)

This is the shape in the LaunchDarkly Strands integration guide. Use it when the call site is already async and you want token extraction split out from duration tracking.

from ldai.tracker import TokenUsage


def track_strands_metrics(tracker, result):
    """Record token usage from a Strands AgentResult on the LD tracker."""
    usage = getattr(result.metrics, "accumulated_usage", {}) or {}
    input_tokens = usage.get("inputTokens", 0)
    output_tokens = usage.get("outputTokens", 0)
    total = usage.get("totalTokens", 0) or (input_tokens + output_tokens)
    if total > 0:
        tracker.track_tokens(
            TokenUsage(input=input_tokens, output=output_tokens, total=total)
        )


async def run_turn(agent, tracker, user_input):
    try:
        result = await tracker.track_duration_of(lambda: agent.invoke_async(user_input))
        tracker.track_success()
        track_strands_metrics(tracker, result)
        return result.message["content"][0]["text"]
    except Exception:
        tracker.track_error()
        raise

What this tracks:

  • Duration — from the track_duration_of wrapper around invoke_async.
  • Tokens — from accumulated_usage, including any tool-calling round trips inside the turn.
  • Success / error — explicit, in the try/except.

Tier 3 — Single-call track_metrics_of_async variant

If you prefer the single-call form that matches the rest of the provider-tracking references, fold the extractor into an LDAIMetrics return and use track_metrics_of_async:

from ldai.providers.types import LDAIMetrics, TokenUsage


def strands_extractor(result) -> LDAIMetrics:
    usage = getattr(result.metrics, "accumulated_usage", {}) or {}
    input_tokens = usage.get("inputTokens", 0)
    output_tokens = usage.get("outputTokens", 0)
    total = usage.get("totalTokens", 0) or (input_tokens + output_tokens)
    return LDAIMetrics(
        success=True,
        tokens=TokenUsage(input=input_tokens, output=output_tokens, total=total),
    )


async def run_turn(agent, tracker, user_input):
    # Exceptions are tracked automatically — track_metrics_of_async catches
    # exceptions, records tracker.track_error(), and re-raises.
    result = await tracker.track_metrics_of_async(
        strands_extractor,
        lambda: agent.invoke_async(user_input),
    )
    return result.message["content"][0]["text"]

Pick the style that matches the rest of the codebase — the two variants record the same metrics.

Provider dispatch stays in your code

Strands model classes are provider-specific (AnthropicModel, OpenAIModel, BedrockModel). To serve more than one provider from a single config key, dispatch on agent_config.provider.name before constructing the Agent. See agent-mode-frameworks.md § Strands Agent for the create_strands_model dispatcher, including the rule that parameters.tools must be dropped before being passed into the Strands model class (tools flow through the Agent constructor, not through model params).

Always flush before exit

Strands examples are commonly short-lived scripts (python run_agent.py ...). Trailing analytics events can be lost if the client closes before flushing. Always call ldclient.get().flush() (and ldclient.get().close() on exit) after the last turn.

Node / TypeScript caveat

The Strands TypeScript SDK ships BedrockModel and OpenAIModel only — no AnthropicModel. The same Tier-3 pattern applies (custom extractor over result.metrics.accumulatedUsage, then tracker.trackMetricsOf or explicit trackDurationOf + trackTokens), but multi-provider variations that include Anthropic require the Python SDK today.

Source: SKILL.md on GitHub

No alerts2d3 checks · Risk SAFE
  • Gen Agent Trust Hub2d

    The skill provides patterns and best practices for instrumenting AI applications with LaunchDarkly's monitoring capabilities. It correctly suggests using environment variables for secrets and interacts with vendor-owned domains. However, it establishes an attack surface for indirect prompt injection by demonstrating how to wrap LLM provider calls that process untrusted user input without providing examples of input sanitization or boundary enforcement.

  • Socket2d

    No alerts

  • Snyk2d

    Risk: LOW · No issues

Signed by skilld at add614f. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 days ago.

Activeupdated 4 months ago
metadata
{
  "author": "launchdarkly",
  "version": "1.0.0-experimental"
}
Other metadata
compatibility
Requires the LaunchDarkly server-side AI SDK (`launchdarkly-server-sdk-ai>=0.20.0` for Python or `@launchdarkly/server-sdk-ai>=0.20.0` for Node) and an existing config.
  • AI/ML
  • launchdarkly
  • metrics
  • instrumentation
  • tracking
  • monitoring
  • openai
  • langchain
  • anthropic

README badge

README badge for launchdarkly/agent-skills/built-in-metrics

Instruments an existing codebase with LaunchDarkly config tracking by walking a four-tier ladder from managed runner down to raw manual calls, selecting the lowest-ceremony option that still captures duration, tokens, and success/error. Targets Python and Node codebases using OpenAI, LangChain, Vercel AI SDK, Anthropic, Gemini, Bedrock, or custom HTTP providers.

Generated from the current SKILL.md.

What LaunchDarkly SDK version does this skill require?
The skill requires launchdarkly-server-sdk-ai version 0.20.0 or later (Python: `launchdarkly-server-sdk-ai>=0.20.0`, Node: `@launchdarkly/server-sdk-ai>=0.20.0`) and an existing LaunchDarkly config.
Does this skill work with streaming responses?
Yes, but streaming with time-to-first-token (TTFT) tracking requires Tier 4 (raw manual tracking). Node offers `trackStreamMetricsOf` for the streaming wrapper, but TTFT must be tracked explicitly via `trackTimeToFirstToken`.
Which AI providers are supported?
The skill supports OpenAI, LangChain, Vercel AI SDK, AWS Bedrock, Anthropic, Gemini, Google GenAI, Strands Agents, and custom HTTP providers. Provider package availability and tracking tier options differ by framework and language—see the included reference matrix.
Can I use this with chat loops or only one-shot completions?
The skill supports both. Chat loops use Tier 1 (managed runner, highest-priority tier with zero tracker calls), while one-shot completions, agent steps, and other non-chat patterns use Tiers 2–4 depending on available provider packages.
What metrics does this capture?
The skill captures duration, input/output token counts, success/error status, and time-to-first-token for streaming—the four core metrics the LaunchDarkly Monitoring tab displays. The exact tracking method depends on which tier you implement (Tier 1 captures all automatically, Tier 2–3 require minimal code, Tier 4 is fully manual).

Generated from the current SKILL.md. These answers refresh after source changes.