All skills
datadog-labs avatar

/agent-observability-experiment-bootstrap

@47d0cf5 official

Bootstrap a reproducible LLM Observability experiment through the Python ddtrace SDK or the Node dd-trace SDK. Use for experiment, dataset, evaluator, benchmark, regression, or LLM-as-a-judge scaffolding. The legacy Python invocation remains supported.

Use this Skill: https://skilld.dev/gh/datadog-labs/agent-skills/agent-observability-experiment-bootstrap

This session only. Nothing lands on disk.

referencespythonevaluator-stylesclass.md

≈462 tokens on demand. Your agent reads this file only when SKILL.md points to it.

--evaluator-style class (advanced — for evaluators that need state or async I/O)

BaseEvaluator subclasses with evaluate(self, context: EvaluatorContext) -> EvaluatorResult. Always return EvaluatorResult — never a bare value. State-bearing evaluators usually have richer reasoning to surface anyway.

Code to emit

from ddtrace.llmobs import BaseEvaluator, EvaluatorContext, EvaluatorResult

class FaithfulnessJudge(BaseEvaluator):
    def __init__(self):
        super().__init__(name="faithfulness")
        # TODO(user): initialize any client or state here

    def evaluate(self, context: EvaluatorContext) -> EvaluatorResult:
        # context exposes: input_data, output_data, expected_output, metadata
        # TODO(user): replace placeholder logic with your faithfulness check
        passed = context.output_data is not None
        return EvaluatorResult(
            value=1.0 if passed else 0.0,
            reasoning="placeholder — replace with your faithfulness rubric",
            assessment="pass" if passed else "fail",
            metadata={"evaluator_version": "v1"},
        )

Rules

  • Call super().__init__(name=...) in __init__. The name is the column header in the Datadog Experiments UI.
  • evaluate() runs in the experiment's worker pool. Do NOT mutate self from evaluate() (thread safety) — state set in __init__ should be read-only thereafter.
  • For async work (e.g., calling an LLM judge over the network), prefer wrapping with asyncio.run(...) inside evaluate() rather than making evaluate itself async. Keeps the experiment runner sync.

When NOT to use this style

If the evaluator is a one-line check (exact_match, length_under_500), use function style — the class boilerplate adds noise. See references/python/evaluator-styles/function.md.

Source: SKILL.md on GitHub

No alerts1mo3 checks · Risk SAFE
  • Gen Agent Trust Hub1mo

    This skill provides scaffolding for LLM observability experiments using Datadog's official SDKs. It adheres to security best practices by explicitly forbidding the hardcoding of credentials, implementing PII scrubbing for datasets, and utilizing official vendor libraries and documentation.

  • Socket1mo

    No alerts

  • Snyk1mo

    Risk: LOW · No issues

Signed by skilld at 47d0cf5. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated last month

README badge

README badge for datadog-labs/agent-skills/agent-observability-experiment-bootstrap