All skills
datadog-labs avatar

/agent-observability-experiment-bootstrap

@47d0cf5 official

Bootstrap a reproducible LLM Observability experiment through the Python ddtrace SDK or the Node dd-trace SDK. Use for experiment, dataset, evaluator, benchmark, regression, or LLM-as-a-judge scaffolding. The legacy Python invocation remains supported.

Use this Skill: https://skilld.dev/gh/datadog-labs/agent-skills/agent-observability-experiment-bootstrap

This session only. Nothing lands on disk.

referencespythonevaluator-stylesfunction.md

≈466 tokens on demand. Your agent reads this file only when SKILL.md points to it.

--evaluator-style function (default — what the notebooks use)

Plain Python functions with the signature (input_data, output_data, expected_output). Always emit at least three: a trivial boolean (returns bool), a richer rule-based one (returns EvaluatorResult), and an LLM-as-Judge surrogate (a RemoteEvaluator reference or a placeholder).

Code to emit

from ddtrace.llmobs import EvaluatorResult

# Trivial check — bare bool is fine here, the result has no extra signal.
def exact_match(input_data, output_data, expected_output) -> bool:
    return output_data == expected_output

# Richer check — use EvaluatorResult so reasoning/assessment surface in the UI.
def response_well_formed(input_data, output_data, expected_output) -> EvaluatorResult:
    if not isinstance(output_data, str):
        return EvaluatorResult(
            value=False,
            reasoning=f"output_data was {type(output_data).__name__}, expected str",
            assessment="fail",
        )
    if len(output_data) > 500:
        return EvaluatorResult(
            value=False,
            reasoning=f"output exceeded 500 chars (was {len(output_data)})",
            assessment="fail",
            metadata={"length": len(output_data)},
        )
    return EvaluatorResult(value=True, assessment="pass")

When to extend

  • If the user passed --dataset with a structured expected_output, add a JSON-shape check (also returning EvaluatorResult).
  • For LLM-as-Judge surrogates, prefer RemoteEvaluator references (server-side, scalable) over inline LLMJudge calls.

When NOT to use this style

If the evaluator needs persistent state (a model client, a cached lookup, an async I/O resource), use class style instead — BaseEvaluator.__init__ is where you set up state safely. See references/python/evaluator-styles/class.md.

Source: SKILL.md on GitHub

No alerts1mo3 checks · Risk SAFE
  • Gen Agent Trust Hub1mo

    This skill provides scaffolding for LLM observability experiments using Datadog's official SDKs. It adheres to security best practices by explicitly forbidding the hardcoding of credentials, implementing PII scrubbing for datasets, and utilizing official vendor libraries and documentation.

  • Socket1mo

    No alerts

  • Snyk1mo

    Risk: LOW · No issues

Signed by skilld at 47d0cf5. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated last month

README badge

README badge for datadog-labs/agent-skills/agent-observability-experiment-bootstrap