All skills
langchain-ai avatar

/langsmith-online-eval-engineering

@6732f69 official

Iteratively inspect traces, interview the user, and create LangSmith online evaluators one at a time. Use specifically for creating online evaluators for use within LangSmith -- use "eval-engineering" for Harbor-style online evaluations.

Use this Skill: https://skilld.dev/gh/langchain-ai/langchain-skills/langsmith-online-eval-engineering

This session only. Nothing lands on disk.

referencestrace-inspection.md

≈998 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Trace Inspection

How to discover trace field names for use in evaluator variable_mapping (LLM evaluators) and run.inputs/run.outputs (code evaluators).

Always inspect traces before building an evaluator. Never guess field names.

Procedure

  1. Ask the user for their LangSmith project name.
  2. Fetch 3--5 recent traces:
from langsmith import Client

client = Client()
runs = list(client.list_runs(project_name="<project>", limit=5))
  1. For each run, print:

    • run.name -- the run name (often the chain or agent class)
    • run.run_type -- chain, llm, tool, retriever, etc.
    • list(run.inputs.keys()) -- available input field names
    • list(run.outputs.keys()) -- available output field names
    • Truncated samples of run.inputs and run.outputs (cap at ~2000 chars)
  2. Identify which fields carry the data the evaluator needs.

Common field name patterns

Chatbot / conversational agent

inputs:  {"input": "user message"}      or  {"messages": [...]}
outputs: {"output": "assistant reply"}  or  {"messages": [...]}

Typical variable_mapping: {"input": "input", "output": "output"}

RAG / retrieval chain

inputs:  {"input": "user question", "chat_history": [...]}
outputs: {"output": "answer", "context": [...]}

Context may appear as "context", "documents", or "source_documents". Check the actual keys.

Tool-calling agent

inputs:  {"input": "user request"}
outputs: {"output": "final answer"}

Tool calls appear in child runs, not in the top-level run's inputs/outputs. The top-level run still has input and output for the user-facing request and response.

Custom chains

Field names depend on the chain's implementation. There is no universal schema -- this is why inspection is required.

Mapping fields to evaluators

LLM evaluators: variable_mapping

variable_mapping connects prompt template variables to top-level trace fields. Only top-level fields are supported.

# If the prompt template uses {question} and {answer}:
VARIABLE_MAPPING = {
    "question": "input",    # template var -> trace field
    "answer": "output",
}

The template variables must match {placeholders} in the prompt messages. The trace field values must match keys in run.inputs or run.outputs.

Code evaluators: run["inputs"] and run["outputs"]

Code evaluators access trace data through the run dict:

def perform_eval(run, example=None):
    inputs = run.get("inputs") or {}
    outputs = run.get("outputs") or {}
    user_input = inputs.get("input", "")
    agent_output = outputs.get("output", "")
    # ... evaluate ...

run is a plain dict at runtime -- use run.get("inputs"), not run.inputs. Always use .get() with defaults and guard against None outputs.

Handling variations

Nested structures

Some traces nest data inside wrapper keys:

inputs: {"input": {"question": "...", "context": "..."}}

For LLM evaluators, variable_mapping maps to top-level keys only. If data is nested, either:

  • Map to the top-level key and handle the nested structure in the prompt
  • Use a code evaluator instead, which can traverse nested dicts

Inconsistent schemas

Different runs in the same project may have different field names if multiple chain types feed into one project. Inspect several traces to identify the common pattern. If schemas vary, a code evaluator with defensive .get() calls is more robust than an LLM evaluator with fixed variable_mapping.

Missing outputs

Runs that errored may have run["outputs"] set to None. Code evaluators must handle this:

def perform_eval(run, example=None):
    outputs = run.get("outputs")
    if outputs is None:
        return {"key": "my_check", "score": 0, "comment": "No output (run may have errored)"}
    # ... normal evaluation ...

LLM evaluators will receive an empty string for mapped fields when the trace field is missing.

Source: SKILL.md on GitHub

1 warning2mo3 checks · Risk SAFE
  • Gen Agent Trust Hub2mo

    This skill facilitates the iterative development of online evaluators for LangSmith. It features a structured workflow for trace inspection, evaluator design, and deployment using official LangChain tools and APIs. While the skill involves executing generated code and processing trace data, these actions are core to its functionality and are managed within the context of the user's LangSmith environment.

  • Socket2mo

    No alerts

  • Snyk2mo

    Risk: MEDIUM · 1 issue

Signed by skilld at 6732f69. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 days ago.

Activeupdated 2 months ago

README badge

README badge for langchain-ai/langchain-skills/langsmith-online-eval-engineering