Evals
Assert on what a real model does with the app's tools → @skybridge/test
DevTools proves a tool works when called. An eval proves the model calls it, with the right arguments, from a natural prompt. Use one when a tool's name/description/schema changes, when two tools could be confused, or when the user asks how the app behaves in a real conversation. Evals are live model calls: they cost money and need an API key, so they are not unit tests. Keep them few and behavior-focused.
Setup
- Dev dependencies:
@skybridge/test@beta(published on thebetadist-tag only),vitest@^4(vitest 5 is not supported yet),ai, and an AI SDK provider (@ai-sdk/anthropic,@ai-sdk/openai, ...). vite.config.ts:skybridge({ evals: {} }). This registers theexpect.chatmatchers, picks upevals/**/*.eval.ts, raises the per-scenario timeout to two minutes, and loads.env.package.json:"evals": "vitest run evals".- The provider key in
.env(ANTHROPIC_API_KEYfor@ai-sdk/anthropic). If the app hasoauth, its provider env is needed too:setupandoauthresolve on the first request.
The default demo template from create skybridge already has all of this, plus evals/start.eval.ts to copy from; the blank template has none of it.
Scenario
// evals/search-flights.eval.ts
import { anthropic } from "@ai-sdk/anthropic";
import { start } from "@skybridge/test";
import { expect, it } from "vitest";
import { app } from "../src/server.js";
it("searches flights from a natural prompt", async () => {
const chat = await start({ app, model: anthropic("claude-sonnet-4-5") });
await chat.send("I need to fly to Lisbon next weekend");
expect.chat(chat).toHaveCalledToolOnce("search-flights", { destination: "Lisbon" });
expect.chat(chat).toNeverHaveCalledTool("book-flight");
});start serves the app in process: no port, no HTTP server. This only works because src/server.ts exports the app and src/index.ts runs it; never put run() in server.ts. Each send is one user turn, during which the model may call several tools. The session closes with the test.
Matchers
All typed against the app's registry (name autocompletes, args is checked against the tool's inputSchema); all support .not.
| Matcher | Passes when |
|---|---|
toHaveCalledToolOnce(name, args?) |
exactly one successful call, optionally matching args (partial, objectContaining) |
toHaveCalledToolWith(name, args) |
some successful call matched args |
toNeverHaveCalledTool(name) |
no call was attempted |
toHaveFailedToolCall(name) |
a call was refused (auth) or threw |
toHaveSaid(text | RegExp) |
an assistant turn contains it (string match is case- and whitespace-insensitive) |
toHaveCalledToolsInOrder(...names) |
the named tools succeeded in that relative order (subsequence, gaps allowed) |
toHaveCalledNoTools() |
no tool was attempted at all |
await toPassJudgment(criteria, options?) |
a judge model grades the conversation against criteria written in plain English |
On failure the message lists every call the model made, with arguments. Matchers see the whole conversation, not just the last send. chat.toolCalls and chat.assistantTurns are available for custom assertions.
Judgments
toPassJudgment is the only async matcher: await expect.chat(chat).toPassJudgment("stays inside the app's scope and explains the tool result"). The judge reads every turn and tool call with its result, runs at temperature 0 on the chat's own model, and its reasoning lands in the failure message. options takes model (another judge) or judge, a callback receiving { criteria, transcript } and returning { pass, reasoning? }, which lets an evaluation model or a scoring service grade instead of a language model. A judge that throws, the provider or your own callback, raises judge unavailable: <error> instead of reporting a fail. Use it only for criteria no other matcher can express. The default judge is a live model call, so it costs money and is not reproducible; a custom judge costs and varies only as much as whatever it calls, and a local heuristic is free and deterministic.
Stubs
stubs answers a tool from the scenario instead of the app, for results that move over time (relative dates, a live catalogue):
const chat = await start({
app,
model: anthropic("claude-sonnet-4-5"),
stubs: { "search-flights": ({ to }) => (to === "LIS" ? lisbonFixture : undefined) },
});Arguments are typed against the registry. Returning undefined falls through to the real server. A stubbed call still appears in chat.toolCalls, but never reaches the handler, so its schema validation and scope checks do not run for that call.
Authenticated apps
Claim an identity per session; only token verification is skipped, per-tool auth and scope checks run for real:
const chat = await start({
app,
model: anthropic("claude-sonnet-4-5"),
authInfo: { token: "eval", clientId: "evals", scopes: ["orders:read"], extra: { subject: "user-1" } },
});Omit authInfo to test the anonymous path: a gated tool then shows up as toHaveFailedToolCall.
Defaults
evals: { temperature, systemPrompt, maxSteps, timeout } in the Vite plugin sets what every scenario starts from (temperature 0, maxSteps 8, timeout 120s). temperature, systemPrompt and maxSteps can be overridden per start.
Pitfalls
- Assert on tool calls and arguments, not on exact wording; use
toHaveSaidwith a loose pattern when the answer matters. - A failing eval usually means the tool
descriptionor schema.describe()text is unclear to the model, not that the handler is wrong. Fix the prompt surface first. - Do not run evals in a loop while iterating on UI; run them once a tool's contract changes.