All skills
alpic-ai avatar

/chatgpt-app-builder

@3340b34
by Alpicalpic-ai/skybridge2.1k stars
140

Guide developers through creating and updating ChatGPT apps. Covers the full lifecycle: brainstorming ideas against UX guidelines, bootstrapping projects, implementing tools/views, debugging, running dev servers, deploying and connecting apps to ChatGPT. Use when a user wants to create or update a ChatGPT app / MCP server for ChatGPT, or use the Skybridge framework.

Use this Skill: https://skilld.dev/gh/alpic-ai/skybridge/chatgpt-app-builder

This session only. Nothing lands on disk.

referencesevals.md

≈1.5k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Evals

Assert on what a real model does with the app's tools → @skybridge/test

DevTools proves a tool works when called. An eval proves the model calls it, with the right arguments, from a natural prompt. Use one when a tool's name/description/schema changes, when two tools could be confused, or when the user asks how the app behaves in a real conversation. Evals are live model calls: they cost money and need an API key, so they are not unit tests. Keep them few and behavior-focused.

Setup

  1. Dev dependencies: @skybridge/test@beta (published on the beta dist-tag only), vitest@^4 (vitest 5 is not supported yet), ai, and an AI SDK provider (@ai-sdk/anthropic, @ai-sdk/openai, ...).
  2. vite.config.ts: skybridge({ evals: {} }). This registers the expect.chat matchers, picks up evals/**/*.eval.ts, raises the per-scenario timeout to two minutes, and loads .env.
  3. package.json: "evals": "vitest run evals".
  4. The provider key in .env (ANTHROPIC_API_KEY for @ai-sdk/anthropic). If the app has oauth, its provider env is needed too: setup and oauth resolve on the first request.

The default demo template from create skybridge already has all of this, plus evals/start.eval.ts to copy from; the blank template has none of it.

Scenario

// evals/search-flights.eval.ts
import { anthropic } from "@ai-sdk/anthropic";
import { start } from "@skybridge/test";
import { expect, it } from "vitest";
import { app } from "../src/server.js";

it("searches flights from a natural prompt", async () => {
  const chat = await start({ app, model: anthropic("claude-sonnet-4-5") });
  await chat.send("I need to fly to Lisbon next weekend");

  expect.chat(chat).toHaveCalledToolOnce("search-flights", { destination: "Lisbon" });
  expect.chat(chat).toNeverHaveCalledTool("book-flight");
});

start serves the app in process: no port, no HTTP server. This only works because src/server.ts exports the app and src/index.ts runs it; never put run() in server.ts. Each send is one user turn, during which the model may call several tools. The session closes with the test.

Matchers

All typed against the app's registry (name autocompletes, args is checked against the tool's inputSchema); all support .not.

Matcher Passes when
toHaveCalledToolOnce(name, args?) exactly one successful call, optionally matching args (partial, objectContaining)
toHaveCalledToolWith(name, args) some successful call matched args
toNeverHaveCalledTool(name) no call was attempted
toHaveFailedToolCall(name) a call was refused (auth) or threw
toHaveSaid(text | RegExp) an assistant turn contains it (string match is case- and whitespace-insensitive)
toHaveCalledToolsInOrder(...names) the named tools succeeded in that relative order (subsequence, gaps allowed)
toHaveCalledNoTools() no tool was attempted at all
await toPassJudgment(criteria, options?) a judge model grades the conversation against criteria written in plain English

On failure the message lists every call the model made, with arguments. Matchers see the whole conversation, not just the last send. chat.toolCalls and chat.assistantTurns are available for custom assertions.

Judgments

toPassJudgment is the only async matcher: await expect.chat(chat).toPassJudgment("stays inside the app's scope and explains the tool result"). The judge reads every turn and tool call with its result, runs at temperature 0 on the chat's own model, and its reasoning lands in the failure message. options takes model (another judge) or judge, a callback receiving { criteria, transcript } and returning { pass, reasoning? }, which lets an evaluation model or a scoring service grade instead of a language model. A judge that throws, the provider or your own callback, raises judge unavailable: <error> instead of reporting a fail. Use it only for criteria no other matcher can express. The default judge is a live model call, so it costs money and is not reproducible; a custom judge costs and varies only as much as whatever it calls, and a local heuristic is free and deterministic.

Stubs

stubs answers a tool from the scenario instead of the app, for results that move over time (relative dates, a live catalogue):

const chat = await start({
  app,
  model: anthropic("claude-sonnet-4-5"),
  stubs: { "search-flights": ({ to }) => (to === "LIS" ? lisbonFixture : undefined) },
});

Arguments are typed against the registry. Returning undefined falls through to the real server. A stubbed call still appears in chat.toolCalls, but never reaches the handler, so its schema validation and scope checks do not run for that call.

Authenticated apps

Claim an identity per session; only token verification is skipped, per-tool auth and scope checks run for real:

const chat = await start({
  app,
  model: anthropic("claude-sonnet-4-5"),
  authInfo: { token: "eval", clientId: "evals", scopes: ["orders:read"], extra: { subject: "user-1" } },
});

Omit authInfo to test the anonymous path: a gated tool then shows up as toHaveFailedToolCall.

Defaults

evals: { temperature, systemPrompt, maxSteps, timeout } in the Vite plugin sets what every scenario starts from (temperature 0, maxSteps 8, timeout 120s). temperature, systemPrompt and maxSteps can be overridden per start.

Pitfalls

  • Assert on tool calls and arguments, not on exact wording; use toHaveSaid with a loose pattern when the answer matters.
  • A failing eval usually means the tool description or schema .describe() text is unclear to the model, not that the handler is wrong. Fix the prompt surface first.
  • Do not run evals in a loop while iterating on UI; run them once a tool's contract changes.

Source: SKILL.md on GitHub

1 alert6d5 checks · Risk SAFE
  • Gen Agent Trust Hub6d

    The skill is a comprehensive development guide for the Skybridge framework, which helps build ChatGPT applications. It is generally safe and follows security best practices, such as documenting CSP and OAuth configurations. It identifies a minor risk of indirect prompt injection as it processes user-provided specification files to guide the development process.

  • Socket6d

    No alerts

  • Snyk6d

    Risk: LOW · No issues

  • Runlayer7mo

    8/23 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 3340b34. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 days ago.

Activeupdated last week

README badge

README badge for alpic-ai/skybridge

Guides developers through the complete lifecycle of building ChatGPT apps—from ideation and project setup through implementation, debugging, and deployment to the ChatGPT Directory. Covers the Skybridge MCP server framework, tool integration, custom UI views, state management, OAuth, and the dual-user model where both the human and ChatGPT LLM interact with the same view.

Generated from the current SKILL.md.

What is a ChatGPT app and how does it work?
A ChatGPT app is a conversational experience that extends ChatGPT through tools and custom UI views, built as an MCP server. The view serves as a shared surface where both the human user and the ChatGPT LLM interact and collaborate.
Do I need a SPEC.md file before building?
Yes. You should read discover.md first to create a SPEC.md that documents your app's requirements and design decisions before implementing anything.
What does this skill cover?
The skill guides the full lifecycle: brainstorming ideas against UX guidelines, bootstrapping projects, implementing tools and views, debugging, running dev servers, deploying via Alpic, and publishing to the ChatGPT Directory.
Does this skill provide API documentation?
Yes. Full API docs are available at https://docs.skybridge.tech/api-reference.md, and the skill includes implementation references for specific tasks like OAuth, state management, and CSP configuration.

Generated from the current SKILL.md. These answers refresh after source changes.