All skills
aktsmm avatar

/vscode-extension-guide

@7ae9078
by yamapanaktsmm/agent-skills26 stars
4

Guide for creating VS Code extensions and plugins from scratch through Marketplace publication. Use when developing a VS Code extension/plugin, adding commands or keybindings, building TreeView or Webview UI, publishing to Marketplace, or troubleshooting activation and packaging issues.

Use this Skill: https://skilld.dev/gh/aktsmm/agent-skills/vscode-extension-guide

This session only. Nothing lands on disk.

referencestesting.md

≈4.5k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Testing VS Code Extensions

Set up and run tests using @vscode/test-electron.

Setup

npm install -D @vscode/test-electron mocha @types/mocha glob

Project Structure

my-extension/
├── src/
│   └── extension.ts
├── test/
│   ├── runTest.ts           # Test runner entry
│   └── suite/
│       ├── index.ts         # Mocha configuration
│       └── extension.test.ts # Test file
├── tsconfig.json
└── tsconfig.test.json

tsconfig.test.json

{
  "extends": "./tsconfig.json",
  "compilerOptions": {
    "rootDir": ".",
    "outDir": "out"
  },
  "include": ["src/**/*", "test/**/*"]
}

test/runTest.ts

Treat engines.vscode as the supported API/runtime floor, not the developer's installed version. Pin @types/vscode to that floor; derive test-host and isolated-install versions from the validated manifest range rather than separate constants. Only strip a caret after validating a simple ^major.minor.patch range; use a range parser for other forms. Keep lockfile and compatibility docs aligned. A host rejecting activation on its engine check does not prove an API failure: test a proposed lower floor with its types and real Extension Host before lowering it. Newer versions inside the declared range need no blanket untested-version warning; report actual failures with a manual, data-free Issue link instead of uploading diagnostics automatically.

import * as path from "path";
import { readFileSync } from "node:fs";
import { runTests } from "@vscode/test-electron";

async function main() {
  try {
    const extensionDevelopmentPath = path.resolve(__dirname, "../../");
    const extensionTestsPath = path.resolve(__dirname, "./suite/index");
    const manifest = JSON.parse(
      readFileSync(path.join(extensionDevelopmentPath, "package.json"), "utf8"),
    );
    const range = manifest.engines?.vscode;
    if (typeof range !== "string" || !/^\^\d+\.\d+\.\d+$/.test(range)) {
      throw new Error("This runner requires a simple caret engine range.");
    }

    await runTests({
      extensionDevelopmentPath,
      extensionTestsPath,
      version: range.slice(1),
      // Optional: open specific workspace
      // launchArgs: ['--disable-extensions', path.resolve(__dirname, '../../test-workspace')],
    });
  } catch (err) {
    console.error("Failed to run tests");
    process.exit(1);
  }
}

main();

test/suite/index.ts

import * as path from "path";
import Mocha from "mocha";
import { glob } from "glob";

export async function run(): Promise<void> {
  const mocha = new Mocha({
    ui: "tdd",
    color: true,
    timeout: 10000,
    failZero: true,
  });

  const testsRoot = path.resolve(__dirname, ".");
  const files = await glob("**/**.test.js", { cwd: testsRoot });

  files.forEach((f) => mocha.addFile(path.resolve(testsRoot, f)));

  return new Promise((resolve, reject) => {
    mocha.run((failures) => {
      if (failures > 0) {
        reject(new Error(`${failures} tests failed.`));
      } else {
        resolve();
      }
    });
  });
}

test/suite/extension.test.ts

import * as assert from "assert";
import * as vscode from "vscode";

suite("Extension Test Suite", () => {
  vscode.window.showInformationMessage("Start all tests.");

  test("Extension should be present", () => {
    const ext = vscode.extensions.getExtension("publisher.extension-name");
    assert.ok(ext, "Extension not found");
  });

  test("Extension should activate", async () => {
    const ext = vscode.extensions.getExtension("publisher.extension-name");
    await ext?.activate();
    assert.ok(ext?.isActive, "Extension not activated");
  });

  test("Command should be registered", async () => {
    const commands = await vscode.commands.getCommands();
    assert.ok(commands.includes("myExt.hello"), "Command not registered");
  });

  test("Command should execute without error", async () => {
    await assert.doesNotReject(vscode.commands.executeCommand("myExt.hello"));
  });
});

package.json Scripts

{
  "scripts": {
    "compile": "tsc -p ./",
    "compile-tests": "tsc -p tsconfig.test.json",
    "pretest": "npm run compile && npm run compile-tests",
    "test": "node ./out/test/runTest.js"
  }
}

Running Tests

# Run all tests
npm test

# Tests will:
# 1. Download VS Code (if needed)
# 2. Launch VS Code with extension loaded
# 3. Execute test suite
# 4. Exit with result code

Risk-Based Regression Checks

Typecheck explicitly when compile only bundles (for example, esbuild); then run the smallest behavior check and the full suite before release.

Extension Host assertions, command resolution, CI success and package identity do not prove that a downstream service produced a user-visible result. For changes to Chat, LM Tools, authentication, model selection or other external integrations, run the affected workflow before release with the same candidate VSIX in disposable --user-data-dir and --extensions-dir roots. Use disabled/disposable data, verify readback and persisted state, observe the real downstream result (for example, an exact Chat response marker), and remove the profile afterward. Never install over a normal profile that can contain enabled schedules or user data.

If the isolated account does not expose the target model or capability, report that path as unverified and keep a static guard; do not select a hidden model or treat a nearby model's successful smoke as proof for it.

Change area Extra checks
Commands / settings / views in package.json Verify manifest consistency, command IDs, setting keys, menu when clauses, and README setting tables
Manifest/runtime localization Treat these as separate systems: compare every manifest package.nls*.json key set; verify runtime bundle.l10n.*.json covers each vscode.l10n.t message with identical placeholder sets; confirm both sets ship in the VSIX
Runtime logging / diagnostics Verify logs go through an Output Channel logger instead of direct console.* calls in extension runtime paths
Resource scanners / providers Test with extension-host APIs available and with missing/empty roots; avoid relying on local filesystem guesses; if you scan installed extensions, cover both known resources/* roots and manifest-declared chatAgents / chatPromptFiles paths
Selectors / quick actions / saved options Hide internal, test, deprecated, stale, or unsupported candidates; preserve newly introduced normal candidates; confirm hidden saved values do not reappear from settings, cache, or fallback paths
Chat / LM Tools / model configuration Query the installed tool catalog, create only disabled disposable state, execute one explicit ID, observe the actual response, and confirm selected model/configuration plus disabled-state persistence
Installer / updater / index merge logic Run focused regression scripts plus a broader smoke test because these paths often cross manifest, filesystem, and network boundaries
Generated marker sections Test duplicate marker handling and confirm the final file contains exactly one generated section pair
Worker or secondary module entry points Guard that the runtime path (for example path.join(__dirname, "x.js")) resolves to a real source file, a compiled output, and an entry in the packaged payload allowlist; packaging can pass while analysis fails only at runtime

Filesystem and realpath behavior can classify the same missing resource differently across developer machines and hosted runners. Assert the user-facing contract first (for example, fallback source/status/payload), and only pin an internal error reason when the test controls the exact failure branch. If multiple reasons are specification-equivalent, use an explicit allowed set instead of one environment-dependent value.

For shared manifest, installer, updater or scanner changes, add a focused regression instead of relying only on manual verification.

Reliability Gotchas

  • Execute actual Webview submit/source-change handlers, not only source-token assertions. Textareas normalize CRLF to LF; normalize comparisons without inferring a provenance change. Keep file-to-inline conversion explicit, disable hidden required fields, and focus validation errors on visible controls.
  • Test rejected create/update requests with multiple changed fields and invalid paths, including NUL in cached local/global references. Assert that memory, persisted payload and metadata remain unchanged; read-time path rejection does not prove that invalid input cannot be saved.
  • A synchronous storage mock cannot prove queued persistence ordering. Hold the first write with a deferred promise, enqueue the next, and assert it cannot start or read stale state before the first settles. Cover rejected-write recovery, preserved entries and direct versus best-effort error propagation; release and await all pending work before fixture cleanup, without sleeps.
  • A UI automation timeout, a missing busy indicator or a disabled Send button is not proof that a mutating LM Tool did not run. Inspect the Chat transcript and persisted readback before retrying; uncertain delivery can otherwise create duplicates.
  • For bounded automatic dispatch, test oldest-due selection, reserved daily capacity and slot release on both execution and persistence failures. Keep unclaimed work pending. Name what callback completion proves: a Chat command resolving may confirm dispatch, not completion of the model response; a per-window limiter is not a cross-window guarantee.
  • Do not assume npm test -- --grep <pattern> reaches Mocha. Parse supported options before editor download/launch, reject unknown arguments and empty/invalid regex, and forward the unchanged pattern through extensionTestsEnv. With no CLI filter, explicitly clear the inherited filter so CI runs the full suite. Verify selected-test counts, zero-match exit failure, and unfiltered execution with an impossible ambient filter. If PowerShell drops options through npm.ps1, use npm.cmd; for shell-sensitive regex, compile first and invoke the Node runner directly.
  • Make each direct test entry point compile or clean first. An explicit list such as node --test out/test/a.test.js can silently pass against stale out/ while a newly added source test never runs. Give every public test script a matching npm lifecycle hook (for example, pretest:unit) and add a guard that every source *.test.ts has a compiled path in the test script, or use deterministic discovery.
  • For privacy-sensitive opt-outs around asynchronous file reads, clearing a cache is not enough. Increment a generation on opt-out/disposal, check it after every await and before cache/UI writes, close any view showing the disabled data, and use an injected delayed filesystem in tests to prove an in-flight read cannot repopulate state after opt-out.
  • For scanners, test pure merging and the actual registered create/change/delete callbacks: prove cache invalidation, in-flight cancellation and displayed-state updates, including another open item. Frequent writes should update one entry; deletion may use a debounced full scan when sibling files share an ID. Before optimizing selected-item refresh, measure filesystem call counts and guard new/deleted files, sibling formats, cold caches and ordinary full refresh. Synthetic call reductions are not wall-clock speedups; state when other items' cached timestamps refresh.
  • Test optional builtins through the real lazy loader, not only an injected replacement. Load the compiled module in an isolated VM with controlled require; assert zero loads on import, one load across repeated successful or failed reads, safe failure results, and isolation between module instances. Exercise the default read path too. Avoid production reset APIs or changing global module caches for tests; account for cross-realm prototypes in object comparisons.
  • Async scans need a generation token and a disposed guard so an older completion cannot overwrite newer state or update UI after deactivation. Route fire-and-forget promises through one rejection handler and assert that watcher/timer entry points use it.
  • If behavior depends on ExtensionContext.storageUri, run the Extension Host suite both with a folder argument and without one. Empty windows can have different storage roots and otherwise remain an unexecuted branch.
  • Size Extension Host fixtures by what the assertion discriminates, not by the production limit. Building a multi-megabyte payload inside the host can trip the Extension host is unresponsive watchdog even when the suite still passes; shrink the fixture until it is fast and still exercises the branch.
  • On Windows, @vscode/test-electron can fail before extension activation while VS Code setup holds the global vscode-updating mutex. A downloaded archive is not an Inno Setup installation: use downloadAndUnzipVSCode, verify the executable resolves inside the dedicated .vscode-test cache, set only that copy's product.json#win32VersionedUpdate to false, re-read it to confirm the value, and pass its vscodeExecutablePath to runTests. Archive layouts may add a build-hash directory before resources/app/product.json; resolve candidate realpaths under the cache root, require exactly one candidate, require a JSON object with a boolean win32VersionedUpdate, and never patch the machine installation.
for (const folderArgs of [[extensionDevelopmentPath], []]) {
  await runTests({
    extensionDevelopmentPath,
    extensionTestsPath,
    launchArgs: [...folderArgs, "--disable-extensions"],
  });
}

Treat the archive patch as test infrastructure with static guards: assert cache containment, the single-candidate product lookup, and both workspace/empty-window launches. The main-process log can still print a harmless instance-mutex warning; completion is decided by extension-host assertions and exit code.

Common Test Patterns

Testing with Documents

test("Should modify document", async () => {
  const doc = await vscode.workspace.openTextDocument({
    content: "hello",
    language: "plaintext",
  });
  const editor = await vscode.window.showTextDocument(doc);

  await editor.edit((editBuilder) => {
    editBuilder.insert(new vscode.Position(0, 5), " world");
  });

  assert.strictEqual(doc.getText(), "hello world");
});

Testing Settings

test("Should read configuration", () => {
  const config = vscode.workspace.getConfiguration("myExt");
  const value = config.get<string>("greeting");
  assert.strictEqual(value, "Hello");
});

Waiting for Events

test("Should handle file save", async () => {
  const doc = await vscode.workspace.openTextDocument({ content: "test" });

  const savePromise = new Promise<void>((resolve) => {
    const disposable = vscode.workspace.onDidSaveTextDocument((saved) => {
      if (saved === doc) {
        disposable.dispose();
        resolve();
      }
    });
  });

  await doc.save();
  await savePromise;
});

Terminal Readiness and Owned Cleanup

  • Diagnose terminal.shellIntegration, its cwd, and terminal.state.shell separately. Active integration does not guarantee a detected shell type; log only readiness flags and an allowlisted shell label, not commands, output or environment values.
  • If shell detection is absent, do not guess from a profile label, switch shells silently or stretch readiness timeouts to pass a gate. A platform-specific fallback needs bounded read-only observation of the owned terminal's actual process, PID validation and fail-closed handling of errors or ambiguous children; it must not kill processes or bypass execution-time authorization checks.
  • Use isolated local fixtures to assert real start/output/end events, literal argv/stdin, cancellation and unrelated-process survival. Unit mocks or a blocked outcome do not substitute for successful real-shell gates. Retry only after new evidence or a corrective change.
  • Resolve owned Webview tabs from the current tabGroups immediately before cleanup rather than retaining stale Tab objects across asynchronous UI changes. Limit cleanup to the view opened by the test; never close unrelated user tabs.

CI Integration

.github/workflows/test.yml:

name: Test
on: [push, pull_request]
jobs:
  test:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: 20
      - run: npm ci
      - run: xvfb-run -a npm test

Note: xvfb-run is required on Linux for headless VS Code testing.

Source: SKILL.md on GitHub

1 warning8d5 checks · Risk SAFE
  • Gen Agent Trust Hub8d

    The skill provides a comprehensive and secure guide for developing and publishing VS Code extensions. It follows security best practices, particularly concerning secret management and webview implementation.

  • Socket8d

    No alerts

  • Snyk8d

    Risk: LOW · No issues

  • Runlayer7mo

    9/9 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 7ae9078. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated last week
argument-hint
作りたい拡張機能、追加したい機能、困っている点
user-invocable
true
metadata
{
  "author": "yamapan (https://github.com/aktsmm)"
}

README badge

README badge for aktsmm/agent-skills/vscode-extension-guide