All skills

Regression test rules for AI-assisted work. Test API routes without a database. Add a test for each fixed bug. Check that sandbox and live paths return the same shape.

  • 1 file
  • 12.6 KB
  • Updated last month
  • GitHub

Use this Skill: https://skilld.dev/gh/agenticluke/ai-regression-guard-plus/skill

This session only. Nothing lands on disk.

SKILL.md

≈44 tokens always: the name and description. ≈3.2k when used: this file.

AI Regression Testing

AI tools can make the same mistake more than once. This happens when one AI writes the code and reviews it. Both steps may use the same wrong idea.

Tests give the code a separate check.

Use This Skill When

Use this skill when:

  • An AI changes an API route or server code.
  • A bug is found and fixed.
  • The app has sandbox, mock, or test data.
  • The app has more than one code path.
  • A code review or /bug-check starts.
  • A field is added to a query or API reply.
  • A UI change updates data before the server replies.

Main Rule

Write a test for the bug that was found.

Do not test only the normal case. Make the test fail in the exact way the old bug failed.

Use this order:

  1. Write a test that shows the bug.
  2. Run it and confirm that it fails.
  3. Fix the code.
  4. Run the test and confirm that it passes.
  5. Run all tests.
  6. Run the build or type check.

A test that passed before the fix may not prove that it can catch the bug.

Common AI Blind Spot

An AI may write and review code with the same wrong idea:

AI writes a fix
  -> AI reviews the fix
  -> The code looks right to the AI
  -> The bug is still there

A common case looks like this:

Fix 1: Add notification_settings to an API reply
  -> The database query does not load that field

Fix 2: Add the field to the query
  -> The type file does not include the field

Fix 3: Fix the live path
  -> The sandbox path still misses the field

Fix 4: Add a test for the field
  -> The test catches the bug on the next change

The key risk is a mismatch between sandbox and live code.

Test Without a Database

Use sandbox or mock data when you can. These tests are fast and safe. They should not call a real database or outside service.

Tests must not:

  • Use live keys or user data.
  • Send email, posts, or payments.
  • change live records.
  • depend on the network.
  • depend on test order.
  • share changed state between tests.

Reset mock data and environment values after each test when a test changes them.

Vitest and Next.js Setup

// vitest.config.ts
import path from "node:path";
import { defineConfig } from "vitest/config";

export default defineConfig({
  test: {
    environment: "node",
    globals: true,
    include: ["__tests__/**/*.test.ts"],
    setupFiles: ["__tests__/setup.ts"],
    restoreMocks: true,
  },
  resolve: {
    alias: {
      "@": path.resolve(__dirname, "."),
    },
  },
});
// __tests__/setup.ts
// Force local sandbox mode. Do not use a real database.
process.env.SANDBOX_MODE = "true";
process.env.NEXT_PUBLIC_SUPABASE_URL = "";
process.env.NEXT_PUBLIC_SUPABASE_ANON_KEY = "";

Read environment values inside the route when possible. A value read at file load time may not change during a test.

Request Helpers

// __tests__/helpers.ts
import { NextRequest } from "next/server";

type RequestOptions = {
  method?: string;
  body?: unknown;
  headers?: Record<string, string>;
  sandboxUserId?: string;
};

export function createTestRequest(
  url: string,
  options: RequestOptions = {},
): NextRequest {
  const {
    method = "GET",
    body,
    headers = {},
    sandboxUserId,
  } = options;

  const fullUrl = url.startsWith("http")
    ? url
    : `http://localhost:3000${url}`;

  const requestHeaders = new Headers(headers);

  if (sandboxUserId) {
    requestHeaders.set("x-sandbox-user-id", sandboxUserId);
  }

  const init: RequestInit = {
    method,
    headers: requestHeaders,
  };

  if (body !== undefined) {
    requestHeaders.set("content-type", "application/json");
    init.body = JSON.stringify(body);
  }

  return new NextRequest(fullUrl, init);
}

export async function parseJsonResponse(response: Response) {
  const text = await response.text();

  return {
    status: response.status,
    json: text ? JSON.parse(text) : null,
  };
}

Using body !== undefined keeps valid values such as false, 0, and null.

Test the API Contract

An API contract is the set of fields and value types that callers need.

// __tests__/api/user/profile.test.ts
import { describe, expect, it } from "vitest";
import { GET } from "@/app/api/user/profile/route";
import {
  createTestRequest,
  parseJsonResponse,
} from "../../helpers";

const REQUIRED_FIELDS = [
  "id",
  "email",
  "full_name",
  "phone",
  "role",
  "created_at",
  "avatar_url",
  "notification_settings",
] as const;

describe("GET /api/user/profile", () => {
  it("returns the required fields", async () => {
    const response = await GET(
      createTestRequest("/api/user/profile"),
    );
    const { status, json } = await parseJsonResponse(response);

    expect(status).toBe(200);
    expect(json.data).toBeTypeOf("object");

    for (const field of REQUIRED_FIELDS) {
      expect(json.data).toHaveProperty(field);
    }
  });

  it("keeps notification_settings defined for BUG-R1", async () => {
    const response = await GET(
      createTestRequest("/api/user/profile"),
    );
    const { json } = await parseJsonResponse(response);

    expect(json.data.notification_settings).not.toBeUndefined();

    const value = json.data.notification_settings;
    expect(value === null || typeof value === "object").toBe(true);
  });
});

Check more than field names when the value matters. Check its type, allowed null value, default value, and nested fields.

Also test:

  • Missing or bad input.
  • Missing user access.
  • Empty lists.
  • Not-found replies.
  • Server errors.
  • Bad JSON.
  • Duplicate requests, if they may happen.
  • Large or odd values.
  • Fields that may be null.
  • Correct status codes.

Test Sandbox and Live Paths

The best test runs both paths with fake data and compares their shapes.

Do not connect to the live database. Mock the data layer for the live path.

function keysOf(value: Record<string, unknown>) {
  return Object.keys(value).sort();
}

it("uses the same reply fields in sandbox and live paths", async () => {
  const sandboxData = {
    id: "user-001",
    email: "sam@example.com",
    name: "Sam",
    notification_settings: null,
  };

  const fakeLiveData = {
    id: "user-001",
    email: "sam@example.com",
    name: "Sam",
    notification_settings: { email: true },
  };

  expect(keysOf(sandboxData)).toEqual(keysOf(fakeLiveData));
});

If both paths cannot run in one test, create one shared list of required fields. Use that list in tests for both paths.

For list replies, do not skip the check when the list is empty. Seed at least one fake row first.

it("includes partner_name in each sandbox message", async () => {
  const response = await GET(
    createTestRequest("/api/user/messages", {
      sandboxUserId: "user-001",
    }),
  );
  const { status, json } = await parseJsonResponse(response);

  expect(status).toBe(200);
  expect(json.data.length).toBeGreaterThan(0);

  for (const message of json.data) {
    expect(message).toHaveProperty("partner_name");
  }
});

Bug Check Steps

Run checks in this order:

  1. Run the test for the fixed bug.
  2. Run the full test set.
  3. Run the build or type check.
  4. Review each changed code path.
  5. Add or update a test for every fixed bug.

Example command file:

# Bug Check

## Step 1: Run checks

Run:

    npm run test
    npm run build

Stop the review if either command fails.

Report a failed test or build as a top bug. Include the file, test name, and short error text.

## Step 2: Review the code

Check:

1. Sandbox and live paths return the same shape.
2. API replies match what the UI reads.
3. Each database query loads every used field.
4. Error paths clear old data.
5. Early returns use the right status and reply shape.
6. Fast UI updates roll back after a failed request.
7. New fields are covered by types, mocks, and tests.

## Step 3: Add tests

For each fixed bug:

1. Add a test that fails on the old code.
2. Name the bug in the test.
3. Run the test before and after the fix.
4. Run all tests and the build.

Common Bug Patterns

1. Sandbox and Live Paths Do Not Match

Bad:

if (isSandboxMode()) {
  return {
    data: { id, email, name },
  };
}

return {
  data: {
    id,
    email,
    name,
    notification_settings,
  },
};

Good:

if (isSandboxMode()) {
  return {
    data: {
      id,
      email,
      name,
      notification_settings: null,
    },
  };
}

return {
  data: {
    id,
    email,
    name,
    notification_settings,
  },
};

Check every path, including feature flags, user roles, cached replies, and old API versions.

2. A Query Misses a Field

Bad:

const { data } = await supabase
  .from("users")
  .select("id, email, name")
  .single();

return {
  data: {
    ...data,
    notification_settings: data.notification_settings,
  },
};

Good:

const { data } = await supabase
  .from("users")
  .select("id, email, name, notification_settings")
  .single();

Use an exact field list when you can. It shows which data the route needs. If you use select("*"), check that extra private fields are not sent to the client.

Also update:

  • Generated types.
  • Mock rows.
  • Sandbox rows.
  • API reply types.
  • Tests.
  • UI code that reads the field.

3. Old Data Stays on an Error

Bad:

catch {
  setError("Failed to load");
}

Good:

catch {
  setReservations([]);
  setError("Failed to load");
}

Test that old data is gone after a failed load. Also clear old errors before a new try.

4. A Fast UI Update Has No Rollback

Bad:

const handleRemove = async (id: string) => {
  setItems((oldItems) =>
    oldItems.filter((item) => item.id !== id),
  );

  await fetch(`/api/items/${id}`, {
    method: "DELETE",
  });
};

Good:

const handleRemove = async (id: string) => {
  const oldItems = items;

  setItems((currentItems) =>
    currentItems.filter((item) => item.id !== id),
  );

  try {
    const response = await fetch(`/api/items/${id}`, {
      method: "DELETE",
    });

    if (!response.ok) {
      throw new Error("Delete failed");
    }
  } catch {
    setItems(oldItems);
    setError("Could not remove the item");
  }
};

Test both success and failure. If two edits can run at once, an old saved list may undo a newer edit. Use a request ID, a cache tool, or a safe state update for that case.

5. Empty Data Hides a Bad Test

Bad:

if (json.data.length > 0) {
  expect(json.data[0]).toHaveProperty("partner_name");
}

This test passes when no rows exist.

Good:

expect(json.data.length).toBeGreaterThan(0);

for (const item of json.data) {
  expect(item).toHaveProperty("partner_name");
}

Seed the row that the test needs.

6. A Mock Is Better Than the Real Code

A mock may include a new field even when the real query does not load it. This can make a bad change look safe.

Keep mock rows close to the real data shape. Add a test for the query field list when possible. Do not mock the code that the test is meant to check.

Concrete Example

A bug report says:

The profile page has no notification settings in sandbox mode.

Use this plan:

  1. Add notification_settings to the required field list.
  2. Add a test named keeps notification_settings defined for BUG-R1.
  3. Run only that test. Confirm that it fails.
  4. Add the field to the sandbox reply.
  5. Check that the live query also loads the field.
  6. Update types and fake rows.
  7. Run:
npm run test -- __tests__/api/user/profile.test.ts
npm run test
npm run build
  1. Keep the test after the fix. It now guards the bug.

Test Naming

Use a name that states the rule and the old bug.

Good:

it("returns notification_settings in sandbox mode for BUG-R1", async () => {
  // Test code
});

Weak:

it("works", async () => {
  // Test code
});

A short bug ID helps people find the old report. Do not put secret data or user names in test files.

Final Check

Before marking the work done, confirm:

  • The new test failed on the old code.
  • The new test passes on the fixed code.
  • The full test set passes.
  • The build or type check passes.
  • Sandbox and live paths have the same reply shape.
  • Queries, types, mocks, and UI code agree.
  • Error tests check status codes and reply shapes.
  • Tests use no live database or outside service.
  • Each fixed bug has a test that can catch it again.

Aim for useful tests, not a coverage score. Add tests where real bugs were found and where two code paths can drift apart.

Source: SKILL.md on GitHub

No third-party reports yet.

Signed by skilld at 6777b2f. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last month.

Activeupdated last month
origin
ECC

README badge

README badge for agenticluke/ai-regression-guard-plus