ui-test — Agentic UI Testing Skill
Adversarial UI testing that catches what Playwright can't. Analyzes git diffs to test only what changed, or explores the full app to find bugs. Runs in a real browser via the browse CLI.
Install
npx skills add browserbase/ui-testQuick Start
"Test the UI changes in my PR" → diff-driven (Workflow A)
"Explore my app and find bugs" → exploratory (Workflow B)
"QA the staging site in parallel" → parallel Browserbase sessions (Workflow C)How It Works
- Analyzes the diff (or explores the app) to decide what to test
- Opens a real browser — clean isolated local browser for localhost, Browserbase for deployed sites
- Tries to break things — adversarial inputs, rapid clicks, keyboard-only, empty states, XSS
- Runs deterministic checks — axe-core, console errors, broken images, form labels
- Reports structured results —
STEP_PASS|id|evidenceorSTEP_FAIL|id|expected → actual - Generates an HTML report — standalone file with embedded screenshots, shareable with reviewers
What It Tests
| Category | How | What Playwright Misses |
|---|---|---|
| Accessibility | axe-core + keyboard nav | WCAG violations, focus rings, screen reader semantics |
| Visual Quality | Screenshot + Claude judgment | Layout balance, typography, spacing, empty states |
| Responsive | Viewport sweep (375px, 768px, 1440px) | Mobile overflow, touch targets, content reflow |
| Console Health | browse eval injection |
Hydration errors, failed requests, runtime exceptions |
| Error States | Navigate to empty/error states | Missing empty states, broken error recovery |
| Adversarial | XSS, empty submit, rapid click, long input | Edge cases developers don't write tests for |
| Exploratory | Navigate freely, try to break things | Bugs you didn't think to test for |
Browser Execution
which browse || npm install -g browse- Localhost →
browse open <url> --local(no API key needed) - Need existing local login/cookies/state on localhost →
browse open <url> --auto-connect(requires an existing debuggable local Chrome) - Need explicit local CDP attach →
browse open <url> --cdp <port|url> - Deployed sites →
browse open <url> --remote(uses Browserbase cloud browsers) - Parallel →
BROWSE_SESSION=<name>for independent concurrent sessions
For default localhost QA, pass --local on the first browse open for clean, reproducible runs.
Project Structure
ui-test/
├── SKILL.md # Skill definition — workflows, assertion protocol, budget
├── EXAMPLES.md # 9 worked examples with exact commands
├── README.md
└── references/
├── adversarial-patterns.md # Adversarial test patterns (forms, modals, nav, keyboard)
├── browser-recipes.md # Copy-paste browse CLI recipes for deterministic checks
├── design-consistency.md # Design consistency checking methodology
├── design-system.example.md # Example design system template (copy to design-system.md)
├── exploratory-testing.md # Guide for agent-driven exploratory QA
├── parallel-testing.md # Parallel testing with named Browserbase sessions
├── report-template.html # HTML report template with embedded screenshots
└── ux-heuristics.md # 6 evaluation frameworks (Laws of UX, Nielsen's, etc.)Philosophy
Traditional tests verify intentions. This skill finds blind spots.
No YAML files, no generated test suites, no artifacts. The agent reads the diff (or explores the app), opens a browser, tries to break things, and reports what it found. Like a human QA tester with perfect knowledge of every design principle.
Requirements
browseCLI (npm install -g browse)- For remote testing:
BROWSERBASE_API_KEYenvironment variable - A running web app (localhost or deployed URL)