All skills
simota avatar

/siege

@35ffd55
by shingo imotasimota/agent-skills85 stars
15

Verifying system resilience via load testing, contract testing, chaos engineering, and mutation testing. Use for limit verification, non-functional testing, or reliability validation.

Use this Skill: https://skilld.dev/gh/simota/agent-skills/siege

This session only. Nothing lands on disk.

referencetest-strategy-2026.md

≈4.3k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Test Strategy 2026

Purpose: load this when designing a test strategy from scratch, evaluating a team's current test mix, or recommending which complementary tools to add. Consolidates the 2026 picture across unit / integration / E2E / load / chaos / contract / mutation / property / metamorphic / trace-based testing — and shows how radar, voyager, siege, attest, and radar slot together.

Contents

  1. Shape of the test mix (pyramid / diamond / trophy)
  2. Coverage as floor, mutation as ceiling
  3. The seven complementary test layers
  4. Property-based + Mutation + Metamorphic
  5. Trace-based testing (Tracetest)
  6. Schemathesis (stateful API fuzz)
  7. Production traffic replay (GoReplay / Speedscale)
  8. Chaos engineering (LitmusChaos / Gremlin / AWS FIS)
  9. Contract testing (Pact / PactFlow HaloAI)
  10. Synthetic monitoring × E2E convergence
  11. The 2026 anti-patterns
  12. Skill-to-layer mapping

1. Shape of the Test Mix

Shape Layer When to use
Pyramid unit > integration > E2E Backend with stable contracts and slow CI; the classic Mike Cohn shape
Diamond integration > unit + E2E Service-heavy systems where unit tests test mocks more than logic
Trophy (Kent C. Dodds) static + unit + integration (largest) + E2E Modern web stacks with cheap static analysis and cheap integration
Honeycomb (Spotify) integration tests cluster + unit perimeter Microservice meshes with mature contract tests

The 2026 consensus is "no single shape fits all" — pick by team's CI budget, change frequency, and existing contract surface. The pyramid still anchors backend-heavy stacks; the trophy is the right shape for frontend / RSC / Vite codebases.

Source: kentcdodds.com/blog/the-testing-trophy-and-testing-classifications; martinfowler.com/articles/practical-test-pyramid.html.

2. Coverage as Floor, Mutation as Ceiling

Coverage is a Goodhart-vulnerable floor metric — target → tautological tests. Mutation Score is the ceiling metric — it measures whether tests actually catch defects.

Metric Role Recommended threshold Why
Line coverage floor ≥ 80% Guarantees the test suite at least executes most lines
Branch coverage floor ≥ 70% Both true and false outcomes of conditionals
Mutation score (Stryker / mutmut / Pitest) ceiling break: 50 / low: 60 / high: 80 Defects caught vs defects possible

Teams hitting mutation high: 80 in CI report ~70% fewer production bugs vs coverage-only teams. Apply mutation gate to changed files only (incremental mutation) to keep CI under 5 minutes.

Source: stryker-mutator.io/docs; medium.com/@jaychopra05 — 100% Code Coverage Is a Lie.

3. The Seven Complementary Test Layers

# Layer Catches Tooling 2026 Skill
1 Unit + property-based logic errors, branch gaps Vitest / pytest + Hypothesis / fast-check / proptest radar
2 Mutation weak assertions Stryker / mutmut / Pitest radar
3 Metamorphic oracle-problem inputs hand-rolled relations + PBT radar
4 Integration + Contract inter-service drift Pact / PactFlow HaloAI / Schemathesis siege + attest
5 Trace-based internal-behaviour gaps Tracetest + OTel siege + attest
6 E2E + Visual + a11y user-flow regressions Playwright Test Agents / Maestro / axe-core + IGT / Argos voyager
7 Load + Chaos + Replay non-functional limits k6 v1.0 / LitmusChaos / GoReplay siege

Each layer catches a class the previous layers miss. Skip layers deliberately; document the rationale (e.g. "no inter-service traffic, contract layer skipped").

4. Property-Based + Mutation + Metamorphic

The three reinforce each other:

  • Property-based testing generates inputs from a domain hypothesis (e.g. for all x: sort(x) is sorted AND sort(x) is a permutation of x).
  • Mutation testing verifies the suite is strong enough to catch the bugs it claims to prevent.
  • Metamorphic testing supplies the oracle when the expected output is intractable: sort(reverse(xs)) ≡ sort(xs), serialize(deserialize(s)) ≡ s (round-trip), f(x + 0) ≡ f(x).

PBT generates inputs; Metamorphic Relations supply the oracle; Mutation testing verifies they both bite. Adoption of Metamorphic Relations is still low in the LLM-testing literature (4 of 36 oracle-automation studies), making it a high-leverage axis to introduce.

Source: dl.acm.org/doi/10.1145/3798226; arxiv.org/html/2405.12766v1.

5. Trace-based Testing (Tracetest)

When the requirement is internal behaviour — "on submit, the audit log is written AND the cache is invalidated AND no PII is logged" — an HTTP-only verifier cannot prove it. Tracetest asserts on individual OpenTelemetry spans inside the trace.

# tracetest spec
spec:
  - selector: span[name="audit.write"]
    assertions: [ "attr:audit.event = 'order.created'" ]
  - selector: span[name="cache.invalidate"]
    assertions: [ "exists" ]
  - selector: span[name="log.write"]
    assertions: [ "not(attr:body ~= 'user.email')" ]

Combine with Playwright (Voyager) for UI-driven trace assertions. The right tool when "the response was 200" hides a broken internal call.

Source: tracetest.io; oneuptime.com/blog/post/2026-02-06-tracetest-playwright-browser-testing/view.

6. Schemathesis (Stateful API Fuzz)

Drive verification from the OpenAPI / GraphQL spec. The property-based engine explores state transitions automatically; published benchmarks show 1.4-4.5× more defects than peer tools.

# Stateful test: explore valid state transitions from the spec
schemathesis run --checks all --hypothesis-database=:memory: \
  --stateful=links openapi.yaml --base-url https://staging.example.com

Two complementary verifications:

  • Schemathesis output → spec-vs-implementation conformance
  • Hand-authored BDD → intent-vs-implementation conformance

Both are required for full verdict. Schemathesis catches spec violations the team did not write a scenario for; BDD catches intent violations the spec did not enumerate.

Source: schemathesis.io; apideck.com/blog/openapi-testing.

7. Production Traffic Replay (GoReplay / Speedscale)

Synthetic generators miss edge cases and skew. Production traffic replay solves this:

Tool Strength
GoReplay OSS, HTTP/gRPC, low-overhead passive capture
Speedscale Commercial, PII auto-scrub (GDPR-safe), Gartner-listed

Use for any service whose load shape is hard to model:

  • Authentication services (skewed cohort distributions)
  • Search / recommendation (long-tail query distributions)
  • Webhook receivers (bursty, unpredictable shape)

Replay also serves dual duty as a fixture source (radar strength): record once, replay across staging or test seeds.

Source: goreplay.org/shadow-testing; speedscale.com/blog/definitive-guide-to-traffic-replay.

8. Chaos Engineering

Tool selection 2026:

Tool Strength When
LitmusChaos Kubernetes-native, MCP Server (2026) K8s clusters with agent-driven experiments
AWS FIS Native to AWS, hardware-fault primitives AWS-heavy stacks
Gremlin Commercial, broad cloud coverage Multi-cloud, regulated industries
Steadybit Service catalogue + experiment library Service-mesh maturity already in place
Chaos Mesh OSS, Kubernetes CRD-driven K8s, prefer YAML-first

LitmusChaos's MCP Server is the 2026 differentiator for Claude / Codex / Cursor hosts — a natural-language inject pod-failure into payment-service for 5m launches and observes the experiment. CNCF Q4 2025 update confirms active development.

Source: litmuschaos.io/blog — Making Chaos Engineering Accessible (MCP Server); cncf.io — LitmusChaos Q4 2025.

9. Contract Testing (Pact / PactFlow HaloAI)

Consumer-Driven Contract testing remains the gold standard for inter-service drift detection. 2026 evolution:

  • Pact core remains the OSS standard
  • PactFlow HaloAI generates and maintains contracts from OpenAPI + observed traffic; reports ~60% maintenance-time reduction
  • Bi-directional contracts verify both consumer expectations and provider spec, catching the "spec drift" case Pact alone misses

Adoption rule of thumb: introduce contract testing when there are ≥ 3 independent services with distinct teams. Below that, integration tests suffice.

Source: pactflow.io/ai/; pactflow.io/bi-directional-contract-testing/.

10. Synthetic Monitoring × E2E Convergence

In 2026 the boundary between "E2E test in CI" and "synthetic check in production" collapses. Checkly runs Playwright suites as production synthetic checks; OpenTelemetry ties the synthetic span to the backend trace.

┌────────────── one Playwright suite ──────────────┐
│                                                  │
│   CI run       │   Synthetic run (Checkly)       │
│   ────────     │   ──────────────────────────    │
│   on PR        │   every 5 min in production     │
│   blocks merge │   pages oncall on failure       │
│   catches      │   catches production-only       │
│   regressions  │   issues (geo, infra, 3rd-party)│
└──────────────────────────────────────────────────┘

Voyager owns the suite design; Beacon owns the production monitoring deployment. Escalate the synthetic-deployment plan via Voyager → Beacon handoff.

Source: checklyhq.com; usenix.org/publications/loginonline — Synthetic Monitoring & E2E Testing: Two Sides of the Same Coin.

11. The 2026 Anti-Patterns

Anti-pattern Why it fails Fix
100% coverage as a goal Drives tautological tests; mutation score may be ~20% Coverage as floor, mutation as ceiling
AI generates both code and tests Blind spots are shared; mutation collapses Generator-Evaluator separation: different model or different role writes the test
Snapshot as the sole oracle Locks in current behaviour, including bugs At least one behavioural assertion per public path
"Retry: 2" for flaky tests Hides the cause; flake re-emerges in production FlakyGuard-class auto-repair + 6-class root-cause taxonomy
Page Object Mega-Class Inheritance bloat, implicit waits, brittle Screenplay Pattern for complex flows, POM for simple
"Test all the buttons" E2E Slow, flaky, low-defect-yield Trophy/diamond shape with integration as the bulk
Synthetic data with no FK preservation Orphan rows, false-pass tests MOSTLY AI / Gretel for FK-preserving generation

12. Skill-to-Layer Mapping

Layer Primary Skill Secondary
Unit + property + mutation + metamorphic radar builder (writes code that meets the properties)
Integration + contract siege attest (conformance verdict)
Trace-based siege + attest beacon (deploys the tracing)
E2E + Visual + a11y voyager radar (component-level testing)
Load + chaos + replay siege triage / mend (incident replay)
Test data + fixtures radar siege (load profile generation)
Spec compliance attest scribe[unified] / scribe (spec authoring)

Use this table when classifying an incoming test request — the primary skill owns the design, the secondary skill is the typical downstream handoff.


Tooling Recommendations Long Form (SKILL.md excerpt)

  • Default to k6 v1.0 with TypeScript-native execution for new load tests. v1.0 GA removed the xk6-ts requirement, executes .ts scripts directly, integrated browser load testing, and reports ~70% lower CPU vs v0.x. Recommend k6 v1.0 over Gatling / Artillery for greenfield unless a specific feature is missing (e.g. Gatling Java DSL). [Source: grafana.com/docs/k6/latest/using-k6/javascript-typescript-compatibility-mode/; infoq.com — Grafana k6 Releases]

  • Use Schemathesis for stateful API fuzz driven by OpenAPI/GraphQL specs. Property-based generation explores state transitions automatically; published benchmarks show 1.4-4.5× more defects than peer tools. Pair with Pact-style consumer-driven contract tests — Schemathesis covers spec-vs-implementation, Pact covers consumer-vs-provider. [Source: schemathesis.io; apideck.com/blog/openapi-testing]

  • Adopt trace-based testing (Tracetest) for distributed assertions. Tracetest asserts on individual OpenTelemetry spans, not just the HTTP response — proving that the auth service was called once, the cache was hit, the DB query took under N ms. Combine with Playwright for UI→trace assertions. The right tool when "the response was 200" hides a broken internal call. [Source: tracetest.io; oneuptime.com/blog/post/2026-02-06-tracetest-playwright-browser-testing/view]

  • Production traffic replay (GoReplay, Speedscale) as a load source. Passively record production HTTP/gRPC, replay against staging with PII auto-scrubbing (Speedscale, Gartner-listed). Replay-derived load profiles capture edge cases and skew that synthetic generators miss; recommend for any service whose load shape is hard to model. [Source: goreplay.org/shadow-testing; speedscale.com/blog/definitive-guide-to-traffic-replay]

  • LitmusChaos MCP Server for chaos via natural language. A Claude-compatible MCP client can launch and observe chaos experiments by name. Add to the chaos-tool selection table alongside AWS FIS / Gremlin / Steadybit when the host is an MCP-capable agent. [Source: litmuschaos.io/blog — Making Chaos Engineering Accessible: Introducing the LitmusChaos MCP Server; cncf.io/blog/2026/01/22 — LitmusChaos Q4 2025 Update]

  • PactFlow HaloAI / AI-augmented contract test maintenance. Generates and maintains Pact contracts from OpenAPI specs and observed traffic; reports ~60% time reduction for contract upkeep vs hand-written. Recommend on consumer-driven contract programmes where the maintenance burden is the bottleneck. [Source: pactflow.io/ai/]

  • MSW v2 as the contract-mock standard for frontend. Standard Fetch API handlers (http.get / http.post returning Response) make the mock the same shape as the contract under test; lets the same handler power Vitest unit tests, Cypress CT, and Storybook visual regression. Replace legacy nock/fetch-mock references with MSW v2 when the codebase already targets the Fetch API. [Source: mswjs.io/blog/introducing-msw-2.0/]

Per-Recipe Behavior Notes (SKILL.md excerpt)

  • load: Select LOAD mode. Verify throughput, latency, capacity, spike, and soak with k6/Locust/Artillery. Always report p50/p95/p99/max.
  • contract: Select CONTRACT mode. Verify consumer/provider contracts with Pact v4+ or Specmatic. Integrate into the CI gate.
  • chaos: Select CHAOS mode. Define steady state first, minimize blast radius, then inject faults. Always prepare a kill switch.
  • mutation: Select MUTATE mode. Generate mutants → classify survivors → evaluate coverage thresholds (60% project-wide / 75%+ recommended).
  • fuzz: Coverage-guided fuzzing of parsers, decoders, and security-sensitive surfaces with AFL++/libFuzzer/go-fuzz/cargo-fuzz/Jazzer. Always pair with a sanitizer (ASan+UBSan default), seed from a real corpus, and minimize+dedupe crashes before reporting. For unit-test coverage gaps use Radar; for test-data factory shapes use Radar; for deeper DAST on security-critical crashes hand off to Probe/Sentinel.
  • property: Property-based testing of invariants (round-trip, idempotent, monotonic, model-based) with fast-check/Hypothesis/jqwik/PropEr/proptest. Compose generators from primitives (no filter-heavy strategies), cap 100-1000 runs at PR tier, commit shrunk counter-examples as regression tests. For example-based unit tests use Radar; for realistic factory data use Radar; for AC-level conformance use Attest; for byte-level parser crashes use fuzz.
  • concurrency: Hunt invisible defects — race conditions (TSan/Helgrind), memory leaks (ASan/Valgrind/MSan/heap-diff), resource leaks (file/connection/goroutine), deadlocks (lock-order analysis), atomic-ordering bugs (Rust loom exhaustive interleavings, Java jcstress, C++ atomics). Pair with chaos to induce exhaustion conditions and property for ordering invariants. Output: defect class + reproduction trace + minimal repro + fix recommendation to Builder. Use when symptoms are flaky-only-under-load, "works on my machine", or sporadic CI failures.
  • smoke: Minimum viable post-deploy gate, 8-15 checks, ≤3 min budget, serial by default, synthetic-check-capable. Emits PROMOTE/HOLD/ROLLBACK verdict tied to deploy SHA. For full user-journey E2E use Voyager; for unit coverage use Radar; for AC compliance use Attest; for SLO ownership and long-term synthetic monitoring topology use Beacon.

Source: SKILL.md on GitHub

1 warning13d5 checks · Risk SAFE
  • Gen Agent Trust Hub13d

    The 'siege' skill is a comprehensive framework for system resilience testing using industry-standard tools. It identifies a potential indirect prompt injection surface since the skill processes external, untrusted data such as API responses and trace logs, which is typical for this type of utility. All external dependencies and tool references point to trusted organizations and well-known services.

  • Socket13d

    No alerts

  • Snyk13d

    Risk: LOW · No issues

  • Runlayer6mo

    3/10 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 35ffd55. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 days ago.

Activeupdated 2 weeks ago

README badge

README badge for simota/agent-skills/siege