All skills
simota avatar

/matrix

@c805268
by shingo imotasimota/agent-skills85 stars
15

Controlling combinatorial explosion across multi-dimensional axes: minimum coverage sets, execution plans, test/deploy/UX/risk prioritization. Use when scoping multi-axis combinations.

Use this Skill: https://skilld.dev/gh/simota/agent-skills/matrix

This session only. Nothing lands on disk.

referencefault-interaction-statistics.md

≈1.7k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Fault Interaction Statistics

Purpose: Use this file when deciding whether 2-way, 3-way, 4-way, or mixed-strength coverage is justified.

Contents

  • Interaction distributions
  • Strength selection
  • Escalation strategy
  • Mixed strength
  • Coverage targets

Interaction Distributions

NIST-style evidence summary:

Domain <=2-way <=3-way <=4-way <=5-way <=6-way
Medical devices 97% 99% 100% - -
Browsers 70% 90% 95% 99% 100%
Servers 76% 95% 99% 100% -
General software 70-95% 90-99% 97-100% - -

Operational takeaway:

  • 2-way is enough for many normal systems.
  • 3-way+ is justified once history, regulation, or criticality says pairwise is insufficient.

Strength Selection

Strength Typical detection Relative cost Use when
2-way 70-95% baseline normal applications
3-way 90-99% 2-3x high-quality or interaction-heavy areas
4-way 97-100% 3-5x safety-critical or regulated systems
5-way+ 99-100% very high exceptional regulatory cases

Escalation Strategy

Use this escalation path:

  1. start with 2-way
  2. if additional higher-order defects appear, move to 3-way
  3. if 3-way still exposes meaningful new defects, move to 4-way
  4. stop escalating when the next step yields no meaningful new defects

Mixed Strength

Use mixed strength when only part of the model is high risk.

Example:

  • authentication x privilege x data sensitivity -> 3-way
  • browser x OS -> 2-way
  • locale or theme -> 1-way sampling

This often delivers better cost efficiency than applying 3-way to the entire matrix.

Coverage Targets

Domain Minimum Preferred
General web application 2-way 100% 2-way 100%
Finance 2-way 100% 3-way 100%
Medical or safety-critical 3-way 100% 4-way 100%
IoT / embedded 2-way 100% 3-way 100%
Security testing 2-way 100% mixed strength with 3-way for high-risk areas

Quality gates:

  • 2-way coverage below 100% -> warning
  • safety-critical domain + 2-way only -> escalate recommendation

2025–2026 Research Updates

High-Strength CIT for Configurable Systems

The ISSTA 2024 paper "Beyond Pairwise Testing: Advancing 3-wise Combinatorial Interaction Testing for Highly Configurable Systems" (ACM SIGSOFT) introduces ScalableCA, a constrained covering array generator that produces 3-wise arrays 38.9% smaller than prior SOTA while running 1–2 orders of magnitude faster. Techniques: fast invalidity detection, uncovering-guided sampling, remainder-aware local search. Source: https://dl.acm.org/doi/10.1145/3650212.3680309

The ICSE 2025 paper "Towards High-Strength Combinatorial Interaction Testing for Highly Configurable Software Systems" extends scalable CCAG to 4-wise and 5-wise for large parameter models where prior algorithms were intractable. Source: https://dl.acm.org/doi/10.1109/ICSE55347.2025.00113

Implication: For highly configurable software (feature flags, plugin systems, product-line architectures), 3-way+ CIT is now computationally feasible at production scale — escalate from 2-way sooner when historical defect patterns or domain criticality justify it.

Combinatorial Security Testing — 10 Years Later (2026)

Simos, Leithner, Kuhn, Garn, Kacker & Lei published a 2026 retrospective in IEEE Security & Privacy documenting a decade of combinatorial security testing (CST) in practice. Key updates:

  • CST scope has expanded from input validation to cloud configurations, IoT firmware, and API security surfaces.
  • Mixed-strength models (3-way for auth × privilege × data-sensitivity, 2-way elsewhere) remain the recommended practice.
  • Constraint modeling quality is the dominant factor in CST effectiveness — over-constrained models produce false confidence. Source: NIST CSRC project page — https://csrc.nist.gov/projects/automated-combinatorial-testing-for-software

AI/ML Dataset Coverage (2025)

Kuhn, Raunak & Kacker, "Measuring and Visualizing Dataset Coverage for Machine Learning", IEEE Computer vol 58 no 4, Mar 2025 — introduces visualization methods for feature-interaction frequency distributions in training data, making data skew detectable before model training. Source: https://www.nist.gov/publications/combinatorial-testing-metrics-machine-learning

NIST CSRC Apr 2025, "Data Frequency Coverage Impact on AI Performance" — pilot study shows: (1) performance may increase or decrease with data skew; (2) feature importance methods do not predict skew impact; (3) adding more data does not reliably mitigate skew effects. Use combinatorial frequency coverage, not raw dataset size, as the quality gate for ML training sets. Source: https://csrc.nist.gov/pubs/conference/2025/04/15/data-frequency-coverage-impact-on-ai-performance/final


Core Contract Long Form (SKILL.md excerpt)

  • Apply the NIST interaction rule: 93% of real-world faults are triggered by ≤ 2-way interactions, 98% by ≤ 3-way, nearly 100% by ≤ 6-way (Kuhn, Wallace & Gallo 2004; NASA/NIST empirical data across distributed systems, medical devices, browser, and server applications). Use this to justify strength selection.

  • For AI/ML dataset coverage, use data frequency coverage — not just tuple presence — to detect training data skew. Simple combinatorial coverage misses imbalanced feature interaction frequencies that degrade model performance (Kuhn, Raunak & Kacker, IEEE Computer Mar 2025, "Measuring and Visualizing Dataset Coverage for Machine Learning"; NIST CSRC Apr 2025, "Data Frequency Coverage Impact on AI Performance").

  • For highly configurable systems requiring 3-way+ coverage, apply scalable CCAG algorithms (e.g., ScalableCA from ISSTA 2024) that deliver 3-wise arrays 38.9% smaller than prior SOTA with 1–2 orders of magnitude faster construction — making high-strength CIT practical for large parameter models (ICSE 2025: "Towards High-Strength CIT for Highly Configurable Software Systems").

  • When applying combinatorial security testing, reference the decade of field evidence: CST has expanded from input validation to cloud, IoT, and firmware surfaces; the 2026 "Combinatorial Security Testing—10 Years Later" review (Simos et al., IEEE Security & Privacy) updates deployment guidance.

  • When parameter modeling is expensive or incomplete, AI-assisted parameter extraction (e.g., Hexawise AI Guidance / Sembi iQ, 2025) can draft parameter/value models from specification documents, accelerating the PARSE phase without replacing engineer review. Treat AI-generated models as first-draft; validate constraints before optimizing.

Source: SKILL.md on GitHub

1 warning13d5 checks · Risk SAFE
  • Gen Agent Trust Hub13d

    The Matrix skill is a comprehensive tool for combinatorial testing design, providing robust frameworks for pairwise and high-strength interaction testing based on NIST and academic standards. It focuses on generating optimized execution plans and risk-weighted coverage sets without possessing any capabilities for code execution, network exfiltration, or unauthorized file access. No security risks were identified within the skill's instructions or reference materials.

  • Socket13d

    No alerts

  • Snyk13d

    Risk: LOW · No issues

  • Runlayer6mo

    1/11 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at c805268. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 3 days ago.

Activeupdated 2 weeks ago

README badge

README badge for simota/agent-skills/matrix