All skills
sanity-io avatar

/content-experimentation-best-practices

@ad50ea4 official
by Sanitysanity-io/agent-toolkit187 stars
30

Content experimentation and A/B testing guidance covering experiment design, hypotheses, metrics, sample size, statistical foundations, CMS-managed variants, and common analysis pitfalls. Use this skill when planning experiments, setting up variants, choosing success metrics, interpreting statistical results, or building experimentation workflows in a CMS or frontend stack.

Use this Skill: https://skilld.dev/gh/sanity-io/agent-toolkit/content-experimentation-best-practices

This session only. Nothing lands on disk.

referencesstatistical-foundations.md

≈1.3k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Statistical Foundations

Understanding basic statistics prevents misinterpreting experiment results.

Table of Contents

  • Key concepts
  • Sample size calculation
  • Common statistical mistakes
  • Interpreting results
  • Alternative approaches
  • When to trust results

Key Concepts

Statistical Significance

A measure of whether observed differences are likely real or due to chance.

  • p-value < 0.05: "Statistically significant" at 95% confidence
  • Means: If there were no real difference, there's less than a 5% chance of seeing results this extreme
  • Does NOT mean: The change is important or meaningful
  • Common misconception: The p-value is NOT "the probability the result is due to chance." It's the probability of observing data this extreme assuming the null hypothesis is true.

Confidence Interval

A range of plausible values for the true effect.

Example: "Conversion rate increased by 5% (95% CI: 2% to 8%)"

  • Best estimate: 5% improvement
  • Could be as low as 2% or as high as 8%
  • Narrower intervals = more certainty

Statistical Power

The ability to detect a real effect when it exists.

  • Standard: 80% power
  • Higher power = larger sample size needed
  • Low power = might miss real improvements

Minimum Detectable Effect (MDE)

The smallest improvement worth detecting.

  • Smaller MDE = larger sample size needed
  • Be realistic: Can you act on a 0.5% improvement?

Sample Size Calculation

Before running a test, calculate required sample size:

Required per variant = 16 × σ² / MDE²

Where:
- σ² = variance (for conversion rate: p × (1-p))
- MDE = minimum detectable effect (absolute)

For a 5% baseline conversion rate, detecting a 1% absolute lift (5% → 6%):

  • σ² = 0.05 × 0.95 = 0.0475
  • MDE² = 0.01² = 0.0001
  • n = 16 × 0.0475 / 0.0001 = 7,600 per variant
  • Total: ~15,200 visitors minimum

Common Statistical Mistakes

Multiple Comparisons Problem

Testing 10 variants increases false positive rate.

Solution: Adjust significance threshold (Bonferroni correction) or use sequential testing methods.

Peeking Problem

Checking results daily and stopping when significant.

Why it's wrong: Significance fluctuates. Early "winners" often regress.

Solution: Pre-commit to sample size and duration. Use sequential testing if you must peek.

Simpson's Paradox

Overall results hide segmented truths.

Example:

  • Overall: Variant B wins
  • Mobile users: Variant A wins
  • Desktop users: Variant A wins
  • How? Different traffic mix per variant

Solution: Always segment by major factors (device, traffic source).

Survivorship Bias

Only analyzing users who completed the funnel.

Solution: Include all visitors, not just converters.

Interpreting Results

Significant + Meaningful

Clear win. Implement the change.

Significant + Trivial

Statistically different but tiny effect. Consider if worth the complexity.

Not Significant + Large Effect

Might be real but underpowered. Extend the test or accept uncertainty.

Not Significant + Small Effect

No detectable difference. Either no real effect or test was underpowered.

Alternative Approaches

Bayesian A/B Testing

An alternative to traditional (frequentist) hypothesis testing. Bayesian methods provide:

  • Direct probability statements: "There's a 95% probability Variant B is better" (more intuitive than p-values)
  • No peeking problem: Continuous monitoring is built in — you can check results at any time
  • Credible intervals: Directly interpretable as "the true value falls in this range with X% probability"

Bayesian methods are offered by platforms like VWO and are useful when you need to make decisions with limited traffic or want more intuitive reporting for stakeholders.

Multi-Armed Bandits

Dynamically allocate more traffic to winning variants while still learning:

  • Thompson Sampling: Balances exploration (learning) with exploitation (serving the best variant)
  • Best for: Ongoing optimization where you want to minimize regret during the test
  • Trade-off: Faster convergence to the winner, but less statistical rigor than fixed-allocation A/B tests

Consider bandits for content recommendations, personalization, or situations where the cost of showing a losing variant is high.

Sequential Testing

For teams that need to monitor experiments continuously:

  • Group sequential designs (O'Brien-Fleming, Lan-DeMets) allow pre-planned interim analyses
  • Always-valid p-values let you check results at any time without inflating false positive rates
  • Use when you must balance the peeking problem with business pressure to act on results quickly

When to Trust Results

Checklist before declaring a winner:

  • Reached pre-calculated sample size
  • Ran for full business cycle (1-2 weeks minimum)
  • p-value < 0.05 (or your chosen threshold)
  • Effect size is meaningful for business
  • Results consistent across major segments
  • No external factors contaminated results

Source: SKILL.md on GitHub

No alerts16d5 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    The skill provides comprehensive guidelines and best practices for content experimentation and A/B testing. It includes illustrative code snippets for CMS integration and statistical calculation. No security risks were identified.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: LOW · No issues

  • Runlayer6mo

    5 files scanned · No issues

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at ad50ea4. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 weeks ago.

Activeupdated 6 months ago
  • a-b-testing
  • content-experimentation
  • statistical-analysis
  • cms
  • conversion-optimization
  • multivariate-testing
  • metrics
  • hypothesis-testing

README badge

README badge for sanity-io/agent-toolkit/content-experimentation-best-practices

Provides guidance on A/B testing, multivariate testing, and statistical analysis for content experiments, including experiment design, metrics selection, sample sizing, CMS integration patterns, and common pitfalls. Use when setting up experimentation infrastructure, designing content variants, or interpreting test results in a headless CMS or frontend stack.

Generated from the current SKILL.md.

Does this skill cover statistical rigor for A/B tests?
Yes. The skill includes statistical foundations covering p-values, confidence intervals, power analysis, and Bayesian methods to help interpret results correctly.
Can I use this skill to set up experiments in a headless CMS?
Yes. The skill includes guidance on CMS-managed variants and field-level variants, with patterns for integrating experimentation into CMS workflows.
What common mistakes does this skill help avoid?
The skill documents 17 common pitfalls across statistics, design, execution, and interpretation to help teams avoid typical experimentation errors.
Does this cover multivariate testing or just A/B tests?
Both. The skill covers A/B testing, multivariate testing, and how to design experiments that test multiple variables simultaneously.

Generated from the current SKILL.md. These answers refresh after source changes.