All skills
sanity-io avatar

/content-experimentation-best-practices

@ad50ea4 official
by Sanitysanity-io/agent-toolkit187 stars
30

Content experimentation and A/B testing guidance covering experiment design, hypotheses, metrics, sample size, statistical foundations, CMS-managed variants, and common analysis pitfalls. Use this skill when planning experiments, setting up variants, choosing success metrics, interpreting statistical results, or building experimentation workflows in a CMS or frontend stack.

Use this Skill: https://skilld.dev/gh/sanity-io/agent-toolkit/content-experimentation-best-practices

This session only. Nothing lands on disk.

referencescommon-pitfalls.md

≈1.7k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Common Experimentation Pitfalls

Avoid these mistakes that invalidate results or lead to wrong conclusions.

Statistical Mistakes

1. Stopping Early (Peeking)

The problem: Checking results daily and stopping when you see significance.

Why it's wrong: Statistical significance fluctuates. At any point during a test, you might see "significance" that disappears with more data. This is called the "peeking problem" or "repeated significance testing."

The fix:

  • Pre-calculate required sample size
  • Commit to running until you reach it
  • If you must peek, use sequential testing methods that account for multiple looks

2. Underpowered Tests

The problem: Running tests without enough traffic to detect realistic effect sizes.

Why it's wrong: You'll conclude "no difference" when there actually is one—you just couldn't detect it.

The fix:

  • Calculate required sample size before starting
  • Be realistic about minimum detectable effect (can you act on a 0.5% improvement?)
  • If traffic is low, test bigger changes

3. Multiple Comparisons

The problem: Testing many variants or metrics and celebrating any that reach significance.

Why it's wrong: With 20 metrics, you expect 1 false positive at 95% confidence—by chance alone.

The fix:

  • Define ONE primary metric before starting
  • Use Bonferroni correction or similar for multiple comparisons
  • Treat secondary metrics as directional, not conclusive

4. Ignoring Segments

The problem: Only looking at aggregate results.

Why it's wrong: Simpson's Paradox—overall winner might be loser for your key segments.

The fix:

  • Always segment by device, traffic source, user type
  • Check if results are consistent across segments
  • If segments differ dramatically, investigate why

Design Mistakes

5. Testing Too Many Things

The problem: Changing headline, image, CTA, and layout simultaneously.

Why it's wrong: You won't know which change caused the result. And each variable multiplies required sample size.

The fix:

  • Test one variable at a time (A/B testing)
  • If testing multiple, use proper multivariate testing with adequate sample size
  • Prioritize highest-impact changes first

6. Vague Hypothesis

The problem: "Let's see if this new design is better."

Why it's wrong: Without a hypothesis, you can't learn WHY something worked (or didn't).

The fix:

  • State: "We believe [change] will [impact metric] because [reasoning]"
  • Even if you're wrong, you learn something

7. No Control

The problem: Changing the control during the test, or not having one.

Why it's wrong: You need a stable baseline to compare against.

The fix:

  • Never modify the control mid-test
  • If you must change it, start a new test
  • Document exactly what the control is

Execution Mistakes

8. External Contamination

The problem: Running a test during a sale, holiday, or major event.

Why it's wrong: External factors affect both variants differently, contaminating results.

The fix:

  • Avoid tests during unusual periods
  • If unavoidable, note it and extend the test past the event
  • Compare to the same period historically

9. Selection Bias

The problem: Testing on a non-representative sample (e.g., only logged-in users).

Why it's wrong: Results won't generalize to your full audience.

The fix:

  • Test on representative traffic
  • Be explicit about who's included/excluded
  • Note limitations when reporting results

10. Implementation Bugs

The problem: Variants don't render correctly, tracking fires incorrectly, assignment is biased.

Why it's wrong: You're not testing what you think you're testing.

The fix:

  • QA both variants thoroughly before launch
  • Verify tracking events fire correctly
  • Check assignment distribution matches weights

Interpretation Mistakes

11. Celebrating Trivial Wins

The problem: Implementing a change because it was "statistically significant" even though the effect was tiny.

Why it's wrong: Statistical significance ≠ practical significance. A 0.01% improvement isn't worth the complexity.

The fix:

  • Define minimum meaningful effect before starting
  • Consider implementation cost vs. benefit
  • Don't over-optimize

12. Ignoring Confidence Intervals

The problem: Only reporting point estimates ("5% improvement!").

Why it's wrong: The true effect could be anywhere in the confidence interval.

The fix:

  • Report confidence intervals: "5% improvement (95% CI: 2%-8%)"
  • Base decisions on the lower bound for conservative estimates
  • Wider intervals = more uncertainty

13. Not Documenting Learnings

The problem: Running tests but not recording what you learned.

Why it's wrong: You'll repeat mistakes, forget context, lose institutional knowledge.

The fix:

  • Document every test: hypothesis, results, learnings
  • Include what surprised you
  • Build a searchable knowledge base

Organizational Mistakes

14. HiPPO (Highest Paid Person's Opinion)

The problem: Running experiments but ignoring results when leadership disagrees.

Why it's wrong: Defeats the purpose of data-driven decision making.

The fix:

  • Get buy-in before testing that results will be honored
  • Present data clearly to stakeholders
  • Frame as "learning" not "winning/losing"

15. Testing Everything

The problem: Running experiments on trivial changes that don't matter.

Why it's wrong: Wastes resources, creates testing fatigue, delays important experiments.

The fix:

  • Prioritize tests by potential impact
  • Not everything needs a test—use judgment for low-risk changes
  • Focus experimentation resources on high-value decisions

16. Sample Ratio Mismatch (SRM)

The problem: The actual traffic split doesn't match the intended split (e.g., you expect 50/50 but observe 52/48).

Why it's wrong: SRM is a strong signal of an implementation bug — broken randomization, bot contamination, or redirect issues. Results from experiments with SRM cannot be trusted.

The fix:

  • Check the actual split ratio against expected before analyzing results
  • Use a chi-squared test to detect statistically significant mismatches
  • If SRM is detected, investigate the root cause before drawing any conclusions
  • Common causes: bot traffic, browser redirects dropping users, bucketing bugs

17. Novelty and Primacy Effects

The problem: Users react differently to new designs initially, and the effect fades over time.

Why it's wrong: Short experiments may show inflated effects that don't persist. Returning users may click more simply because something looks new.

The fix:

  • Run experiments for at least 2 full business cycles
  • Segment results by new vs. returning users
  • If possible, check whether the effect holds in the second week vs. the first

Source: SKILL.md on GitHub

No alerts16d5 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    The skill provides comprehensive guidelines and best practices for content experimentation and A/B testing. It includes illustrative code snippets for CMS integration and statistical calculation. No security risks were identified.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: LOW · No issues

  • Runlayer6mo

    5 files scanned · No issues

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at ad50ea4. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 weeks ago.

Activeupdated 6 months ago
  • a-b-testing
  • content-experimentation
  • statistical-analysis
  • cms
  • conversion-optimization
  • multivariate-testing
  • metrics
  • hypothesis-testing

README badge

README badge for sanity-io/agent-toolkit/content-experimentation-best-practices

Provides guidance on A/B testing, multivariate testing, and statistical analysis for content experiments, including experiment design, metrics selection, sample sizing, CMS integration patterns, and common pitfalls. Use when setting up experimentation infrastructure, designing content variants, or interpreting test results in a headless CMS or frontend stack.

Generated from the current SKILL.md.

Does this skill cover statistical rigor for A/B tests?
Yes. The skill includes statistical foundations covering p-values, confidence intervals, power analysis, and Bayesian methods to help interpret results correctly.
Can I use this skill to set up experiments in a headless CMS?
Yes. The skill includes guidance on CMS-managed variants and field-level variants, with patterns for integrating experimentation into CMS workflows.
What common mistakes does this skill help avoid?
The skill documents 17 common pitfalls across statistics, design, execution, and interpretation to help teams avoid typical experimentation errors.
Does this cover multivariate testing or just A/B tests?
Both. The skill covers A/B testing, multivariate testing, and how to design experiments that test multiple variables simultaneously.

Generated from the current SKILL.md. These answers refresh after source changes.