Designing experiments that produce trustworthy results: a pre-launch guide to validity threats
Blog post from GrowthBook
Reliable A/B test results depend largely on decisions made before data collection, with experimental validity encompassing internal validity, external generalizability, appropriate metric construction, and sound statistical conclusions. The framework recommends pre-registering hypotheses, metrics, sample sizes, stopping rules, guardrails, and decision criteria; matching randomization and analysis units; re-randomizing users for new tests; ensuring treatment and control differ only in the intended change; and using cluster, geographic, or time-based designs when users influence one another. Teams should select one primary metric linked to the intended mechanism and longer-term organizational goals, calculate sample size through power analysis, set a sufficient runtime and conversion window, and account for concurrent experiments that may interact. Pre-launch checks such as A/A tests, balance assessments, instrumentation verification, overlap reviews, and runtime projections can identify design problems early, while monitoring for novelty effects, network interference, and treatment interactions remains important during and after testing. GrowthBook is presented as a platform that supports these practices through required checklists, power calculations, statistical health checks, sequential testing, and embedded decision frameworks.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.