Two ways to measure the cumulative impact of experiments
Blog post from Datadog
Cumulative impact from experimentation programs cannot be accurately measured by simply adding observed lifts from statistically significant winning A/B tests, because the winner’s curse causes selected results to be inflated by sampling noise. A randomized holdout offers the most direct estimate by comparing users who retain the original product experience with users who receive all shipped winning variants, and it can capture long-term effects and interactions among changes, but it requires advance setup, sustained feature flags, traffic allocation, and time. Datadog’s Cumulative Impact feature provides a faster model-based alternative for historical or ongoing experiments by using empirical Bayes shrinkage to correct individual estimates and aggregate their likely true effects. This approach depends on experiments being comparable, their effect distribution remaining stable, and treatment interactions being limited, making it suitable when a holdout is impractical but its assumptions are reasonable. Teams may use model-based estimates for routine reporting while reserving holdouts for high-stakes decisions, long-term measurement, or validation.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.