July 2026 Summaries
25 posts from GrowthBook
Filter
Month:
Year:
Post Summaries
Back to Blog
The recent episode of The Experimentation Edge podcast features Fabian Hans, founder of Cogniteer, who discusses the challenges and nuances of running effective experimentation programs. Hans, drawing on his background in behavioral psychology, emphasizes that while increasing the volume of experiments is easy, gaining meaningful insights from them is much more challenging. He critiques the common practice of mass-produced tests, highlighting that they can lead to a superficial win rate without deeper understanding. Hans's experience shows that most e-commerce issues, like high drop-off rates, are structural rather than company-specific, often due to a mismatch between the product and the online shopping channel. He advocates for tailoring the user interface to match different buyer needs, such as providing more context to new users while maintaining a streamlined experience for returning customers. Ultimately, Hans argues that the key to a successful experimentation program is not just about having the right tools but fostering a deep understanding of user behavior to design experiments that yield genuine insights, thereby turning experimentation into a compounding growth mechanism.
Jul 31, 2026
1,281 words in the original blog post.
In experimentation, achieving significant results on a mature product is challenging due to the small effects most changes have on target metrics, as shown by 1,450 experiments on Microsoft's Bing. Such small effects are difficult for standard A/B tests to detect, often leaving results undecided by evidence. To enhance experiment sensitivity and detect these small effects without waiting for larger samples, practitioners can use variance reduction techniques, which include metric choice, winsorization, CUPED, post-stratification, and triggered analysis. Each technique addresses different aspects of variance, such as choosing metrics closer to the treatment for clarity, capping outliers to prevent dominant effects, using pre-experiment data to adjust outcomes, stratifying users to manage group differences, and filtering analyses to focus on exposed users. GrowthBook implements these techniques to streamline the experimentation process and reduce variance, empowering businesses to make quicker, more informed decisions.
Jul 30, 2026
2,613 words in the original blog post.
In January 2026, Monzo experienced a two-hour outage due to its mobile banking app malfunctioning, which highlighted the importance of modernizing fintech infrastructure to protect customer trust and comply with regulations. Fintech companies, responsible for handling sensitive financial data, face significant risks with each software deployment, and feature flags have become a critical tool to mitigate these risks. Feature flags allow for more controlled, progressive rollouts and automated rollbacks, ensuring that deployment can be decoupled from release and handled within the constraints of compliance requirements. They offer a way to manage deployments safely by enabling the monitoring of key metrics and ensuring that financial data remains secure. GrowthBook is one example of a platform that provides these capabilities, allowing fintech companies like Upstart to streamline their processes, reduce technical debt, and enhance their experimentation efforts while maintaining compliance. Feature flags help fintech organizations maintain operational resilience, manage jurisdiction-based feature releases, and implement AI models for tasks like fraud detection, all while providing a safeguard against potential disruptions.
Jul 28, 2026
2,070 words in the original blog post.
Fin, an AI support agent developed by Intercom, has improved its capabilities through a rigorous experimentation framework led by Principal Machine Learning Scientist Pedro Tabacof. The AI agent, which now generates over $100 million in annual recurring revenue and serves more than 10,000 customers, emphasizes the importance of large-scale A/B testing to refine its functionalities, given the non-deterministic nature of AI systems where traditional unit tests fall short. Through continuous testing, even for minor changes, Fin has discovered counterintuitive insights, such as the benefits of increased latency, which can enhance user perception by making interactions appear more thoughtful and human-like. A notable experiment involved increasing the conversation history context to improve answer quality, initially leading to unintended hallucinations like false refund promises. By revisiting and refining prompts, the team managed to retain the positive improvements without the downsides, demonstrating that failed experiments often have a path to success through targeted adjustments. Fin's culture of experimentation, viewed as essential for quality rather than a hindrance to speed, allows for innovation and learning from failures, underpinned by leadership that values data-driven decision-making.
Jul 28, 2026
1,357 words in the original blog post.
On July 19, 2024, a logic error in a configuration update deployed by CrowdStrike to all its Falcon sensors simultaneously led to a global outage affecting 8.5 million Windows systems and costing Fortune 500 companies an estimated $5.4 billion. The incident highlighted the risks associated with deploying software updates without staged rollouts and underscored the importance of adopting canary releases, which expose changes to a small user group before a full rollout, thus limiting potential damage. The article explains the concept of canary releases and their implementation, using feature flags to manage deployment risks effectively by allowing rapid rollback and minimizing disruption. It discusses the advantages and challenges of canary releases, such as operational complexity and the need for robust metrics and monitoring, contrasting infrastructure and feature flag canaries for different deployment needs. Additionally, it outlines strategies for effective canary releases, emphasizing the importance of metrics monitoring, avoiding common pitfalls, and advocating for a phased release approach to prevent large-scale outages and improve deployment confidence.
Jul 27, 2026
2,654 words in the original blog post.
GrowthBook 5.0 is designed to streamline and enhance the process of running experiments for teams at various stages of experimentation, addressing common friction points such as complex setup and bottlenecked resources. By reducing the number of fields required to start an experiment, the platform makes it easier for more team members to initiate experiments without needing specialized expertise, fostering a culture of collaboration. The introduction of AI tools like the AI Assistant and AI Visual Editor further democratizes experiment creation, allowing broader participation while maintaining standards through mandatory custom fields and linked feature flags. Advanced analytical techniques and infrastructure improvements, including quantile treatment effects and CUPED, reduce compute costs and improve efficiency for high-volume data environments. This comprehensive approach ensures that as teams grow and scale their experimentation efforts, they can do so with maintained safety, reduced costs, and enhanced learning opportunities, ultimately aiding in better product development.
Jul 24, 2026
1,144 words in the original blog post.
GrowthBook's 5.0 release emphasizes enhanced governance for feature flags, aiming to facilitate faster yet safe code deployment. As AI accelerates coding and organizations release more changes, ensuring each modification is valid and contextually appropriate becomes crucial. The update introduces key governance features to catch issues during the change-making process, manage rule interactions, and facilitate final decision-making with comprehensive context. These include schema validation for string and number flags, feature-scoped Custom Hooks, soft warnings, and sparse patches for JSON rules, all designed to identify potential errors early and maintain focus on essential changes. Additionally, GrowthBook provides tools to detect rule conflicts, ensuring that rules are evaluated effectively without unintentional traffic overlap. The new Review & Publish tab consolidates the review process, allowing teams to manage changes with greater clarity and confidence, thus enhancing the quality of published work without hindering productivity.
Jul 23, 2026
1,265 words in the original blog post.
GrowthBook's Product Analytics, now generally available, expands upon its existing experimentation framework by utilizing the same trusted metric definitions to monitor KPIs, explore trends, and analyze funnels. It provides tools such as dashboard creation and AI-powered analytics, allowing users to generate actionable insights and visualizations without needing deep technical expertise. The platform's Metric Explorer enables visualization of predefined metrics or raw data, supporting various chart types and integration with major data warehouses like BigQuery and Snowflake. GrowthBook also offers a SQL Explorer for custom queries, which can be saved and integrated into dashboards alongside metric explorations. Funnel analysis features help identify conversion rate drop-offs, thus facilitating hypothesis-driven experimentation. Dashboards are customizable, supporting a range of data blocks to align with team-specific narratives and share insights. The Product Analytics platform builds on existing data infrastructure, allowing seamless integration with GrowthBook's experiment data while offering programmatic access through a REST API.
Jul 22, 2026
876 words in the original blog post.
GrowthBook 5.0 introduces an AI Visual Editor that empowers growth and marketing teams to create and launch website experiments from plain-language prompts without requiring engineering resources. This tool addresses the common issue of limited testing capacity due to engineering bandwidth, enabling teams to implement and test ideas like new headlines, landing pages, and audience-specific messaging quickly and efficiently. The AI Visual Editor allows users to generate images, update designs, and run tests such as multi-arm bandits, all within a WYSIWYG interface, thereby reducing the friction traditionally associated with experimentation. By lowering the cost of implementation, teams are encouraged to test more ideas, validate concepts before committing to significant development, and personalize experiences for different audiences, ultimately fostering a culture of rapid learning and creativity. The editor, available as a Chrome extension, maintains transparency, allowing technical teams to audit experiments and ensuring that engineering involvement is minimized without sacrificing rigor or reliability.
Jul 21, 2026
1,278 words in the original blog post.
GrowthBook 5.0 introduces significant enhancements to its platform, enabling companies to improve their products efficiently and at scale. This major update builds on the work done since the last release, GrowthBook 4.0, and includes a complete rewrite of feature flags, the introduction of Product Analytics, faster experimentation, and AI integration throughout the platform. The update has expanded experimentation capabilities to the entire team, allowing more team members to ship flags, run tests, and make data-driven decisions across various interfaces such as apps, terminals, code editors, and browsers. Throughout the week, GrowthBook is highlighting new features daily, including Agents with a CLI for direct platform operations, an AI Visual Editor for browser-based experimentation, general availability of warehouse-native Product Analytics, customizable governance for feature flags, and enhancements for faster, scalable experiments. GrowthBook 5.0 is available on Cloud and for self-hosted deployments, with further details provided in the release notes and a planned live Office Hours session.
Jul 20, 2026
349 words in the original blog post.
Fyxer, a company that conducted 541 experiments last year, has developed an innovative workflow integrating agents to streamline and manage their experimental processes using GrowthBook. This approach involves initiating experiments through a form submission, which triggers an agent to create and manage flags and experiments, contributing to a significant ARR growth from $1M to $35M despite only 25% of experiments being successful. GrowthBook 5.0 enhances this agent-driven experience, allowing users to create, launch, and manage experiments with ease across different platforms, supported by 25 open-source skills and a comprehensive REST API. The system ensures that agent-driven changes are reviewed by a team member before going live, maintaining governance and quality control. GrowthBook's consistent platform experience across various tools, supported by AI-driven insights and dashboards, promotes seamless experimentation and decision-making, with plans for further enhancements like an AI Visual Editor to simplify browser-based experimentation.
Jul 20, 2026
636 words in the original blog post.
GrowthBook's AI Visual Editor is a Chrome extension designed to facilitate rapid experimentation across organizations by enabling users to create and launch A/B tests without writing code. By allowing team members to describe desired changes in plain language, the editor generates and applies these modifications directly to the webpage, integrating them into GrowthBook experiments. It supports the adjustment of text, images, fonts, colors, and layouts, and includes a manual editor for fine-tuning. Additionally, the tool allows for importing designs from Figma and testing various elements like landing pages and promotional banners, promoting collaboration among marketing, growth, and product teams. The AI Visual Editor ensures existing workflows remain intact, maintaining governance while accelerating the testing process, and is currently available on GrowthBook Pro and Enterprise plans in beta form.
Jul 17, 2026
896 words in the original blog post.
GrowthBook's webinar, featuring experts Ronny Kohavi and Luke Sonnet, focused on the challenges of ensuring A/B test results are reliable, emphasizing that while generating data is straightforward, obtaining trustworthy data is complex. Drawing insights from Kohavi's extensive experience at companies like Amazon and Microsoft, and Sonnet's leadership at GrowthBook, the webinar highlighted common pitfalls in A/B testing such as underpowered tests, misinterpreted p-values, and sample ratio mismatches. The discussion stressed the importance of rigorous setup practices like power calculations, setting realistic Minimum Detectable Effects (MDEs), and choosing appropriate significance thresholds based on the cost of errors. Post-experiment checks such as verifying sample ratios and computing false positive risks were recommended to enhance result reliability. The session underscored that A/B tests, while a gold standard in experimentation, require careful planning and analysis to avoid misleading conclusions, and GrowthBook offers tools to facilitate these best practices.
Jul 17, 2026
2,618 words in the original blog post.
LaunchDarkly is a proprietary enterprise feature management platform that has expanded its offerings to include observability and analytics, catering to companies adopting AI-native development processes. Despite its strengths in release governance and feature management, users have raised concerns about its unpredictable usage-based pricing, the separate and costly experimentation module, lack of self-hosting options, reliability issues, and vendor lock-in due to its closed-source nature. The platform's integration post-2025 acquisitions, such as Highlight and Houseware, has broadened its capabilities, yet also increased the overall cost and complexity. LaunchDarkly's SaaS-only deployment model and limited support for non-Snowflake warehouse-native capabilities add to the challenges faced by teams with specific compliance and data-residency requirements. Users seeking alternatives often consider platforms that offer open-source options, self-hosting, integrated experimentation and analytics, and more predictable pricing structures.
Jul 17, 2026
8,505 words in the original blog post.
Experimentation at Kargo is not just encouraged but deeply embedded in its culture, as emphasized by James Falzone, Director of Product Management. Kargo's approach to advertising technology, which involves real-time auctions for ad placements, necessitates constant experimentation and learning from failures to stay competitive. Falzone argues that the true value of experimentation lies in distinguishing between a "bad result" and a "bad experiment," with the former offering insights into incorrect assumptions and the latter being hindered by excessive caution. A failed experiment involving Kargo's click optimization model highlighted the importance of context-specific implementation rather than a one-size-fits-all approach. Kargo fosters an environment where discussing failures is normalized, allowing teams to learn collectively and innovate freely. While embracing AI for its potential to democratize experimentation, Falzone maintains that these technologies should complement robust machine learning and infrastructure engineering. Ultimately, Kargo's philosophy prioritizes improving processes over merely expanding resources, ensuring that experimentation remains a strategic advantage in the fast-paced ad tech industry.
Jul 16, 2026
1,113 words in the original blog post.
GrowthBook 4.4 introduces a comprehensive framework that integrates AI coding agents, such as Claude Code, Cursor, and Codex, to automate the entire product development lifecycle, including ideation, feature flag creation, building, and testing within a single platform. By leveraging AI, the process of coding, which traditionally took significant time and resources, is now expedited, allowing for rapid experimentation and iteration. The platform addresses the need for consistency and rigor by providing experiment templates and a decision framework that ensure the trustworthiness of results, thereby preventing errors like flawed metrics and unexpected side effects. Despite automation, critical decisions such as quality assurance and the final rollout remain under human control, ensuring that AI serves as a tool for efficiency rather than replacing human judgment. GrowthBook's open-source skills, available on GitHub, allow teams to customize their experimentation framework, making the process adaptable and tool-agnostic across different coding agents.
Jul 15, 2026
1,119 words in the original blog post.
Dan Layfield, Director of Product Management at Diligent, shares insights from his extensive experience in product management, emphasizing the complexities of interpreting experimental data and knowing when to persist with inconclusive results. Throughout his career, including roles at Codecademy and Uber Eats, Layfield has highlighted the importance of experimentation and the recent role of AI in expediting data synthesis. He distinguishes between high-volume conversion optimization and more uncertain, high-stakes experiments, illustrating the latter with a case at Codecademy where persistent adjustments led to a significant increase in conversion rates. Layfield also critiques the "feature factory" approach in B2B settings, advocating for a disciplined, goal-oriented product management strategy with a clear alignment between business objectives and team metrics. At Diligent, he stresses the importance of aligning engagement strategies with the natural use cases of their products. Layfield acknowledges AI's ability to streamline data analysis, allowing teams to focus on meaningful insights and reducing the temptation to prematurely abandon promising projects.
Jul 14, 2026
1,239 words in the original blog post.
Danielle Oleen, Director of E-commerce at Box, emphasizes that her role is defined by revenue generation rather than adherence to a feature roadmap, highlighting the importance of experimentation in product development. With over 15 years of experience in e-commerce, Oleen has consistently utilized A/B testing as a crucial tool, proving that experimentation is essential for all product teams, not just those managing checkout flows. Her experience at Box involves transforming the company from a cloud storage provider to an AI platform, where experimentation, such as simplifying and testing the pricing page, has led to surprising insights like the "wine effect," where unexpected outcomes provide valuable learning opportunities. Oleen stresses the necessity of building a culture that embraces both wins and losses, fostering psychological safety and encouraging teams to test and iterate continuously. Her approach demonstrates that the most successful teams are those willing to explore and learn from unexpected results, rather than relying solely on initial hypotheses.
Jul 13, 2026
1,245 words in the original blog post.
A/B testing often results in higher than expected false positive rates, even at companies with advanced experimentation programs like Microsoft and Airbnb, where rates can range from 6% to 26%. This occurs when tests are run too quickly or improperly, leading to incorrect conclusions about the effectiveness of changes. The primary causes of false positives include premature peeking at results, testing multiple variables simultaneously, sample ratio mismatches, and underpowered tests. These issues can lead to significant inefficiencies and resource wastage, as teams may mistakenly interpret random variations as meaningful improvements. To mitigate such errors, best practices include pre-committing to sample sizes and test durations, using A/A tests for calibration, applying sequential testing to manage peeking, checking for sample ratio mismatches, and employing multiple testing corrections like Holm-Bonferroni and Benjamini-Hochberg. Additionally, validating results through causal chains and applying variance reduction techniques such as CUPED can further reduce false positives. These strategies are part of the comprehensive approach offered by GrowthBook, which integrates various features to ensure more reliable and accurate experimentation outcomes.
Jul 10, 2026
2,309 words in the original blog post.
Kevin Yang's exploration of experimentation at JPMorgan Chase reveals a perspective that values the insights gained from failed experiments as much as, if not more than, the successes. While his team estimates that successful experiments have generated over a billion dollars, Yang emphasizes that the true value lies in the lessons learned from failures, which prevent potentially harmful changes from being implemented at scale. At Chase, where his team supports about 100 product teams running 300 experiments annually, the focus is on building a culture that appreciates the importance of control groups and pre-planned responses to failure, which help avoid confirmation bias and ensure informed decision-making. This approach is particularly crucial in the AI era, where rapid iteration and customization demand rigorous measurement to avoid compounding mistakes. Yang's insights underscore the importance of a balanced decision framework that values trust and customer satisfaction over raw engagement metrics, highlighting the necessity of planning for failure to foster innovation.
Jul 08, 2026
1,338 words in the original blog post.
GrowthBook's Visual Editor is designed to address common shortcomings of traditional visual editors by offering an AI-first approach that allows users to make changes to a website for A/B testing without coding or relying on engineering teams. Unlike other editors that often require reverting to CSS or HTML for complex tasks, GrowthBook enables users to describe changes in natural language, which the AI then applies instantly on the page. It also incorporates image editing and generation, allowing users to create or modify images directly in the editor. The platform supports imports from Figma or static mockups, facilitating the seamless integration of designs into experiments. GrowthBook's editor is engineered to work with modern web platforms, using durable selectors to avoid issues with dynamically generated class names, ensuring variations persist through site updates. It includes features like a manual mode for precise control, global code editors for advanced customizations, and mechanisms to eliminate flicker, ensuring experiments run smoothly without bias. The interface supports multiple languages and offers a transparent change-tracking system, enhancing usability and control for diverse teams.
Jul 08, 2026
933 words in the original blog post.
The text discusses the concept of the "winner's curse" in the context of experimental results, highlighting how summing significant results from multiple experiments can lead to an overestimation of their true impact. This bias arises because experiments that pass the significance threshold often do so due to favorable random noise, causing inflated estimates. The text uses examples, such as Airbnb's findings, to illustrate how much the reported impact can differ from reality. To address this, it suggests using a "holdout" group—a subset of users not exposed to new changes—to get a more accurate measure of impact, as well as considering replication and Bayesian adjustments to correct for bias. The piece emphasizes the importance of understanding the mechanics of hypothesis testing and the limitations of aggregating results based solely on significance, warning against over-reliance on these inflated figures for decision-making.
Jul 08, 2026
2,172 words in the original blog post.
Holdout testing is an experimental method used to measure the overall impact of multiple features released over a period by comparing a group of users who did not receive the updates against those who did. This method addresses the inflation of individual A/B test results due to the winner's curse and potential feature interactions that might cancel each other out. The text examines four configurations of holdout testing, each providing a slightly different perspective on the cumulative effect of recent product changes: full-terminal, full-incremental, split-incremental, and split-terminal, with a fifth, reverse holdout, as an after-the-fact option. These configurations differ based on whether they focus on the long-term settled value of features or the real-time experience of users, how they handle novelty effects, and their approach to accounting for interactions between features. The right configuration depends on the desired outcomes, such as understanding the net effect of all features or identifying which features or teams drive more impact. The choice between configurations considers factors like interaction weight, novelty effects, and the power cost of maintaining separate test groups.
Jul 08, 2026
3,996 words in the original blog post.
As software teams increasingly adopt AI-assisted code and face higher deployment frequencies, the complexity and risk of deployments have risen, prompting a shift towards using feature flags to mitigate these risks. Feature flags allow teams to decouple deployment from release, providing control over who sees what features and enabling quick rollbacks if issues arise, thereby reducing the risk associated with all-or-nothing deployments. Progressive rollouts and attribute-based targeting further minimize exposure to potential problems by gradually introducing features and selectively targeting users based on specific attributes. Tools like GrowthBook facilitate these processes by offering platforms that support feature flags, kill switches, and dark launches, allowing for more controlled and flexible deployment strategies. This approach not only helps manage the surface area for mistakes but also improves recovery times by enabling quicker responses to deployment failures, as evidenced by various case studies, including AWS and CrowdStrike. The emphasis on feature flags as a derisking mechanism reflects a broader industry trend towards more robust and resilient deployment pipelines in the face of increasing code volumes and complexity.
Jul 06, 2026
1,938 words in the original blog post.
In a discussion on The Experimentation Edge podcast, Arun Bodapati, director of data science at Twitch, emphasizes the critical importance of preventing false negatives in experimentation, which can lead to potentially valuable ideas being disregarded and shelved for years. Unlike false positives, which typically get scrutinized and corrected, false negatives often go unnoticed and can have a lasting detrimental impact by institutionalizing incorrect conclusions. Bodapati advocates for rigorous preparation before running experiments, including clearly defining hypotheses in plain English, ensuring reliable enrollment triggers, and using broad "explore" experiments to avoid over-narrowing that could limit statistical power. He highlights the necessity of understanding the mechanisms behind positive results and warns against relying solely on numerical outcomes without a clear explanation of user behavior. A case study at Twitch regarding subscription pricing demonstrates how methodical experimentation and causal inference have shifted pricing strategy from being an untouchable area to a continuously adjustable lever, informed by reliable data and analysis.
Jul 01, 2026
1,280 words in the original blog post.