Home / Companies / GrowthBook / Blog / April 2026

April 2026 Summaries

8 posts from GrowthBook

Filter
Month: Year:
Post Summaries Back to Blog
Chess.com, a platform with 10 million daily active users ranging from beginners to FIDE-rated competitors, faces unique challenges in its experimentation program due to the broad skill variance. The company conducted 400 experiments in 2024 and aims for 1,000 in 2025, focusing on designing user-specific experiments rather than simply increasing volume. Nafis Shaikh, Director of Product Management, highlights the importance of segmenting users by skill level to avoid misleading aggregate results and emphasizes a four-dimension framework for measuring metrics: inflows, engagement, retention, and monetization. Chess.com prioritizes understanding user behavior beyond whether a KPI moved, emphasizing narrative-driven insights to capture organizational learning. A notable experiment involved changing the Game Review feature's focus from mistakes to celebrating good moves, resulting in a 25% increase in feature engagement and improved subscription conversions. This approach underscores the importance of framing in user experience, suggesting that consumer-facing products might benefit from highlighting user successes.
Apr 27, 2026 1,450 words in the original blog post.
The evaluation of AI models through traditional benchmarks often fails to reflect their performance in real-world applications, as demonstrated by the disconnect between models like GPT-5's high scores on coding benchmarks and the actual preference for Anthropic's models by developers for practical use. Benchmarks, which are designed to measure model performance on standardized tasks, are criticized for not aligning with the specific needs and constraints of production environments, such as cost, latency, and task-specific performance. Instead, rigorous A/B testing with real users and workloads is advocated as a more reliable method for selecting and optimizing large language models (LLMs), as it allows businesses to assess metrics that truly drive value, such as task completion rates and cost efficiency. The "portfolio approach," which employs a variety of models tailored to different tasks, is highlighted as effective for optimizing performance and cost. Ultimately, the true measure of an AI model’s success is its ability to solve user problems within operational constraints, not just its scores on standardized benchmarks.
Apr 27, 2026 1,337 words in the original blog post.
In a webinar hosted by GrowthBook, Ronny Kohavi, a renowned figure in computer science and experimentation, alongside Luke Sonnet, delved into the principles of designing experiments for long-term growth. They emphasized the critical role of experimentation in shifting decision-making from intuition to evidence-based processes, highlighting the importance of setting clear success metrics and understanding the potential for high failure rates. Kohavi and Sonnet discussed the necessity of defining Overall Evaluation Criteria (OECs) to guide experimentation and avoid flawed decision-making, using real-world examples from companies like Airbnb and Bing to illustrate how poor metrics can lead to misguided strategies. They warned against "shipping flat," or implementing changes despite inconclusive experimental results, as this can lead to significant resource wastage. The speakers advocated for a culture shift from celebrating shipping to celebrating learning, positioning experimentation as a tool for smarter, faster growth through a structured framework that aligns short-term metrics with long-term goals. They also stressed the importance of defining shipping criteria before experiments to eliminate bias and ensure consistent decision-making.
Apr 27, 2026 2,620 words in the original blog post.
The Average Treatment Effect (ATE) is a commonly used metric in randomized experiments to measure the difference in outcomes between treatment and control groups, but it often oversimplifies complex individual responses by averaging them into a single figure. While ATE provides a quick overview, it can mask the diverse effects on different user segments, potentially leading to misguided decisions if not further analyzed. Different scenarios, such as universal slight benefits, effects driven by a specific subgroup, or a mixture of positive and negative impacts, can all yield the same average effect, but require different strategic responses. To gain a more nuanced understanding, it's essential to examine the distribution of treatment effects across user segments, which can unveil underlying dynamics and inform more targeted approaches. Experimentation platforms like GrowthBook offer tools to explore these variations, enabling more informed decisions and tailored user experiences by identifying which segments benefit from specific changes. Understanding the heterogeneity of treatment effects is crucial for maximizing the value of experimentation and ensuring that results are not just averaged but actionable insights.
Apr 27, 2026 1,525 words in the original blog post.
GrowthBook, an open-source, warehouse-native experimentation platform, is gaining traction among companies seeking an alternative to Statsig following its acquisition by OpenAI. Concerns about data governance and rising costs have prompted engineering teams to switch to GrowthBook, which offers predictable pricing, full data ownership, and transparent, reproducible results. Unlike Statsig, which routes event data through its own servers, GrowthBook's architecture utilizes existing data warehouses like Snowflake or BigQuery, ensuring that sensitive data remains under the user's control and complies with regulations such as GDPR and HIPAA. GrowthBook also supports a variety of statistical methodologies, including Bayesian and frequentist testing, and offers comprehensive tools for experimentation and feature flagging without incurring per-evaluation costs. The platform provides a seamless migration process, including an AI-powered assistant to help transfer existing setups from Statsig, and allows users to run experiments and manage features with extensive visibility and control, making it a compelling choice for teams prioritizing data sovereignty and cost-effective scaling.
Apr 27, 2026 1,827 words in the original blog post.
Security certifications alone are insufficient to prevent data breaches, as demonstrated by the breach at Evolve Bank & Trust, despite their compliance with SOC 2 Type II, HIPAA, HITRUST CSF, and PCI DSS standards. The incident underscores the importance of reducing the attack surface, which refers to the number of potential entry points for cyber attackers. In highly regulated sectors like fintech, minimizing this surface is crucial, as evidenced by the 566 data breaches in the financial sector in 2022. Self-hosting tools, which keep infrastructure within a private network, offer a more secure alternative by limiting exposure to the internet, without compromising on innovation or speed. GrowthBook exemplifies this approach, providing self-hosted solutions that ensure enterprise-grade security while maintaining flexibility and control.
Apr 27, 2026 454 words in the original blog post.
Khan Academy's journey to effectively measure and improve the quality of their AI tutor, Khanmigo, highlights the challenges of evaluating AI features in educational contexts. Initially relying on subjective assessments, the team transitioned to a rigorous A/B testing approach by developing a metric for cognitive engagement based on the ICAP framework. This involved creating a rubric, labeling student interactions, and training an LLM-as-judge to automate the evaluation process at scale. With this reliable metric, Khan Academy could conduct controlled experiments to enhance Khanmigo's tutoring capabilities, shifting their perception of experimentation from a hurdle to an essential safety net. This experience underscores the importance of establishing solid evaluation frameworks when dealing with complex AI outputs, enabling teams to make informed, data-driven improvements to their AI-driven products.
Apr 21, 2026 1,252 words in the original blog post.
Fyxer, an AI email assistant company, achieved remarkable growth from $1M to $35M in annual recurring revenue by adopting a culture of rapid experimentation, running 541 experiments in a year, primarily driven by a small growth engineering team led by Kameron Tanseli. This approach involved treating every product change as a hypothesis to validate and leveraging AI tools like Cursor, Claude, and GrowthBook to expedite the experimentation process, enabling the team to operate at a scale previously unimaginable. The company's strategy emphasized the importance of A/B testing not just as an optimization tool but as a means of learning and adapting to new market dynamics, resulting in a 25% experiment success rate that enabled significant improvements in user conversion and retention. By integrating AI to streamline research, development, and data analysis, Fyxer could efficiently test and iterate on product features, ultimately leading to counterintuitive successes such as increasing free-to-paid conversion rates and enhancing customer retention through innovative pricing strategies. As Fyxer aims for further growth, the focus remains on expanding the growth engineering team and deepening investment in AI-driven processes to amplify the impact of their high-velocity experimentation strategy.
Apr 04, 2026 1,858 words in the original blog post.