Home / Companies / GrowthBook / Blog / October 2026

October 2026 Summaries

5 posts from GrowthBook

Filter
Month: Year:
Post Summaries Back to Blog
Michal Lenik describes Grubhub’s experimentation approach as an upstream, design-led process focused less on traditional A/B testing and more on testing small cohorts, prototypes, and concepts with real users before broad releases. Her teams position design as a business function that helps define solutions early, using AI to accelerate prototyping and research synthesis while allowing designers to focus on design engineering, craft, and product strategy. Examples include a post-order review screen that was gradually tested to reduce customer-care contacts without hurting orders, and merchant-dashboard concepts that shifted direction after merchants preferred a simpler experience over more metrics. Grubhub defines success metrics across design, product, engineering, and business outcomes, then combines quantitative results with follow-up qualitative interviews to understand user behavior and improve underperforming features. Lenik argues that organizations should intentionally decide whether to de-risk products through rapid production iteration or more extensive upstream validation, based largely on how quickly their engineering teams can respond to feedback.
Oct 07, 2026 1,297 words in the original blog post.
GrowthBook rebuilt its MCP server from a large collection of product-specific tools into a thin adapter that routes agents to a typed REST API for product capabilities and reusable skills for workflow guidance, reducing maintenance duplication across its API, CLI, MCP clients, and in-app assistant. The API supports operations involving feature flags, experiments, metrics, and configuration, while skills instruct agents on how to investigate context, follow experimentation procedures, identify assumptions, and return evidence for human review; permissions, validation, approval policies, review states, and audit logs remain the product-layer controls rather than security functions of the skills themselves. GrowthBook also uses experiment history and curated Learnings as context, enabling agents to generate and develop test ideas grounded in prior results, while requiring people to select proposals and approve consequential actions such as customer exposure changes. The company argues that autonomy should be determined by risk, reversibility, and organizational controls rather than simply whether an operation is a read or write, and that trust should be built progressively through inspectable outputs, limited credentials, sandbox work, and reviewable changes. By treating MCP as one interface among many rather than the core product, GrowthBook aims to support engineers, product managers, analysts, and growth teams across editors, terminals, chat clients, and its application while preserving a shared experimentation model and governance process.
Oct 07, 2026 3,019 words in the original blog post.
MilliporeSigma’s A/B testing program, coordinated by Senior Analyst Dorothy Crepin, supports a complex life sciences e-commerce site serving scientists, procurement buyers, and university users across varying regulatory markets. A notable experiment replaced product-specific type-ahead search suggestions with broader search-term suggestions, producing double-digit increases in search feature use and add-to-cart activity by allowing customers to explore rather than steering them toward a single assumed product. Crepin emphasizes measuring experiments against page-specific leading metrics, such as product discovery and add-to-cart behavior, rather than treating revenue as the sole indicator of success, while carefully sequencing overlapping tests to preserve valid results. The company combines A/B testing, which identifies what users do, with domain-expert user testing to better understand why they behave that way, and encourages teams to begin with validated customer problems rather than assumptions about preferred features. Looking ahead, the program is exploring agent-based personas to simulate browsing behavior, more varied statistical evaluation methods, and anomaly detection beyond standard primary and secondary metrics as it targets roughly 12 experiments per quarter.
Oct 06, 2026 890 words in the original blog post.
GrowthBook has introduced beta support for using LLM trace data as metrics in controlled online experiments, helping teams evaluate model, prompt, agent, and parameter changes with real users before broad releases. By connecting trace information such as token costs, latency, errors, tool calls, and evaluation scores with product metrics including revenue, feature adoption, activation, retention, and user feedback, teams can assess both AI quality and business impact. The platform assigns users to experiment variations, tags their traces with the assigned version, retrieves those traces from integrations such as Langfuse and Arize Phoenix, and compares results using frequentist or Bayesian methods. GrowthBook provides default metrics including cost per user, p95 latency, error rate, token usage, calls per user, and traces per user, while allowing custom evaluation and feedback metrics. Suggested applications include testing cheaper models without sacrificing quality, improving assistant helpfulness, dynamically routing traffic among prompts, measuring whether larger models increase sales, and gradually rolling out agents with performance guardrails.
Oct 05, 2026 1,155 words in the original blog post.
Running experiments sequentially limits annual test volume, while large technology companies achieve thousands of tests by running independently randomized experiments concurrently, allowing users to participate in multiple tests without generally biasing results. Deterministic per-experiment hashing and sticky assignment help ensure that treatment and control groups are balanced across other experiments, although conflicts can arise when changes affect the same product element, parameter, funnel stage, pricing offer, or technical implementation. Known conflicts can be managed through mutual exclusion groups or namespaces, but exclusion reduces available sample size and can slow testing. Unanticipated problems are commonly identified through health checks such as sample ratio mismatch, multiple-exposure warnings, and pairwise interaction analyses. As experimentation expands, programs must also address false positives, multiple comparisons, repeated result checking, and validation of downstream metrics through statistical corrections and sequential testing. The account argues that operational capacity—including automated metric calculation, standardized setup and review processes, searchable experiment records, and monitoring of portfolio performance—typically constrains scale more than experiment interactions, citing DoorDash, Microsoft, Google, Booking.com, Chess.com, and GrowthBook as examples of platforms and organizations supporting high-volume experimentation.
Oct 01, 2026 2,768 words in the original blog post.