What I Learned from Khan Academy About A/B Testing AI
Blog post from GrowthBook
Khan Academy's journey to effectively measure and improve the quality of their AI tutor, Khanmigo, highlights the challenges of evaluating AI features in educational contexts. Initially relying on subjective assessments, the team transitioned to a rigorous A/B testing approach by developing a metric for cognitive engagement based on the ICAP framework. This involved creating a rubric, labeling student interactions, and training an LLM-as-judge to automate the evaluation process at scale. With this reliable metric, Khan Academy could conduct controlled experiments to enhance Khanmigo's tutoring capabilities, shifting their perception of experimentation from a hurdle to an essential safety net. This experience underscores the importance of establishing solid evaluation frameworks when dealing with complex AI outputs, enabling teams to make informed, data-driven improvements to their AI-driven products.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 6 | 5,932 | 1,046 | 223 | -2% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.