Home / Companies / GrowthBook / Blog / Post Details
Content Deep Dive

What I Learned from Khan Academy About A/B Testing AI

Blog post from GrowthBook

Post Details
Company
Date Published
Author
Ashley Stirrup
Word Count
1,252
Company Posts That Month
8
Language
English
Hacker News Points
-
Post removed?
No
Summary

Khan Academy's journey to effectively measure and improve the quality of their AI tutor, Khanmigo, highlights the challenges of evaluating AI features in educational contexts. Initially relying on subjective assessments, the team transitioned to a rigorous A/B testing approach by developing a metric for cognitive engagement based on the ICAP framework. This involved creating a rubric, labeling student interactions, and training an LLM-as-judge to automate the evaluation process at scale. With this reliable metric, Khan Academy could conduct controlled experiments to enhance Khanmigo's tutoring capabilities, shifting their perception of experimentation from a hurdle to an essential safety net. This experience underscores the importance of establishing solid evaluation frameworks when dealing with complex AI outputs, enabling teams to make informed, data-driven improvements to their AI-driven products.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 6 5,932 1,046 223 -2%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.