October 2026 Summaries
2 posts from Hume
Filter
Month:
Year:
Post Summaries
Back to Blog
Hume AI describes the science, training data, and evaluation methods behind its Expression API, which provides real-time measurements of vocal and facial expression for applications such as data annotation, AI evaluation, and interaction analysis. Its speech emotion model assigns scores across 414 emotional tags, while its voice descriptor model measures 190 vocal characteristics, and its facial model recognizes 48 emotional categories plus 27 visible facial features. The models were trained using human-rated resources including more than 370,000 facial images with over one million ratings and more than 260,000 speech recordings. In evaluations against human judgments and public benchmarks, Hume reports that its speech emotion model achieved the highest overall emotion-ranking AUC of 77.1 among 11 compared systems across over 9,000 recordings from 32 datasets, ahead of Gemini 3.8 Flash at 73.0, though its performance was similar to Gemini on an internal evaluation of finer-grained emotional expressions across 16 languages.
Oct 06, 2026
1,263 words in the original blog post.
Hume AI has introduced Expression APIs that convert vocal and facial cues into structured measurements intended to help developers and researchers capture context beyond spoken words. The Audio Expression API assesses 414 expression tags across 23 emotion categories and 391 detailed descriptions, plus 190 voice characteristics such as breathiness, speed, monotony, and raspiness, and supports more than 50 languages. The Video Expression API measures 48 emotional-expression categories and 27 visible facial features, drawing on human-rated data from participants across several countries. Designed for both real-time use and analysis of recorded material, the APIs provide timestamped scores that can support customer-service agents, market research, robots, data retrieval, evaluation, and model training. Hume states that its models were trained on more than 260,000 speech recordings and 370,000 human-rated facial-expression images and match or exceed other tested systems against human judgments.
Oct 06, 2026
986 words in the original blog post.