Home / Companies / Hume / Blog / May 2026

May 2026 Summaries

2 posts from Hume

Filter
Month: Year:
Post Summaries Back to Blog
Humans possess the ability to separate emotion from vocal delivery, an area where voice models still struggle due to the entanglement problem, where emotions and delivery are learned as fixed pairings from natural training data. Common emotional cues like anger with shouting or boredom with monotone delivery dominate training datasets, causing models to sound narrow or unrealistic when tasked with more subtle combinations like angry whispers or energetic boredom. Recent architectural advancements, such as factorized codecs and self-distillation, aim to address this issue but remain limited by the data distribution. To provide a more robust solution, cross-product sampling is proposed, encouraging models to learn emotions and delivery as independent dimensions by sampling from a grid of emotion and voice categories. This method aims for better coverage of rare combinations, such as confident disappointment or articulate pain, by collapsing attributes into parent categories and applying z-normalization to balance attribute prominence. Cross-product sampling has shown promising results in maximizing expressivity and reducing mutual information, thus offering a structured way to expand expressive capabilities in voice models, and highlights the importance of diverse training data for achieving independent control over emotion and delivery.
May 27, 2026 1,783 words in the original blog post.
Research over the past decade has challenged the traditional six-category model of emotions, revealing instead a high-dimensional space of emotional expression, particularly evident in vocal cues. This new understanding underscores the importance of interpreting emotions as continuous, blending states rather than discrete categories, which has significant implications for emotion AI systems. Voice, as a separate channel from spoken words, carries rich emotional information through prosody, which includes pitch, intonation, pauses, and vocal bursts. Effective emotion AI must move beyond transcript analysis to capture these nuances, requiring audio-based measurements anchored in human judgment. A scientifically grounded approach to mapping emotional expressions enables more accurate and meaningful AI interpretations, which are essential for developing systems that align with human perceptions of vocal emotion.
May 04, 2026 816 words in the original blog post.