Home / Companies / AssemblyAI / Blog / Post Details
Content Deep Dive

Node.js voice agent with AssemblyAI Universal-3.5 Pro Realtime

Blog post from AssemblyAI

Post Details
Company
Date Published
Author
Kelsey Foster
Word Count
1,335
Company Posts That Month
19
Language
English
Hacker News Points
-
Post removed?
No
Summary

This tutorial provides a comprehensive guide on building a real-time voice agent in Node.js using the AssemblyAI Universal-3.5 Pro Realtime model for speech-to-text, without the need for Python or heavy framework dependencies. The setup includes two modes: a terminal agent that uses mic input and plays TTS audio in the terminal, and a browser server utilizing Node.js WebSocket with a user interface. The AssemblyAI model offers features like punctuation-based turn detection, context carryover, and mid-session keyterm prompting, enhancing the transcription process and eliminating the need for a separate VAD library. The tutorial emphasizes the model's efficiency in handling real-time conversations with low word error rates (WER) and provides detailed instructions on connecting to the AssemblyAI WebSocket, streaming audio, and fine-tuning turn detection. It also highlights the ability to update conversation context and keyterms mid-session without needing to restart the connection, thereby optimizing the real-time voice agent's functionality.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.