Home / Companies / Daily / Blog / Post Details
Content Deep Dive

Announcing Pipecat PhoneLLM Alpha 1

Blog post from Daily

Post Details
Company
Date Published
Author
Marcus
Word Count
2,109
Company Posts That Month
1
Language
English
Hacker News Points
-
Post removed?
No
Summary

Pipecat has released PhoneLLM Alpha 1, a BSD-licensed open-weights voice-agent model fine-tuned from NVIDIA’s Nemotron 3 Nano 30B-A3B, designed to deliver accurate tool use and fast multi-turn conversations for customer-service and outbound calling applications. With 3.5 billion active mixture-of-experts parameters, the model is intended to offer lower latency and operating costs than larger general-purpose models while remaining deployable on private infrastructure. The company also introduced PhoneBench v1, which evaluates phone-agent models using calibrated LLM judges across factors including speaking style, tool-call accuracy, factual grounding, conversation coherence, escalation behavior, latency, and estimated per-minute cost. Pipecat argues that voice agents require roughly 1,500 ms voice-to-voice response times, making low time-to-first-token performance essential, and reports that PhoneLLM can achieve high concurrency on NVIDIA B200 hardware, including sub-100 ms single-request P95 time-to-first-token. The weights are available through Hugging Face and can be served with SGLang, vLLM, or Modal AutoEndpoints, while the team positions the release as part of a broader shift toward smaller, specialized open models that can be continually refined with targeted evaluations, production data, and inference optimizations.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.