Announcing Pipecat PhoneLLM Alpha 1
Blog post from Daily
Pipecat has released PhoneLLM Alpha 1, a BSD-licensed open-weights voice-agent model fine-tuned from NVIDIA’s Nemotron 3 Nano 30B-A3B, designed to deliver accurate tool use and fast multi-turn conversations for customer-service and outbound calling applications. With 3.5 billion active mixture-of-experts parameters, the model is intended to offer lower latency and operating costs than larger general-purpose models while remaining deployable on private infrastructure. The company also introduced PhoneBench v1, which evaluates phone-agent models using calibrated LLM judges across factors including speaking style, tool-call accuracy, factual grounding, conversation coherence, escalation behavior, latency, and estimated per-minute cost. Pipecat argues that voice agents require roughly 1,500 ms voice-to-voice response times, making low time-to-first-token performance essential, and reports that PhoneLLM can achieve high concurrency on NVIDIA B200 hardware, including sub-100 ms single-request P95 time-to-first-token. The weights are available through Hugging Face and can be served with SGLang, vLLM, or Modal AutoEndpoints, while the team positions the release as part of a broader shift toward smaller, specialized open models that can be continually refined with targeted evaluations, production data, and inference optimizations.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.