Home / Companies / Daily / Blog / August 2026

August 2026 Summaries

1 posts from Daily

Filter
Month: Year:
Post Summaries Back to Blog
Pipecat has released PhoneLLM Alpha 1, a BSD-licensed open-weights voice-agent model fine-tuned from NVIDIA’s Nemotron 3 Nano 30B-A3B, designed to deliver accurate tool use and fast multi-turn conversations for customer-service and outbound calling applications. With 3.5 billion active mixture-of-experts parameters, the model is intended to offer lower latency and operating costs than larger general-purpose models while remaining deployable on private infrastructure. The company also introduced PhoneBench v1, which evaluates phone-agent models using calibrated LLM judges across factors including speaking style, tool-call accuracy, factual grounding, conversation coherence, escalation behavior, latency, and estimated per-minute cost. Pipecat argues that voice agents require roughly 1,500 ms voice-to-voice response times, making low time-to-first-token performance essential, and reports that PhoneLLM can achieve high concurrency on NVIDIA B200 hardware, including sub-100 ms single-request P95 time-to-first-token. The weights are available through Hugging Face and can be served with SGLang, vLLM, or Modal AutoEndpoints, while the team positions the release as part of a broader shift toward smaller, specialized open models that can be continually refined with targeted evaluations, production data, and inference optimizations.
Aug 27, 2026 2,109 words in the original blog post.