Introducing Qwen3.5 4B on LLM Gateway—optimized by AssemblyAI for the fast rewrite tasks at the center of voice products
Blog post from AssemblyAI
AssemblyAI has added the open-source Qwen3.5 4B model to its LLM Gateway as a self-hosted, latency-optimized option for voice-text rewriting tasks such as dictation cleanup, transcript formatting, live formatting, and conversational turn summarization. The company reports that its qwen3.5-4b-32k-fast deployment averaged 612 milliseconds on representative voice rewrite benchmarks, making it 1.9 times faster and 94% less expensive per hour of audio than GPT-4.1, with pricing of $0.10 per million prompt tokens and $0.50 per million completion tokens. The model is intended to remove filler words, add punctuation, resolve spoken corrections, and produce polished text from speech-to-text output, but it supports only token limits, temperature, and streaming rather than tool calling or structured outputs. AssemblyAI recommends pairing its Sync speech-to-text API with the model for dictation features, while directing voice-agent developers needing tool use and multi-step reasoning toward the larger qwen3-next-80b-a3b model or frontier models. Qwen3.5 4B, alongside Claude, GPT, Gemini, and other models, is available through AssemblyAI’s OpenAI-compatible endpoint, with free credits offered to new users.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.