Qwen3.8-2.4T-A95B now available on Modal
Blog post from Modal
Qwen3.8-2.4T-A95B, an open-weights model that improves on Qwen 3.7 in coding, workplace, research, and long-horizon tasks, is now available through Modal’s Auto Endpoints and as an OpenAI-compatible Shared Endpoint. Modal worked with Qwen before launch to support the model using SGLang and a custom DFlash speculative decoding model optimized for Qwen3.8’s architecture. The speculator was trained with an emphasis on tool-call-heavy coding, research, and work sequences to improve drafted-token acceptance rates and accelerate inference. The text-only model will be offered for the next month with token-based pricing.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.