DeepSeek-V4.1-Flash is live on Featherless
Blog post from Featherless
DeepSeek released V4.1-Flash on September 10 as an MIT-licensed, natively multimodal mixture-of-experts model with 763 billion total parameters, 8 billion active during prefill, 16 billion active during decoding, and a native 1 million-token context window, replacing V4-Flash and V4-Flash-Vision-Exp. Its main advancement is substantially reduced memory use: FP4 KV-cache compression and SWA Bounded Replay reportedly reduce cache requirements to roughly one-quarter and then one-eighth of the prior model’s persistent footprint. The architecture combines a 40-layer causal encoder-decoder, sparse attention modes that share KV states across layers, 384 routed experts per MoE layer, and a separate 196 billion-parameter conditional memory system, while multimodal input is supported through a DeepSeek-ViT image encoder. DeepSeek reports competitive results on coding, automation, and reasoning benchmarks, although it trails Claude Opus 5.0 and GPT-5.6 Sol on some evaluations, while Artificial Analysis independently ranked it sixth of 113 models on its Intelligence Index and fourth for output speed. Featherless hosts the FP8-quantized model through an OpenAI-compatible API with a 256,000-token context limit, charging per-token rates for input, cached input, and output, and also promotes dedicated GPU deployments for organizations needing larger-scale or full-context inference.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| AI Model Fine-tuning | 1 | 139 | 28 | 14 | -75% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.