DeepSeek-V4.1-Flash on Fireworks: Astra-level DeepSWE at 1/15th the cost
Blog post from Fireworks AI
Fireworks reports that DeepSeek-V4.1-Flash establishes a strong cost-performance position for autonomous software engineering, achieving a 74.34% DeepSWE pass rate comparable to GPT-6 Astra, Gemini 3.8 Flash, and Claude Opus 5 while costing about $0.43 per task, or roughly 15 times less than Astra. The 552-billion-parameter mixture-of-experts model uses a split activation design with 8 billion active input parameters and 16 billion output parameters, aiming to reduce costs for coding agents that repeatedly consume far more input context than they generate. Improved KV-cache efficiency, including lower HBM and SSD requirements, is presented as especially important because cached input represented most spending in long agent trajectories. On Terminal-Bench 2.1, it reportedly reached 86.5% accuracy, one percentage point below Astra, at about one-twelfth the total cost per task. However, it performed less strongly on Humanity’s Last Exam, scoring 34.52% compared with Astra’s 50.40%, suggesting it is better suited to agentic coding than difficult academic reasoning. An oracle-routing evaluation found that combining both models could reach 54.80% on HLE, indicating that their differing strengths may improve results in multi-model systems.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Cost per task | 4 | 10 | 5 | 5 | -84% |
| Serverless | 2 | 156 | 54 | 28 | -80% |
| LLM | 1 | 747 | 162 | 79 | -85% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.