Home / Companies / Fireworks AI / Blog / Post Details
Content Deep Dive

DeepSeek-V4.1-Flash on Fireworks: Astra-level DeepSWE at 1/15th the cost

Blog post from Fireworks AI

Post Details
Company
Date Published
Author
-
Word Count
1,351
Company Posts That Month
4
Language
English
Hacker News Points
-
Post removed?
No
Summary

Fireworks reports that DeepSeek-V4.1-Flash establishes a strong cost-performance position for autonomous software engineering, achieving a 74.34% DeepSWE pass rate comparable to GPT-6 Astra, Gemini 3.8 Flash, and Claude Opus 5 while costing about $0.43 per task, or roughly 15 times less than Astra. The 552-billion-parameter mixture-of-experts model uses a split activation design with 8 billion active input parameters and 16 billion output parameters, aiming to reduce costs for coding agents that repeatedly consume far more input context than they generate. Improved KV-cache efficiency, including lower HBM and SSD requirements, is presented as especially important because cached input represented most spending in long agent trajectories. On Terminal-Bench 2.1, it reportedly reached 86.5% accuracy, one percentage point below Astra, at about one-twelfth the total cost per task. However, it performed less strongly on Humanity’s Last Exam, scoring 34.52% compared with Astra’s 50.40%, suggesting it is better suited to agentic coding than difficult academic reasoning. An oracle-routing evaluation found that combining both models could reach 54.80% on HLE, indicating that their differing strengths may improve results in multi-model systems.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Cost per task 4 10 5 5 -84%
Serverless 2 156 54 28 -80%
LLM 1 747 162 79 -85%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.