Home / Companies / Hugging Face / Blog / Post Details
Content Deep Dive

Base Optimization Stack: From Open Weights to Frontier On-Device Inference Speed

Blog post from Hugging Face

Post Details
Company
Date Published
Author
Fabian Waschkowski
Word Count
1,338
Company Posts That Month
74
Language
-
Hacker News Points
-
Post removed?
No
Summary

BaseCompute presents its Base Optimization Stack (B:OS), an agent-driven pipeline designed to convert newly released open-weight models into device-specific, optimized BaseRT inference releases through quantization, architecture porting, correctness validation, and kernel-level performance tuning. Using NVIDIA’s 31.6B-parameter Nemotron 3 Nano hybrid MoE model as a demonstration, the company added support for Mamba-2, sparse expert routing, and other previously unsupported components on Apple silicon, while requiring unit tests and perplexity-based accuracy gates for each change. It reports that tuning increased prefill performance by roughly 10.8–12.8 times over an untuned port, while decode improved 1.3 times due to memory-bandwidth limits. In comparisons on the same hardware, BaseRT reportedly exceeded llama.cpp by 1.39–1.76 times on prefill and 1.90 times on decode, and MLX by 1.98–2.55 times on prefill and 1.43 times on decode. Tests involving Claude Fable 5, Kimi K3, and locally run GLM 5.2 suggest that more capable or costly agents can achieve higher optimization scores, although lower-cost and local models retained much of the performance. The company argues that accumulated optimization knowledge can reduce support time and cost across models and hardware, citing a 27% faster process for a subsequent NVIDIA DGX Spark port.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
MLX 6 23 9 4 -28%
LLM 1 4,718 960 222 -38%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.