Base Optimization Stack: From Open Weights to Frontier On-Device Inference Speed
Blog post from Hugging Face
BaseCompute presents its Base Optimization Stack (B:OS), an agent-driven pipeline designed to convert newly released open-weight models into device-specific, optimized BaseRT inference releases through quantization, architecture porting, correctness validation, and kernel-level performance tuning. Using NVIDIA’s 31.6B-parameter Nemotron 3 Nano hybrid MoE model as a demonstration, the company added support for Mamba-2, sparse expert routing, and other previously unsupported components on Apple silicon, while requiring unit tests and perplexity-based accuracy gates for each change. It reports that tuning increased prefill performance by roughly 10.8–12.8 times over an untuned port, while decode improved 1.3 times due to memory-bandwidth limits. In comparisons on the same hardware, BaseRT reportedly exceeded llama.cpp by 1.39–1.76 times on prefill and 1.90 times on decode, and MLX by 1.98–2.55 times on prefill and 1.43 times on decode. Tests involving Claude Fable 5, Kimi K3, and locally run GLM 5.2 suggest that more capable or costly agents can achieve higher optimization scores, although lower-cost and local models retained much of the performance. The company argues that accumulated optimization knowledge can reduce support time and cost across models and hardware, citing a 27% faster process for a subsequent NVIDIA DGX Spark port.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.