August 2026 Summaries
5 posts from Featherless
Filter
Month:
Year:
Post Summaries
Back to Blog
Z.ai has identified the previously anonymous ox-alpha coding model as GLM-5.3-Flash, a 320-billion-parameter open-weight, MIT-licensed, natively multimodal mixture-of-experts model that was previewed using Chinese AI chips and is now available through Featherless. Compared with earlier GLM models, it uses fewer active parameters and layers while introducing combined linear and sparse attention, IndexPool compression for long contexts, and manifold-constrained hyper-connections to reduce compute and KV-cache requirements; it was trained on a 30-trillion-token multimodal corpus and supports image and video understanding. Z.ai reports that the model outperforms GLM-5.2 at roughly one-tenth the cost and posts competitive coding, agentic, and vision results against models including Claude Opus 4.8, GPT-5.6 Terra, and Gemini 3.7 Flash, while third-party Artificial Analysis assigns it an Intelligence Index score equal to Opus 4.8 at substantially lower estimated cost. The reported comparisons carry caveats, including reliance on Z.ai’s internal benchmarks, a “Flash” label referring to pricing rather than inference speed, and the benchmarked Opus version having since been superseded. Featherless serves the model through an OpenAI-compatible API with a 256K-token context window, FP8 quantization, tool calling, adjustable reasoning effort, and a stated no-logging policy, while offering dedicated GPU clusters for larger-scale or full-million-token deployments.
Aug 27, 2026
875 words in the original blog post.
Moonshot’s Kimi K3 and Alibaba’s Qwen3.8 releases represent competing large hybrid-attention mixture-of-experts models with million-token context ambitions, though their publicly available checkpoints differ substantially. Kimi K3 is a 2.8-trillion-parameter multimodal MoE with 104 billion active parameters, a built-in vision encoder, and native 4-bit MXFP4 weights that reduce its storage footprint to roughly 1.4–1.56 TB, while Qwen3.8-Max is an API-only multimodal flagship and the open Qwen3.8-2.4T-A95B checkpoint is text-only, with 2.4 trillion total and 95 billion active parameters; Alibaba also offers the much smaller, Apache 2.0-licensed multimodal Qwen3.8-27B. Both use linear-attention and full-attention layers to make long contexts more practical, but K3 uses more experts and quantization-aware training, whereas Qwen ships primarily at full precision with an FP8 option. Vendor-reported overlapping benchmarks generally place K3 ahead on coding, tool-use, and agent evaluations, while Qwen3.8-Max remains close on reasoning tests and is positioned as stronger for instruction following and long-document retrieval; independent early rankings similarly place K3 slightly ahead among open-weight models. Operating the largest models requires substantial infrastructure, but Qwen3.8-Max’s API pricing is notably lower, particularly for output tokens generated during reasoning, while Featherless provides K3 at up to 256K context and Qwen3.8-27B as a lower-cost deployment option.
Aug 26, 2026
934 words in the original blog post.
Moonshot’s Kimi K3 is an open-weight mixture-of-experts model with 2.8 trillion total parameters, roughly 104 billion active parameters per token, a native one-million-token context window, multimodal capabilities, and tool calling designed for agent workflows; Moonshot reports frontier-level benchmark performance, although some results use differing test harnesses. Because its approximately 1.56 TB weights and recommended 64-plus-accelerator deployment make local operation impractical, Featherless offers it through an OpenAI-compatible API under the model ID moonshotai/Kimi-K3, providing 256K context on its serverless Developer plan and up to one million tokens on dedicated infrastructure. The setup process for OpenCode, Claude Code via Claude Code Router, Cline, Roo Code, Aider, and other compatible agents mainly requires the Featherless base URL, an API key, model ID, and appropriate context-limit settings, while custom agent implementations must preserve reasoning and tool-call data across turns. The guide distinguishes pay-as-you-go serverless use for occasional or individual development from dedicated GPU deployments for sustained, high-concurrency agent workloads, arguing that reserved hardware can offer more predictable latency and lower costs at scale; listed serverless pricing is $2 per million input tokens, $0.30 per million cached tokens, and $10 per million output tokens. Although K3’s weights permit commercial use under a conditional license, its training data and training recipe are not public, and its coding performance is presented as competitive with leading proprietary models but dependent on the task and token usage.
Aug 24, 2026
1,479 words in the original blog post.
AI token consumption is rising rapidly as models gain larger context windows, longer-running agent capabilities, and the ability to coordinate subagents, making cost management increasingly important. Major providers generally charge for input, output, and cached tokens, while the most useful cost metric for agentic systems is increasingly the cost per successfully completed task rather than the cost of individual conversations or tool calls. Recommended optimization approaches include routing simpler work to smaller or specialized models, using lightweight agent harnesses with limited prompts and essential tools, and maintaining organizational memory to avoid repeatedly rediscovering information. The text also highlights capacity planning, suggesting that lower-priority asynchronous tasks could use slower, cheaper inference during off-peak periods. For workloads where these measures are insufficient, it promotes dedicated GPU infrastructure with flat-rate pricing as a potentially cheaper alternative to fully pay-per-token usage, while retaining elastic capacity for traffic spikes.
Aug 14, 2026
1,163 words in the original blog post.
AT&T, Microsoft, and AMD have released OTel 2.0, an open telecommunications-focused language model available through Featherless that is trained on publicly available industry standards from organizations including 3GPP, ETSI, and the GSMA. Built by post-training the open Gemma 4 model, OTel 2.0 is designed for specialized network-engineering tasks where general-purpose models may be less reliable, and it reportedly led the GSMA Open-Telco benchmark’s TeleLogs troubleshooting task against larger general models. The release is presented as evidence that enterprise AI may increasingly shift toward open, domain-specific models trained on industry knowledge, particularly in sectors with extensive technical documentation and standards. OTel 1.0 was downloaded more than 25 million times, while version 2.0 was trained at production scale using AMD hardware, and Featherless argues that similar specialized models could benefit underrepresented but economically significant fields such as transportation, construction, and agriculture.
Aug 13, 2026
640 words in the original blog post.