July 2026 Summaries
4 posts from Featherless
Filter
Month:
Year:
Post Summaries
Back to Blog
Moonshot has introduced Kimi K3, an open-weight AI model boasting nearly 3 trillion parameters, significantly expanding upon its predecessor, K2, which had about 1 trillion parameters. This model is notable for its innovative architecture, including the Kimi Delta Attention mechanism, which enhances scaling efficiency, and its capability for native multimodality, allowing it to process images directly with an upgraded vision system. Kimi K3 has demonstrated competitive performance on benchmarks like Terminal-Bench 2.1 and GPQA Diamond, matching or surpassing closed-weight models such as Claude Fable 5 and GPT-5.6 Sol. Its availability on Featherless, with features like a 32K context and FP8 quantization, aims to make powerful open models accessible to all, with options for dedicated GPU clusters to support larger-scale operations. This development has drawn attention, particularly in Washington DC, as it represents a significant advance from an open-weight model originating in China.
Jul 29, 2026
490 words in the original blog post.
Featherless offers a range of uncensored language models through an OpenAI-compatible API, catering to various use cases such as general reasoning, creative writing, security research, and fast general chat. Notable models include Huihui-Qwen3.5-27B-abliterated for general uncensored work, Gemma-4-26B-A4B-it-uncensored for speed, Huihui-Mistral-Small-3.2-24B-Instruct-2506-abliterated for creative writing and roleplay, and Qwythos-9B-Claude-Mythos-5-1M for security research and coding. These models are available on a flat-rate Chat plan that does not log any interactions, ensuring privacy for users. The abliterated models suppress refusal behavior without retraining, providing an efficient uncensoring method. Featherless supports users with a transparent pricing plan and encourages exploring their catalog to find the most suitable model for specific tasks.
Jul 24, 2026
791 words in the original blog post.
Featherless.ai, in collaboration with Z.ai, has launched GLM 5.2 access worldwide, marking a significant advancement in open-source AI infrastructure. This release is part of Featherless's broader mission to make AI practical and reliable at any scale, supported by a recent $20 million Series A financing round to enhance their infrastructure offerings. Featherless has also become Hugging Face's largest LLM inference provider with over 6,700 models available, offering unlimited token pricing and instant access for open-source AI deployment. Their serverless inference platform democratizes AI access with over 4,000 models, allowing users to start building AI solutions in under three minutes.
Jul 21, 2026
194 words in the original blog post.
Featherless AI offers dedicated GLM 5.2 instances optimized for high-volume, recurring workloads using AMD MI325X GPUs, starting at $7,592 per month. These instances provide significant cost savings, reducing monthly inference costs for major clients by up to 95% compared to previous setups. The GLM 5.2 model, released during a significant industry disruption, offers an open-weight alternative that mitigates regulatory risks and enhances control over AI deployment. Its design allows for reduced costs, reserved capacity, and enhanced privacy, making it a resilient choice amidst geopolitical uncertainties. The model's performance is optimized through techniques like DSpark speculative decoding and custom AMD kernels, ensuring efficient execution and scalability. Featherless AI promotes a flexible, open approach to AI, emphasizing the importance of accessibility and control in a landscape where reliance on closed models poses continuity risks.
Jul 16, 2026
1,521 words in the original blog post.