Home / Companies / Hugging Face / Blog / October 2026

October 2026 Summaries

7 posts from Hugging Face

Filter
Month: Year:
Post Summaries Back to Blog
Microsoft and Hugging Face introduced ThinkingBox and ThinkingBox-Bench, an evaluation environment for AI agents that judges whether they produce the correct final database state and side effects rather than merely generating plausible responses or valid tool calls. Across 507 synthetic enterprise workflows in retail, insurance, travel, banking, and consulting, agents were tested 20 times per task in isolated MCP sessions to measure both single-run performance and repeatability. Results showed that many apparently successful runs still left incorrect, missing, or unintended changes in backend records, while reliability varied sharply among models: Claude Opus 5.5 led overall single-attempt performance at 67.16%, Kimi-K3 solved the most tasks at least once but was less consistent, and Claude Opus models completed the most tasks correctly across all 20 trials. The study also compares cost per successful attempt and per fully dependable task, finding that inexpensive single-run success does not necessarily translate into affordable consistency. Roughly four-fifths of failures were attributed to tool-use problems such as unrecovered errors or failed preconditions, suggesting that validation, retry mechanisms, restricted tool access, and human review for irreversible actions may improve deployed agents. ThinkingBox, its benchmark data, and an OpenEnv-based evaluation interface are publicly available through Hugging Face and Microsoft repositories.
Oct 03, 2026 3,146 words in the original blog post.
Rapidata presents a method for incorporating real-time human preference feedback into the post-training of image and video diffusion models, aiming to avoid limitations of offline preference datasets and learned reward models, including divergence from human judgment and reward hacking. Using Flow-GRPO as an example, the approach replaces automated scoring of groups of generated images with crowd-sourced pairwise comparisons collected through Rapidata Flows, which converts votes into Elo-style scores using a Bradley-Terry model and feeds normalized group-relative advantages back into the training loop. The system is designed to evaluate many image groups asynchronously to reduce GPU idle time, with options to provide prompts as evaluation context and configure response targets and deadlines. The article also discusses practical considerations including LoRA-based fine-tuning, pipelining generation and evaluation, estimated annotation throughput for large GPU deployments, and centralized authentication for distributed training, while noting that the same feedback mechanism could be adapted to other online preference-optimization algorithms.
Oct 02, 2026 2,210 words in the original blog post.
VIDRAFT’s FINAL-Bench project released POCKET-Darwin-180B, a 4-bit GGUF version of the 180-billion-parameter Darwin-180B-RSI mixture-of-experts model, reducing its size from 360 GB to 111 GB while claiming to preserve its 87.65% accuracy on a 2,000-question MMLU-Pro evaluation. The model can run on systems ranging from an 8 GB VRAM laptop with 32 GB RAM, where selected experts are streamed from SSD, to CPU-only servers or 128 GB-memory mini PCs where more weights can remain resident. Its feasibility is attributed to the MoE architecture, which activates roughly 3 billion parameters per token, llama.cpp memory mapping, and a “graft quantization” approach that retains most of an existing quantized base model while replacing 300 tensors altered by self-improvement training. The underlying Darwin model is based on Alibaba’s Qwen3.8-Flash-Next and was trained through recursive self-improvement using only automatically verified model-generated solutions, with its developers reporting strong leaderboard results across reasoning, knowledge, vision, and law benchmarks. The release is presented as useful for organizations requiring local, offline inference for sensitive workloads, and it is available through Hugging Face and ModelScope under the Qwen Community License.
Oct 02, 2026 1,058 words in the original blog post.
llama.cpp now supports decision models through its `/v1/systemone` endpoint, enabling applications to submit text, JSON, screenshots, or chat messages alongside typed questions and receive option probabilities in a single forward pass rather than generated text. Based on TypeSafe’s System One API format, the feature supports choice, scored-scale, and yes/no questions for tasks such as request routing, content moderation, agent-step verification, and action selection. Available models range from the 144M-parameter Julia-1 to the 27B OpenJev, with differing language, image, licensing, and performance capabilities; OpenJev currently supports image inputs. Users can run models locally, load models dynamically in router mode, batch independent questions, and select quantizations, while guidance emphasizes adding clear option descriptions, comparing models, and calibrating confidence thresholds on application-specific examples. The project plans to add newly released open decision models, with Cloudflare’s Clef identified as a forthcoming option.
Oct 02, 2026 1,021 words in the original blog post.
Ai2 has open-sourced AstaBrief 8B, a Qwen3-8B-based model designed to generate cited scientific literature reports from research questions and retrieved excerpts, alongside its weights, training data, and a local PDF-report workflow. Available as Asta’s Fast mode, it produces reports in a single pass rather than through the slower multistep, section-by-section Claude-powered Thinking mode, averaging 51.1 seconds per report versus 178.5 seconds. The model was trained using 47,000 supervised examples derived from filtered real scientific queries and 6,000 preference pairs judged by GPT-4.1 and DeepSeek-R1, with an emphasis on evidence grounding, relevance, report structure, and citation support. Ai2 found that filtering training reports for citation density yielded especially strong improvements, while DPO further improved performance toward that of its proprietary pipeline and DR Tulu in development evaluations. The organization notes that the reported comparisons reflect 2025 models and methods rather than current frontier performance, but argues that the training-data, attribution-filtering, and efficient-serving lessons may generalize. Early Asta usage suggests Fast mode receives feedback comparable to Thinking mode, and open weights enable institutions to run the model locally for sensitive or unpublished research; future work will explore richer preference learning, retrieval-augmented reinforcement learning, multi-tool capabilities, broader data sources, and evaluations of whether reports preserve the scope and strength of scientific evidence.
Oct 02, 2026 2,815 words in the original blog post.
AutoSynthData is a ServiceNow CoreAI pipeline for producing enterprise-specific synthetic training data by identifying where an agent fails, using stronger teacher models to characterize successful behavior, and generating new executable tasks that target those capability gaps. Each task includes a system specification, user prompt, and verifier, and is designed to be feasible in the environment, realistic for enterprise workflows, and difficult enough to provide learning value. The system creates core tasks and varied derivatives, validates reference solutions through positive and negative checks, repairs flawed candidates, and reviews batches for diversity, coverage, and redundancy. As models improve through supervised fine-tuning, the process shifts toward remaining weaknesses, creating an adaptive curriculum near the model’s capability boundary. In EnterpriseOps Gym experiments, fine-tuning Gemma-4-26B-A4B-it on roughly 2,000 generated samples improved Hybrid-domain Pass@1 by 7.2 percentage points, from 63.01% to 68.55% verifier success, and raised ITSM Pass@1 from 18.77% to 27.18%, suggesting that validated synthetic tasks can improve performance in stateful enterprise environments.
Oct 02, 2026 2,242 words in the original blog post.
Ai2 has released Olmo-core 3, an open training framework redesigned to efficiently train mixture-of-experts language models at scales reaching trillions of parameters. The system replaces its earlier fully sharded approach with distributed data parallelism that keeps experts on GPUs and routes tokens to them, improving throughput in a preliminary 47-billion-parameter benchmark by about 2.7 times while allowing expert capacity to increase substantially with limited performance loss. It combines expert, pipeline, and optimizer parallelism with routing and computation optimizations such as GPU-resident routing, grouped matrix operations, and MXFP8 low-precision computation, which raised measured throughput by roughly 21% and reduced memory use in one test. Benchmarks included a 1.2-trillion-parameter configuration across 512 GPUs and a shorter capacity test reaching 2.38 trillion parameters, although these evaluated systems performance using random routing rather than trained-model quality. The accompanying report also examines practical training findings, including pitfalls in balancing expert workloads and the fact that overlapping communication with computation does not always improve speed. Olmo-core 3 will support Ai2’s next MoE-based Olmo model and is publicly available for researchers and developers to adapt, evaluate, and use for their own training experiments.
Oct 01, 2026 1,437 words in the original blog post.