February 2026 Summaries
2 posts from LabelBox
Filter
Month:
Year:
Post Summaries
Back to Blog
AI models are often evaluated for safety using benchmarks that test their ability to refuse harmful requests, but recent research highlights significant shortcomings in these assessments. The study examines the quality of two widely used safety datasets, AdvBench and HarmBench, revealing that they do not accurately reflect real-world adversarial behavior due to their reliance on "triggering cues"—overtly negative or sensitive expressions designed to activate safety mechanisms. This reliance leads to inflated safety evaluations, as models appear safe when they are merely responding to these cues rather than resisting genuine malicious intent. The research introduces the concept of "intent laundering," a technique that removes triggering cues while preserving malicious intent, demonstrating that models considered safe often fail when these cues are absent. This exposes a gap between current safety evaluations and real-world threats, suggesting that AI safety research must develop better benchmarks and alignment techniques to more accurately model harmful behavior and improve the robustness of AI models against realistic misuse scenarios.
Feb 20, 2026
2,250 words in the original blog post.
Labelbox has acquired Upcraft, a company specializing in AI-powered sales automation, to enhance its AI model training capabilities by integrating AI agents into its operations. The acquisition aims to strengthen Alignerr, Labelbox's network of over one million domain experts, by automating complex workflows and improving how the network recruits and engages experts to generate high-quality training data. Upcraft, founded in 2021, brings its expertise in automating sophisticated sales workflows, which will now be leveraged to support AI model development. Greg Caplan, Upcraft's CEO, expressed excitement about joining Labelbox to contribute to their vision of advancing superintelligence and making AI more accessible. Labelbox CEO Manu Sharma emphasized the importance of connecting elite domain experts with development teams to secure high-quality, expert-driven training data, which is crucial as AI labs invest heavily in advancing AI technologies. The integration of Upcraft's capabilities is expected to transform the growth and operation of the Alignerr network, ultimately enhancing the quality and scalability of AI models.
Feb 10, 2026
337 words in the original blog post.