Home / Companies / LabelBox / Blog / June 2025

June 2025 Summaries

2 posts from LabelBox

Filter
Month: Year:
Post Summaries Back to Blog
Agentic AI is transforming how models interact with the world by proactively making decisions and executing complex tasks with minimal guidance, necessitating advanced training techniques and high-quality data. Labelbox collaborates with AI labs to develop data infrastructure that supports the creation of agentic systems, which require detailed feedback, verifiable outcomes, and scalable evaluation pipelines. The company has worked on projects that involve simulating complex tool use, verifying structured reasoning, and benchmarking multi-turn instruction-following, demonstrating the capacity of AI to plan, adapt, and respond autonomously in real-world contexts. These initiatives include developing environments for multi-step API interactions, creating benchmarks for planning tasks with constraints, and evaluating models' ability to follow evolving instructions. Through these projects, Labelbox aims to benchmark and enhance the performance of agentic AI systems, which are crucial for AI assistants and tool-using agents, by providing flexible infrastructure for training and evaluation.
Jun 30, 2025 1,442 words in the original blog post.
Labelbox conducted a comprehensive study to evaluate the performance of three advanced language models with native web search capabilities: Google Gemini 2.5 Pro, OpenAI GPT-4.1, and Anthropic Claude 4.0 Opus. The study aimed to assess their effectiveness in providing accurate, current, and diverse responses to 200 complex and varied queries across multiple domains, including STEM, current events, historical information, and multi-language contexts. The evaluation focused on four key dimensions: source quality, answer relevance, information recency, and multi-language understanding. Results showed that Gemini excelled in recency and scientific source access, GPT-4.1 was strong in synthesis and reasoning, while Claude was noted for its clarity in explanation. However, citation reliability emerged as a common weakness across all models, highlighting the need for improved source verification processes in enterprise applications. The study emphasizes the importance of considering question types and domain specializations when deploying these models for enterprise use.
Jun 13, 2025 1,352 words in the original blog post.