Home / Companies / Arize / Blog / November 2025

November 2025 Summaries

11 posts from Arize

Filter
Month: Year:
Post Summaries Back to Blog
Yongchao Chen, a Research Scientist Intern at Google and a PhD candidate at MIT and Harvard, presents his innovative paper on "TUMIX: Multi-Agent Test-Time Scaling with Tool-Use Mixture," which introduces an ensemble framework known as Tool-Use Mixture (TUMIX). This framework operates by running multiple agents in parallel, each utilizing different tool-use strategies and answer paths, and involves agents iteratively sharing and refining their responses based on questions and prior answers. Experimental results demonstrate that TUMIX significantly outperforms existing state-of-the-art methods in tool augmentation and test-time scaling, showcasing its potential to enhance AI capabilities.
Nov 24, 2025 121 words in the original blog post.
In an exploration of prompt optimization techniques, the blog post details how Claude Code, a leading coding agent using the Claude Sonnet 4-5 model, was enhanced using Prompt Learning, an approach inspired by reinforcement learning. This method focuses on optimizing the system prompts of coding agents based on their performance in handling datasets, specifically using SWE Bench Lite, a benchmark for evaluating coding models. Through a structured process involving meta-prompting and LLM feedback, Claude Code's performance improved significantly, achieving a 5.19% increase in general coding abilities and a 10.87% boost when tailored to specific repositories. This demonstrates the effectiveness of refining prompts without altering model architectures or tools, highlighting the potential of personalized, repository-specific optimization as a valuable asset for developers.
Nov 20, 2025 1,728 words in the original blog post.
Microsoft's AI Red Teaming Agent, integrated with Arize AX, offers a comprehensive approach to enhancing AI security by simulating adversarial attacks and identifying vulnerabilities in AI models. This method moves beyond traditional security testing by focusing on whether AI systems can be manipulated into generating harmful content, covering risk categories such as violence, sexual content, hate, and self-harm. Arize AX adds value by providing observability and evaluation, which allows attacks to be traced, weak points identified, and security improvements to be quantified. The process involves creating a feedback loop where attack data is used to optimize AI prompts, thereby enhancing the AI's defensive capabilities. By automating prompt optimization, the system continuously evolves to counter new attack strategies, resulting in AI models that are progressively safer and more trustworthy. This integrated workflow, emphasizing prompt optimization and continuous monitoring, ensures that AI systems not only withstand adversarial attacks but also improve over time, reinforcing Microsoft's commitment to responsible AI deployment.
Nov 19, 2025 1,557 words in the original blog post.
As AI systems evolve, the focus for enterprises has shifted from mere development to ensuring trust and responsibility in AI outputs. This necessitates a new integrated lifecycle that combines observability, evaluation, and experimentation, moving beyond traditional separate phases of model testing and deployment monitoring. Microsoft Foundry and Arize AX collaborate to provide a robust framework that supports continuous AI quality improvement through an ecosystem of flexible evaluation and observability tools. Microsoft Foundry offers enterprise-grade evaluation capabilities and agent development support, while Arize AX enhances observability and experimentation, allowing organizations to adapt new evaluators and models seamlessly. Together, they enable a feedback loop where data from model interactions is used to drive improvements, ensuring AI systems remain safe, fair, and compliant. This integration facilitates responsible AI at scale, providing automated monitoring, transparent governance, and continuous learning, with tools like Azure's content safety evaluators exemplifying how trace data, dataset benchmarking, and dashboard insights all contribute to a principled AI lifecycle.
Nov 18, 2025 2,211 words in the original blog post.
In the context of emerging software development paradigms, the text explores two distinct prompt optimization approaches, Prompt Learning and GEPA, which seek to enhance large language model (LLM) performance through feedback loops akin to reinforcement learning. Prompt Learning, developed by Arize AI, emphasizes high-quality evaluations and tailored meta-prompts to provide rich feedback without necessitating complete system overhauls, making it suitable for diverse frameworks and real-world production systems. GEPA, on the other hand, incorporates advanced algorithmic strategies such as evolutionary search and Pareto filtering, which are ideal for research environments that operate within controlled pipelines. Despite their differing methodologies, both frameworks aim to iteratively refine prompts by leveraging trace-level reflection and meta-prompting, with benchmarking results indicating that Prompt Learning achieves comparable or superior outcomes to GEPA using fewer rollouts. The text underscores the importance of evaluation quality and meta-prompt specificity over algorithmic complexity in driving meaningful improvements in LLM applications, advocating for a more structured and accessible prompt optimization process.
Nov 17, 2025 2,206 words in the original blog post.
Google's Agent Development Kit (ADK) and Arize AX platform are collaboratively advancing the deployment of multi-agent systems from prototypes to production, particularly demonstrated through a travel concierge system. ADK's modular, code-first architecture facilitates the orchestration of specialized agents, allowing for scalable applications integrated with various tools, while Arize AX offers robust observability and evaluation, ensuring clarity and optimization in agent behavior. This partnership emphasizes a seamless transition from development to production, leveraging comprehensive evaluation methods and real-time observability to enhance the reliability and efficiency of agent systems. By using a combination of OpenTelemetry standards and specialized evaluations, Arize AX captures extensive data on agent decisions, tool usage, and coordination quality, enabling continuous improvement through prompt optimization and experimentation. The travel concierge example showcases the integration of six specialized agents, each responsible for different phases of travel, coordinated by a root orchestrator, highlighting the powerful synergy between ADK's orchestration capabilities and Arize AX's observability infrastructure. Together, they provide a robust foundation for building scalable, reliable, and continuously improving agent systems, positioning organizations to leverage future AI advancements effectively.
Nov 14, 2025 1,811 words in the original blog post.
Prompt management tools have become integral to the effective use of AI language models by providing structured ways to manage, version, organize, and refine prompts, much like how engineers use GitHub for code. These tools are essential as prompts, comprising not just simple instructions but also model parameters, significantly influence AI behavior. The blog outlines the importance of prompt management, likening it to an infrastructure that ensures reproducibility, collaboration, and tracking of AI instructions, akin to software version control. It highlights five leading prompt management tools: Arize AX, Arize Phoenix, PromptLayer, DSPy, and PromptHub, each offering unique features such as sandbox testing, feedback loops, declarative modules, and collaborative workspaces to refine and optimize prompt usage. These tools enable teams to improve AI workflows by providing environments for testing, comparing, and enhancing prompts, ensuring consistent and reliable AI performance while integrating seamlessly with existing observability systems.
Nov 07, 2025 3,091 words in the original blog post.
Prompt management has become a crucial aspect of AI systems, influencing the performance and reliability of language models through structured workflows similar to software engineering practices. As AI prompts dictate model behavior by setting parameters and instructions, tools for managing these prompts have emerged to help teams organize, share, and refine them, ensuring consistent outputs and reproducibility. Central libraries, sandboxes, and feedback loops are key components of prompt management systems, providing environments for storing, testing, and improving prompts. Leading tools like Arize AX, Arize Phoenix, PromptLayer, DSPy, and PromptHub offer various features such as version control, analytics, and modular workflows, catering to diverse needs from enterprise-level deployments to open-source flexibility. The choice of tool depends on a team's specific requirements, whether they prioritize deep observability, hosted solutions, or collaborative environments, each facilitating the effective integration of prompts into AI workflows.
Nov 07, 2025 2,863 words in the original blog post.
Meta AI has developed ARE (Agents Research Environments) and Gaia2 as innovative platforms to enhance the development and evaluation of AI agents in complex, dynamic environments. ARE is designed to create time-driven worlds where agents can operate, adapt, and be verified, shifting the focus from simple tasks to intricate scenarios that unfold asynchronously with events and delays. This platform supports diverse applications and allows for agent collaboration and temporal responsiveness, integrating verification processes to ensure task completion. Gaia2, a benchmark within ARE, introduces 1,120 scenarios across 10 universes in a smartphone-like environment, emphasizing the ability of agents to adapt, collaborate, and respond promptly in real-time situations. The benchmark evaluates agents on various capabilities, including adaptability and the handling of ambiguities, revealing that while stronger reasoning models excel in certain areas, they often struggle with latency and time-sensitive tasks due to slower processing speeds.
Nov 06, 2025 686 words in the original blog post.
In October 2025, Arize AX introduced several new features to enhance AI agent engineering, including tags for organizing and labeling entities, a Data Fabric for synchronizing data with cloud warehouses, and automatic threshold ranges for monitors. Users can now benefit from API-driven monitors, which evaluate based on triggers rather than fixed schedules, and a new timeline tab for tracing that provides a detailed view of execution flows. Additional updates include tracing for authentication failures, a data region selector on the login page for improved compliance and performance, and expanded support for various LLMs, including Claude models and GPT open-source models. The platform has also improved data management with sorting options for datasets and experiments, as well as autocomplete for annotations, streamlining the process of building and managing evaluations.
Nov 04, 2025 567 words in the original blog post.
Hyland is integrating AI agent technology with their platforms such as OnBase, Alfresco, and Nuxeo to enhance document processing capabilities, emphasizing reliability and context-awareness in enterprise environments. By leveraging the Hyland Agent Builder and agentic document processing, the company aims to transition from impressive demonstrations to dependable, real-world applications, focusing on system behavior under actual conditions. Gabriel Keith, Senior Manager of Engineering, highlights the importance of determinism, human-in-the-loop strategies, and robust security measures, including multi-tenant cloud runtime and the Content Federation Service, which allows agents to interact with on-premises data securely. Hyland's approach includes integrating AI agents within existing business processes to accelerate decision-making and improve workflow efficiency. The platform's architecture, consisting of components like the Agent Builder and Core Runtime, facilitates seamless data access and integration with external and internal APIs, reducing the need for federation while maintaining security mappings. The company utilizes Arize AX for online evaluation and observability to ensure trust and consistency in agent performance, even as data landscapes evolve.
Nov 03, 2025 1,035 words in the original blog post.