October 2026 Summaries
9 posts from Arize
Filter
Month:
Year:
Post Summaries
Back to Blog
No summary generated yet.
Oct 09, 2026
1,293 words in the original blog post.
No summary generated yet.
Oct 08, 2026
637 words in the original blog post.
A benchmark of eight recently released decision models assessed their ability to detect hallucinations, evaluate LLM outputs, and route agent tool or skill choices, comparing accuracy, latency, cost, calibration, and deployment options across hosted APIs and locally run open models. Jev, Liquid d1, and Clef 27B were closely matched on core accuracy measures and approached the performance of a far more expensive LLM judge, while Jev generally offered strong speed and calibration among APIs, Kev emerged as the leading local model, and Clef 27B stood out as an open-weights option with consistently competitive accuracy. In the final routing test, Jev and Kev effectively tied, making the preferred deployment model—managed API or self-hosting—the main practical differentiator. A major finding was that several models performed very differently when equivalent evaluation questions were phrased with reversed polarity, with some producing confidently inverted results, underscoring that benchmarks and production evaluators should test prompt wording as carefully as model performance.
Oct 07, 2026
2,802 words in the original blog post.
tracelint is an open-source linter for OpenInference agent traces, including traces collected by Arize Phoenix, designed to detect structural agent failures that can be proven deterministically rather than judged semantically by an LLM. It distinguishes hard defects, such as invalid tool arguments, nonexistent tool calls, or reuse of a declared failed result in a side-effecting action, from candidate issues like repetitive calls that may be legitimate retries, while marking insufficiently evidenced cases as not checked. A release-agent example shows how an agent deployed a Jenkins build marked UNSTABLE despite a policy to deploy only successful pipelines; by declaring in a tools.json file that UNSTABLE represents failure and that deployment has side effects, tracelint can identify the trace as a hard defect and return a CI-failing exit code. The tool can generate initial tool definitions from traces, integrate with pytest and GitHub Actions, and complements rather than replaces agent evaluations, which remain necessary for assessing task-specific correctness, strategy, grounding, and policy compliance.
Oct 06, 2026
1,294 words in the original blog post.
Agent benchmarking is difficult because realistic predeployment evaluation requires reproducible tasks, data, tools, environments, and resettable state, while simplified mocks often omit critical frontend context, external services, and real database behavior. The Phoenix team encountered these limitations while testing its PXI coding agent and adopted Harbor, an open-source sandboxing framework that runs agents in fresh local or cloud environments and uses verifiers to evaluate outcomes. Harbor enabled testing PXI against a real Phoenix server and seeded data, while also supporting evaluations of MCP servers, skills, CLIs, and coding agents such as Claude Code and Codex using shared tasks and outcome-based measures of correctness and efficiency. A Phoenix plugin records Harbor tasks as versioned datasets, agent configurations as experiments, verifier results as annotations, and agent trajectories as traces, allowing teams to inspect tool use, model calls, failures, turns, and task-version changes. Although agents may achieve similarly correct answers, trace analysis can expose substantial differences in their workflows, helping developers refine prompts, skills, and tools to reduce unnecessary steps without reducing performance.
Oct 05, 2026
891 words in the original blog post.
Arize argues that MCP, command-line interfaces, and agent skills should be treated as complementary ways to extend AI agents rather than competing alternatives, a view reflected in its new hosted MCP server for Arize AX alongside its existing CLI and installable skills. Earlier concerns that MCP imposed heavy context costs were valid when tool definitions loaded at session startup, but modern agent harnesses such as Claude Code, Cursor, and Codex now support deferred tool loading and code execution, substantially reducing that overhead. The hosted AX MCP server provides 45 named, authenticated tools for accessing projects, traces, datasets, experiments, evaluators, monitors, and prompts from MCP-compatible clients without requiring local installation or shell access, while the AX CLI supports scripting, automation, CI, and other noninteractive environments. Skills add procedural guidance by teaching agents how to perform multistep tasks using the CLI, with instructions loaded only when relevant. Arize recommends choosing MCP for client-based discovery, structured permissions, and shell-less users, choosing CLI for scriptable and execution-heavy workflows, and combining both with skills for production use, all backed by a shared REST API.
Oct 02, 2026
2,152 words in the original blog post.
Prompt caching can reduce repeated-input costs in multi-turn agents by reusing instructions, context, and conversation history, but benchmark results using a 20-conversation e-commerce assistant workload showed that high cache reuse alone does not guarantee lower overall cost. Harbor and Arize Phoenix compared four model/provider paths across repeated conversations of 5 to 20 turns, finding that DeepSeek had the highest aggregate cache-read rate at 93.6%, the lowest estimated cost at $1.17 per 100 runs, and the fastest average latency at 34.7 seconds, while Claude had an 89.8% cache-read rate but the highest estimated cost of $14.63 because output tokens accounted for much of its spending. Cache reuse increased with conversation length for all models, with GPT showing the largest increase from 34.2% at five turns to 87.0% at 20 turns, whereas DeepSeek began with high reuse even in shorter conversations. The comparison reflects combined model and provider behavior rather than caching efficiency in isolation, and Phoenix’s experiment and trace views allow teams to inspect cache reads, token usage, estimated costs, latency, and individual LLM calls. The results emphasize that organizations should evaluate caching on their own workloads while considering output volume, pricing, latency, and task quality alongside cache-read rates.
Oct 02, 2026
1,416 words in the original blog post.
Anthropic’s Claude API skill introduces build-eval, which helps create evaluations through an interview process, and hillclimb, which iteratively changes an agent configuration, retains improvements based on training scores, and uses held-out tests to detect overfitting. The discussion endorses this approach while arguing that production-grade agent optimization requires evaluation datasets derived from real traces, human selection and labeling of meaningful failures, combinations of deterministic code checks and calibrated LLM judges, isolated test environments that prevent benchmark leakage, and measurement of cost and latency alongside quality. Arize describes its AX platform as extending the workflow by storing traces and experiment results centrally, enabling teams to compare runs, investigate poor outcomes, continuously evaluate live traffic, identify emerging failure patterns, and add production failures back into test sets. The central recommendation is to treat hillclimbing not as a one-time local optimization exercise but as a continuous feedback loop in which production behavior, human judgment, reproducible experiments, and quality-cost tradeoffs guide agent improvements.
Oct 02, 2026
2,577 words in the original blog post.
Arize AX has introduced long-term memory for Alyx, its AI engineering agent, allowing it to retain durable user-specific context such as project goals, workflow conventions, preferences, decisions, and rejected approaches across separate sessions. Enabled by default, the feature is available throughout traces, evaluations, experiments, and other AX workflows, helping users continue work without repeatedly explaining key requirements, such as policy constraints identified during production investigations. Alyx is designed to retrieve current configurations, IDs, and live metrics directly from the platform rather than store potentially outdated operational details, reserving memory for context that is difficult to reconstruct. Memories are scoped to an individual user and Arize space, are not shared with teammates or other spaces, and can be viewed, corrected, or deleted through chat.
Oct 01, 2026
577 words in the original blog post.