Home / Companies / Arize / Blog / Post Details
Content Deep Dive

Trace-Level LLM Evaluations with Arize AX

Blog post from Arize

Post Details
Company
Date Published
Author
Sanjana Yeddula
Word Count
583
Company Posts That Month
11
Language
English
Hacker News Points
-
Post removed?
No
Summary

The text discusses the importance and methodology of trace-level evaluations for Large Language Model (LLM) applications, as opposed to the more common span-level evaluations. While span-level assessments focus on individual steps such as tool calls or LLM responses, trace-level evaluations provide a comprehensive view of the entire workflow, assessing the success, efficiency, and relevance of the final outcome. The tutorial highlights the use of Arize AX for conducting these evaluations and provides an example through a movie recommendation agent, which utilizes multiple tools to deliver a comprehensive answer to user queries. By evaluating the entire sequence of steps, trace-level evaluations help identify whether issues arise from specific components or the overall process. This approach is particularly valuable for multi-step workflows or multi-agent systems, ensuring end-to-end reliability and relevance.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 8 3,922 600 189 -6%
AI Guardrails 1 375 104 49 +60%
Multi-agent systems 1 239 80 45 -38%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.