Home / Companies / New Relic / Blog / Post Details
Content Deep Dive

Introducing AI Evaluation: Close the Loop Across the Full AI Lifecycle

Blog post from New Relic

Post Details
Company
Date Published
Author
David Fabritius, Product Marketing Manager
Word Count
1,553
Company Posts That Month
13
Language
English
Hacker News Points
-
Post removed?
No
Summary

New Relic introduces AI Evaluation, a capability within its AI Observability platform designed to address semantic and safety failures in production generative AI systems that traditional monitoring may not detect. Using asynchronous “LLM-as-a-judge” models, it evaluates sampled prompts and completions for criteria such as accuracy, relevance, hallucinations, prompt injection, PII leakage, toxicity, bias, and RAG faithfulness, assigning normalized scores, pass/fail results, and explanatory reasoning. These results are attached directly to distributed traces, allowing teams to connect low-quality AI outputs with specific prompts, retrieval calls, backend services, infrastructure behavior, and token costs. The service also supports configurable sampling to limit evaluation expense while maintaining oversight, and is positioned for AI/ML engineers, platform and operations teams, and security and compliance personnel seeking to troubleshoot AI failures, compare model cost versus quality, and apply production guardrails without separate monitoring silos.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
AI Guardrails 14 No monthly metrics for this publish month.
LLM 9 No monthly metrics for this publish month.
Observability 6 No monthly metrics for this publish month.
Real-time 6 No monthly metrics for this publish month.
RAG 5 No monthly metrics for this publish month.
Vector Search 3 No monthly metrics for this publish month.
AI Agents 2 No monthly metrics for this publish month.
AI Model Fine-tuning 1 No monthly metrics for this publish month.
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.