Home / Companies / Inngest / Blog / Post Details
Content Deep Dive

Evals in Inngest: A Practical Guide to Measuring and Improving Your Agent

Blog post from Inngest

Post Details
Company
Date Published
Author
Lauren Craigie
Word Count
1,859
Company Posts That Month
6
Language
-
Hacker News Points
-
Post removed?
No
Summary

Inngest’s approach to AI agent evaluations frames evals as individual, targeted measurements of behavior or outcomes, such as user satisfaction, instruction following, tool use, structured-output validity, cost, or completion of an intended action. The article recommends starting with a narrow metric, often using live user feedback such as thumbs-up or thumbs-down scores, and using deferred scoring to connect later events like approvals, ticket closures, or repeat questions back to the original agent run. Inngest supports online evaluations on production traffic and offline evaluations against fixed datasets, with online methods positioned as an accessible starting point and offline testing becoming more valuable for regression checks as systems mature. Scores can be viewed alongside execution data and AI metadata in the Inngest dashboard, while experiments can split traffic between models, prompts, retrieval configurations, or other changes to identify which option produces better outcomes. The article also distinguishes operational observability, including latency, token usage, and model activity, from evaluations of whether an agent’s results were effective, arguing that both are necessary for improvement.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Observability 5 3,175 737 186 -24%
AI Agents 3 5,780 1,243 245 -15%
LLM 1 5,068 1,020 229 -34%
OpenTelemetry 1 757 153 55 -30%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.