Using TypeSafe’s Jev for evals in Datadog Agent Observability
Blog post from Datadog
TypeSafe AI’s Jev is a decision-focused model that accepts structured state and typed questions to return probabilistic yes/no, categorical, and scored judgments without written explanations, positioning it as a lower-cost alternative to using text-generating models for evaluation tasks. The post demonstrates a shared five-question rubric for evaluating a fictional airline support agent’s replies against retrieved policy context, measuring grounding, whether the question was answered, handoff behavior, failure mode, and customer impact in one parallel request. It emphasizes retaining raw probabilities and confidence values rather than reducing results to binary labels, using uncertainty or near-tied categories to identify cases for human review, while keeping thresholds, arithmetic, dates, and composite business rules in application code. For Datadog Agent Observability, Jev can score live spans asynchronously through external evaluations joined by a domain tag and can also power offline experiments, where caching enables multiple evaluator metrics to reuse one Jev request per dataset row. The approach recommends pinning and recording model versions, filtering state to relevant evidence, calibrating thresholds against labeled data, and validating agreement and consistency with human reviewers before operational use.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Jev | 31 | No monthly metrics for this publish month. | |||
| Observability | 7 | 472 | 102 | 54 | -85% |
| LLM | 3 | 747 | 162 | 79 | -85% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.