Evals in CI: How to write your LLM evals as tests with Arize Phoenix
Blog post from Arize
Arize Phoenix offers a structured approach to integrating evaluations (evals) as tests within continuous integration (CI) frameworks like Pytest and Vitest/Jest, especially for applications involving large language models (LLMs). The primary challenge addressed is the non-deterministic nature of LLMs, which necessitates the use of repetitive evaluations to ensure reliability. Evals differ from traditional tests due to the inherent unpredictability and additional complexities such as cost, latency, and the need for qualitative judgment often requiring another LLM. Phoenix provides tools to write evals as standard tests, allowing developers to track performance metrics and debug applications effectively. The process involves defining scenarios, the system under test, and checks on outputs, distinguishing between hard invariants that fail CI tests and quality signals that are monitored over time. Phoenix facilitates the organization and analysis of test data, allowing teams to maintain a source of truth in their test files while providing infrastructure to log and compare results as applications evolve.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 20 | 6,942 | 1,215 | 234 | +11% |
| Developer Experience | 1 | 511 | 247 | 89 | +26% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.