Beyond models: How context and evals make agents work in production
Blog post from Arize
Building AI agents for production environments poses significant challenges, as the transition from controlled demo settings to real-world systems often exposes gaps in context and evaluation. Tobias Leong, CTO of Axium Industries, highlights that the critical issue is not the AI model itself but the surrounding infrastructure, which includes understanding context, system design, and evaluation. Successful deployment requires a deep understanding of the environment, such as supply chain operations, and not just relying on model upgrades. This has led to the emergence of the role of agent engineers, who integrate software engineering with applied AI, emphasizing the importance of structured data, retrieval pipelines, and domain-specific logic. Leong stresses the need for robust evaluation frameworks to ensure consistent performance, suggesting that building internal tools can be resource-intensive, thus recommending platforms like Arize for monitoring and evaluation. The industry focus should shift from upgrading models to enhancing the systems around them to effectively utilize AI intelligence.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Observability | 7 | 4,496 | 812 | 176 | +40% |
| AI Agents | 2 | 4,430 | 1,100 | 236 | -3% |
| LLM | 2 | 5,932 | 1,046 | 223 | -2% |
| Harness engineering | 1 | 164 | 111 | 62 | +6% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.