monday Service + LangSmith: Building a Code-First Evaluation Strategy from Day 1
Blog post from LangChain
Monday.com has developed an evals-driven development framework for its AI Native Enterprise Service Management platform, designed to automate and resolve inquiries across various service departments. By integrating evaluation as a core component from the outset, the company has significantly accelerated feedback loops, enabling comprehensive testing across numerous examples in minutes rather than hours. The framework employs a dual-layered evaluation approach: offline evaluations act as a safety net, using curated datasets to ensure core logic and specific edge cases are robust, while online evaluations provide continuous quality monitoring in real-time production environments. The evaluations are managed as version-controlled code, using GitOps-style CI/CD deployment, which enhances agent observability and ensures high-quality AI interactions. The platform's architecture allows for parallel and concurrent testing, drastically improving evaluation speed and efficiency. This structured approach to evaluations, with tools like LangSmith and Vitest, reflects a commitment to rigorous testing standards akin to production code management, ensuring the AI service workforce remains reliable and adaptable to various enterprise service management use cases.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 7 | 5,138 | 781 | 181 | +34% |
| Real-time | 3 | 5,046 | 1,089 | 214 | +11% |
| MCP | 2 | 3,346 | 363 | 139 | +19% |
| AI Agents | 1 | 3,583 | 743 | 199 | -1% |
| Observability | 1 | 2,816 | 550 | 145 | +34% |
| Platform Engineering | 1 | 368 | 138 | 58 | +24% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.