Home / Companies / TestMu AI / Blog / Post Details
Content Deep Dive

Eval-First QA Agents: Testing Streaming Platforms at Fox Networks [Testμ 2026]

Blog post from TestMu AI

Post Details
Company
Date Published
Author
TestMu AI
Word Count
2,926
Company Posts That Month
113
Language
English
Hacker News Points
-
Post removed?
No
Summary

At Testμ Conf 2026, Fox QA manager Gregory Goldshteyn described an “eval-first” approach to testing AI agents, in which safety assertions are written before capabilities are exposed and untested tools are treated as unavailable. Fox applies this method to in-house streaming QA agents and MCP tools used across devices such as Apple TV, Roku, Fire TV, and mobile platforms, reporting 88 assertions at a 100% internal pass rate across 38 gated MCP tools and helper APIs. Demonstrations showed that models can fail security rules through seemingly routine requests, including documentation pretexts, persona overrides, and a translation request that caused one model to reveal its entire system prompt despite explicit secrecy instructions. A model-based grader then incorrectly judged the leaked content as evidence of safe behavior, underscoring Goldshteyn’s recommendation to pair LLM-graded checks with deterministic structural assertions. The team uses Promptfoo with multiple models and layered checks ranging from simple string and regex tests to LLM rubrics and factuality evaluation, while emphasizing that larger models are not necessarily safer. Goldshteyn also argued that AI defect reports must record prompts, model settings, run counts, reproduction rates, and grading criteria because outputs are probabilistic, and that many significant failures occur in hidden routing, retrieval, permissions, and tool-call behavior rather than in the final response text. Fox reported finding leaks in more than 36% of MCP servers it tested, though without sample-size or scope details, and recommended placing red-team eval suites in CI/CD gates to detect regressions before deployment.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 11 747 162 79 -85%
MCP 8 2,241 148 72 -74%
Real-time 5 649 155 80 -85%
AI Guardrails 1 35 22 12 -94%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.