How we got frontier-level capability without a frontier dependency
Blog post from Box
A security team describes building a proof-of-concept agentic dynamic testing engine to find authorization, tenant-isolation, and other business-logic flaws that signature-based scanners often miss and that AI has made cheaper for both defenders and attackers to investigate. The system uses a constrained, multi-model ensemble that reads source code, generates hypotheses, tests only non-production environments through a policy-gated and rate-limited execution layer, and limits writes to disposable canary objects, while a deterministic oracle independently reproduces exploits, repeats successful tests, rejects false positives, and supplies replayable remediation details. More than a dozen models from six vendors provide diverse perspectives, with frontier models contributing deeper multi-step reasoning and open-weight models delivering broad, low-cost coverage and serving as backstops for provider refusals, timeouts, and failures. The team found that benchmark rankings and pricing did not reliably predict vulnerability-finding performance, making measurement against an organization’s own systems essential. The proposed next step is integration into CI so testing becomes a recurring security control, with the central lesson that enforceable boundaries, grounding in source and documentation, and independent proof matter more than reliance on any single model.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.