First look: GPT 6 Astra is at the frontier of complex enterprise work
Blog post from Box
Box reports that OpenAI’s GPT-6 Astra achieved 77% accuracy on its benchmark of complex, multi-document business workflows, compared with 74% for GPT-5.6 Sol. Tested through an agent harness that requires models to locate and reconcile information across spreadsheets, PDFs, presentations, and images, GPT-6 Astra performed especially well on tasks involving conflicting sources, derived calculations, and domain-specific evidence. Examples included correctly applying tax-incentive adjustments in media profitability rankings, identifying proxy metrics and inconsistent growth claims in technology planning, citing the relevant contracting-policy provisions in legal NDA review, and distinguishing anomalous readings from missing records in energy consumption reporting. Box says these capabilities could reduce review time and decision risk by producing more traceable, decision-ready outputs, and plans to make GPT-6 Astra available in Box AI Studio.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| AI Agents | 1 | 931 | 231 | 103 | -84% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.