Box AI enterprise eval: OpenAI's o3 and o4-mini for data extraction with Box AI
Blog post from Box
OpenAI has released the o3 and o4-Mini reasoning models, which have demonstrated strong capabilities in accurate data extraction from complex enterprise documents, as evaluated using the Box AI Enterprise Eval framework on a subset of the CUAD dataset. The o4-Mini model achieved an 84% correctness rate with an F1 score of 0.85, showcasing its high accuracy and reliability in extracting critical information, making it ideal for tasks requiring precision, such as legal reviews and compliance checks. The o3 model also performed robustly with an 80% correctness rate and an F1 score of 0.81, providing a reliable baseline for a variety of enterprise tasks, especially in high-volume processing scenarios. These models are positioned as valuable tools across various sectors, including legal, finance, sales operations, procurement, and HR, offering organizations the flexibility to choose between peak accuracy or broader processing capabilities depending on their specific needs. OpenAI's integration of these models with Box AI presents powerful new options for tackling complex data extraction challenges, and interested parties can access them through Box AI Studio and APIs.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.