April 2026 Summaries
2 posts from Inference
Filter
Month:
Year:
Post Summaries
Back to Blog
Schematron V2 is the latest iteration of specialized HTML-to-JSON extraction models, offering enhanced performance with two new variants, Schematron V2 Small and Schematron V2 Turbo, which significantly improve upon the previous generation's speed and quality. Optimized for cost and latency, these models maintain high extraction quality while providing faster processing capabilities, making them suitable for large-scale data extraction tasks. Schematron V2 Small nearly matches the quality of the original 8B model with faster performance, while Schematron V2 Turbo focuses on maximizing throughput, achieving 4.14 requests per second, which is 2.5 times faster than its predecessor. Both models are available through a serverless API on Inference.net, with pricing designed to make web-scale extraction more accessible. Future developments include Schematron Pro, which aims to offer even higher accuracy without sacrificing throughput.
Apr 16, 2026
1,041 words in the original blog post.
Catalyst is a newly launched platform designed to optimize production AI applications by integrating monitoring, evaluations, training, and deployment into a single system. It simplifies the process by using actual production data as a training environment, eliminating the need for synthetic environments and reducing costs significantly. Catalyst is compatible with existing OpenAI and Anthropic providers and can be integrated into a project with minimal code adjustments. The platform constructs training and evaluation datasets from real traffic, allowing for a continuous cycle of model improvement through its self-improvement flywheel, which manages data ingestion, training, and deployment. Catalyst's approach addresses the challenges of traditional reinforcement learning environments by providing a more efficient and direct method for improving production AI models. Currently available in public beta, Catalyst covers associated costs and offers users an opportunity to experience its full capabilities in optimizing AI systems using real-world data.
Apr 14, 2026
929 words in the original blog post.