Home / Companies / Weave / Blog / August 2026

August 2026 Summaries

1 posts from Weave

Filter
Month: Year:
Post Summaries Back to Blog
Weave optimizes the routing of requests to machine learning models by selecting the most cost-effective model that can handle each specific task, thereby substantially reducing costs while maintaining performance. By routing entire sessions rather than individual requests, Weave preserves the working cache, ensuring that cost savings are real rather than theoretical. This approach results in significant cost reductions, such as an 88.5% cut on median requests and a 68.7% cut across all routed-down requests, while maintaining reliability with a 98.8% success rate for failover requests. Unlike naive routing practices that switch models to save money but end up increasing costs due to cache loss, Weave's session-based routing ensures that these savings are sustainable in production environments. The system is designed to quickly adapt to new model releases without requiring code changes, allowing users to benefit from improved models and cost savings seamlessly.
Aug 04, 2026 664 words in the original blog post.