Open-source agents with frontier advisors: matching frontier performance through training and harness engineering
Blog post from Fireworks AI
In a study examining the performance of legal AI models using Harvey's Legal Agent Benchmark (LAB), the integration of open-source models with frontier tools and Fireworks-native post-training techniques significantly improved performance and cost efficiency. The hybrid system, featuring an open-source GLM 5.1 worker and Claude Opus 4.7 as an advisor, achieved an 18/100 all-pass rate at a reduced cost of $368, outperforming Opus alone, which had a 14/100 rate at $954. Post-training on the Fireworks platform, utilizing supervised and reinforcement fine-tuning on models like Kimi K2.6, further enhanced performance, demonstrating the potential of open-source models to approach frontier-level quality while maintaining cost-effectiveness. This approach allowed for a seamless transition from research to production, with no discrepancies between training and serving models, emphasizing the competitive edge of open-source solutions in legal AI tasks.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| AI Model Fine-tuning | 9 | 739 | 196 | 71 | +20% |
| Harness engineering | 4 | 255 | 140 | 70 | +38% |
| Multi-agent systems | 2 | 538 | 169 | 80 | -1% |
| LLM | 1 | 6,237 | 1,165 | 246 | -31% |
| Serverless | 1 | 1,010 | 231 | 94 | -44% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.