Open-source agents with frontier advisors: matching frontier performance through training and harness engineering
Blog post from Fireworks AI
In a study examining the performance of legal AI models using Harvey's Legal Agent Benchmark (LAB), the integration of open-source models with frontier tools and Fireworks-native post-training techniques significantly improved performance and cost efficiency. The hybrid system, featuring an open-source GLM 5.1 worker and Claude Opus 4.7 as an advisor, achieved an 18/100 all-pass rate at a reduced cost of $368, outperforming Opus alone, which had a 14/100 rate at $954. Post-training on the Fireworks platform, utilizing supervised and reinforcement fine-tuning on models like Kimi K2.6, further enhanced performance, demonstrating the potential of open-source models to approach frontier-level quality while maintaining cost-effectiveness. This approach allowed for a seamless transition from research to production, with no discrepancies between training and serving models, emphasizing the competitive edge of open-source solutions in legal AI tasks.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| AI Model Fine-tuning | 9 | 762 | 211 | 75 | +14% |
| Harness engineering | 4 | 254 | 141 | 71 | +28% |
| Multi-agent systems | 2 | 556 | 175 | 81 | -7% |
| LLM | 1 | 6,292 | 1,205 | 252 | -36% |
| Serverless | 1 | 1,019 | 237 | 96 | -45% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.