Fine-tuning open LLM judges to outperform GPT-5.2
Blog post from Together AI
The text discusses the use of preference optimization to train open-source language models (LLMs) that outperform GPT-5.2 in aligning with human preferences, using the Reward Bench 2 benchmark. The study highlights that models like GPT-OSS 120B and Qwen3 235B can be fine-tuned to match or surpass GPT-5.2 in human preference alignment through Direct Preference Optimization (DPO), a method that optimizes models based on preference pairs. The concept of LLM-as-a-judge is explored, where LLMs are used to evaluate other LLM outputs by focusing on simpler classification tasks, such as determining which response is better or if a text contains harmful content. The experiment reveals that while Qwen3 235B outperforms GPT-5.2 without tuning, GPT-OSS 120B shows significant improvement post-fine-tuning, particularly in math and subjective response quality. The analysis underscores the cost-effectiveness and flexibility of open-source models, which offer transparency and lower costs compared to closed-source alternatives like GPT-5.2, making them a promising option for production evaluation systems.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 20 | 5,987 | 964 | 233 | +29% |
| AI Model Fine-tuning | 16 | 1,108 | 170 | 74 | +87% |
| Reinforcement learning | 2 | 136 | 62 | 39 | -12% |
| AI Guardrails | 1 | 449 | 167 | 60 | +25% |
| RAG | 1 | 1,791 | 278 | 92 | +70% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.