Home / Companies / Fireworks AI / Blog / Post Details
Content Deep Dive

Using Model-as-a-Judge for Reward in Reinforcement Fine Tuning

Blog post from Fireworks AI

Post Details
Company
Date Published
Author
-
Word Count
765
Company Posts That Month
12
Language
English
Hacker News Points
-
Post removed?
No
Summary

Leveraging a large language model (LLM) as a judge can significantly enhance the performance of policy models in domains that are challenging to quantify, such as creative writing. Using the Fireworks Reinforcement Fine Tuning (RFT) API, the Qwen2.5 32B base model was fine-tuned to achieve a 93.8% win rate on creative writing tasks against its original version. This process involved using the open-source Qwen3 235B model as a judge, which was managed with the Fireworks Build SDK to automatically allocate optimal compute resources. The evaluation methodology employed pairwise comparisons of different rollouts of the same prompt to assign rewards, using a rule-based reward function to assess dimensions like style and coherence. The Arena Hard Auto dataset, which includes creative writing, mathematics, and software engineering tasks, served as the testing ground, and the RFT methodology also showed improvements in more objective domains like mathematics and programming. The study highlights the potential of using LLMs for nuanced evaluation in creative tasks, offering significant improvements over base models.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 6 4,152 612 181 +19%
AI Model Fine-tuning 4 657 141 57 +70%
Serverless 2 889 215 78 +28%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.