Home / Companies / Hugging Face / Blog / Post Details
Content Deep Dive

Direct Preference Optimization Beyond Chatbots

Blog post from Hugging Face

Post Details
Company
Date Published
Author
Erick Lachmann and Pimenta de Freitas Cardoso
Word Count
2,953
Company Posts That Month
94
Language
-
Hacker News Points
-
Post removed?
No
Summary

DharmaOCR, a specialized OCR model for Brazilian Portuguese text, has demonstrated significant improvements in reducing text degeneration rates through a method called Direct Preference Optimization (DPO). The approach leverages the model's own degenerate outputs as negative training signals, using them to create preference pairs that help the model learn to avoid failure modes. Unlike traditional supervised fine-tuning (SFT), which optimizes for correct predictions without explicitly penalizing degeneration, DPO targets this specific failure mode by training on the full output, rather than token-level predictions. The study showed that DPO reduced degeneration rates across various model families by an average of 59.4%, with some models experiencing reductions as high as 87.6%. This methodology suggests that structured generation tasks like OCR can benefit from using model failures as training signals, provided the failures are distinct, scoreable, and numerous, thereby enhancing model reliability without compromising extraction quality.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
AI Model Fine-tuning 6 762 211 75 +14%
LLM 5 6,292 1,205 252 -36%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.