Home / Companies / OpenPipe / Blog / Post Details
Content Deep Dive

Using Reinforcement Learning and $4.80 of GPU Time to Find the Best HN Post Ever (RLHF Part 1)

Blog post from OpenPipe

Post Details
Company
Date Published
Author
Kyle Corbitt
Word Count
2,044
Company Posts That Month
3
Language
English
Hacker News Points
217
Post removed?
No
Summary

This post discusses using reinforcement learning and human feedback (RLHF) to improve the performance of a Large Language Model (LLM) on predicting the upvote count of Hacker News (HN) stories. The author, Kyle Corbitt, founder of OpenPipe, explains how they built a reward model that can predict the upvote count based on the story title, URL, date, and content. The model is trained using a dataset of 114K HN stories with their corresponding upvote counts, and the training process takes around 1.5 hours on an H100 GPU for $4.05. The model achieves a root mean-square error (RMSE) of 1.11, which translates to an accuracy of e^1.11 ≈ 3. The author then runs the model against the entire corpus of HN stories and finds that it consistently over-estimates the score at the low end and under-estimates it at the high end. Despite this, the model identifies some great HN stories and provides interesting insights into what makes a story successful on HN. The author concludes by saying that RLHF gives them a powerful set of techniques to improve post quality, which they will cover in the next post in the series.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Reinforcement learning 13 No monthly metrics for this publish month.
LLM 4 3,598 465 143 -7%
AI Model Fine-tuning 1 897 160 75 +43%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.