Home / Companies / Riza / Blog / Post Details
Content Deep Dive

What GPT-4o Can't Code

Blog post from Riza

Post Details
Company
Date Published
Author
Kyle Gray
Word Count
1,636
Company Posts That Month
2
Language
English
Hacker News Points
-
Post removed?
No
Summary

Riza, a company focused on safely running untrusted code, explored the performance of OpenAI's GPT-4o model on the HumanEval benchmark, which consists of 164 Python programming problems. The model achieved a high Pass@10 rate of 97.0%, indicating it solved 159 problems within ten attempts, and a Pass@1 rate between 88% and 94% for individual samples. However, five problems consistently failed across ten attempts, highlighting areas where the model struggled with logic errors and misinterpretations of problem requirements. Specific cases are analyzed to understand these failures, such as incorrect handling of sentence delimiters, mismanagement of nested structures, and improper ordering based on digit sums. The examination aims to identify and rectify the errors, with the potential for improvements by feeding failure contexts back into the model for future iterations.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 3 3,629 397 137 -13%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.