Light
Home
/
Companies
/
Surge AI
/
Hacker News
Surge AI on HN
50 posts with 1+ points since 2022
Filters
Min points:
1
10
25
50
100
250
500
Since:
2021
2022
2023
2024
2025
2026
Posts by Month (50 total)
Hacker News Posts
Search:
Title
Points
Comments
Date
Is Google Search Deteriorating? Measuring Google's Search Quality in 2022
470
414
2022-01-11
30% of Google's Emotions Dataset Is Mislabeled
334
144
2022-07-14
Evaluation of TikTok vs. Instagram Reels
222
263
2022-09-02
Building a no-code toxicity classifier by talking to GitHub Copilot
212
143
2022-03-25
Are popular toxicity models simply profanity detectors?
183
211
2022-01-25
Generating Children’s Stories Using GPT-3 and DALL·E
138
145
2022-06-29
We asked 100 humans to draw the DALL·E prompts
138
73
2022-05-13
HellaSwag: 36% of this popular large language model benchmark contains errors
49
8
2022-12-06
I wanted burritos. Facebook Search sent me to a dead restaurant 45m …
25
33
2022-06-16
SWE-Bench Failures: When Coding Agents Spiral into 693 Lines of Hallucinations
22
1
2025-09-18
Google Search Is Falling Behind
16
5
2022-04-13
Twitter’s Egregious Content Moderation Failures
15
0
2022-11-10
Move Over, Google: The TikTokification of Next-Gen Search
13
4
2022-10-26
The average number of ads on a Google Search recipe? 8.7
13
3
2022-04-29
DALL·E vs. Imagen, and Evaluating Astral Codex Ten's Bet on AI Progress
13
0
2022-09-30
What if social media optimized for human values? A Facebook case study
12
1
2022-02-11
Explaining Reinforcement Learning with Human Feedback (RLHF)
11
0
2023-01-05
The $250K Inverse Scaling Prize and Human-AI Alignment
11
0
2022-09-28
An Analysis of Omicron Tweets: 30% Are Skeptical of the Medical Establishment
10
2
2022-01-21
How Good is Hugging Face's BLOOM? Human Evaluation of Large Language Models
10
0
2022-07-21
AI Red Teams for Adversarial Training: Making ChatGPT and LLMs More Robust
9
0
2022-12-13
Writing a Super Bowl Worthy Commercial with GPT-3
9
0
2022-02-16
Inter-Annotator Agreement: An Introduction to Krippendorff’s Alpha
9
0
2022-01-06
Is Elon right? We labeled 500 Twitter users to measure the amount …
7
5
2022-05-20
Optimizing Facebook's Algorithms for Human Values Instead of Clicks
7
1
2022-07-29
Extracting text from a pdf broke ChatGPT
7
2
2025-09-16
LMArena Is a Cancer on AI
6
1
2026-01-06
Building Better Developer Search: How Neeva Measures Search Quality
5
0
2022-07-07
Wall Street Experts Tested GPT-5 and Claude. Both Struggled – Even with …
5
1
2025-11-07
Google Search and Gmail are getting overrun by Spam
4
0
2022-06-02
How We Built It: OpenAI's GSM8K Dataset of 8,500 Math Problems
4
0
2022-06-15
Unsexy AI Failures: The PDF That Broke ChatGPT
4
0
2025-10-03
RL Environments and the Hierarchy of Agentic Capabilities
4
0
2025-11-11
ChatGPT vs. Google Search
3
0
2022-12-22
Humans vs. Gary Marcus: The Complexity of Measuring Machine Intelligence
3
0
2022-06-23
LMArena Is a Plague on AI
3
1
2025-12-08
How TikTok Is Evolving the Next Generation of Search
2
1
2022-11-01
Sentiment Analysis Dataset of Social Media Stock Conversations
2
0
2022-06-10
Unsexy AI Failures: Still Confidently Hallucinating Image Text
2
0
2025-09-23
GDP.pdf: Can Frontier Models Master the Documents That Run the World?
2
None
2026-07-11
Riemann-Bench
2
None
2026-06-10
Is Sonnet 4.5 the best coding model in the world?
2
None
2025-10-15
The Obscenity List
1
None
2022-01-18
SurgeAI Blog: Human Evals vs. Academic Benchmarks
1
None
2025-09-04
Evolving Instruction Following Beyond IFEval and "Avoid the Letter C"
1
None
2026-01-24
EnterpriseBench: CoreCraft – Measuring AI Agents in Chaotic RL Environments
1
None
2026-03-07
Hemingway bench AI writing leaderboard
1
None
2026-02-04
A Product Take on Sonnet 4.5
1
None
2025-10-14
The Human/AI Frontier: A Conversation with Bogdan Grechuk
1
None
2025-09-30
AI agents still can't solve 1/3 of SWE-Bench problems. Why not? (A …
1
None
2025-09-22