Home / Companies / Surge AI / Hacker News

Surge AI on HN

50 posts with 1+ points since 2022

Filters
Since:
Posts by Month (50 total)
Hacker News Posts
Title Points Comments Date
Is Google Search Deteriorating? Measuring Google's Search Quality in 2022 470 414 2022-01-11
30% of Google's Emotions Dataset Is Mislabeled 334 144 2022-07-14
Evaluation of TikTok vs. Instagram Reels 222 263 2022-09-02
Building a no-code toxicity classifier by talking to GitHub Copilot 212 143 2022-03-25
Are popular toxicity models simply profanity detectors? 183 211 2022-01-25
Generating Children’s Stories Using GPT-3 and DALL·E 138 145 2022-06-29
We asked 100 humans to draw the DALL·E prompts 138 73 2022-05-13
HellaSwag: 36% of this popular large language model benchmark contains errors 49 8 2022-12-06
I wanted burritos. Facebook Search sent me to a dead restaurant 45m … 25 33 2022-06-16
SWE-Bench Failures: When Coding Agents Spiral into 693 Lines of Hallucinations 22 1 2025-09-18
Google Search Is Falling Behind 16 5 2022-04-13
Twitter’s Egregious Content Moderation Failures 15 0 2022-11-10
Move Over, Google: The TikTokification of Next-Gen Search 13 4 2022-10-26
The average number of ads on a Google Search recipe? 8.7 13 3 2022-04-29
DALL·E vs. Imagen, and Evaluating Astral Codex Ten's Bet on AI Progress 13 0 2022-09-30
What if social media optimized for human values? A Facebook case study 12 1 2022-02-11
Explaining Reinforcement Learning with Human Feedback (RLHF) 11 0 2023-01-05
The $250K Inverse Scaling Prize and Human-AI Alignment 11 0 2022-09-28
An Analysis of Omicron Tweets: 30% Are Skeptical of the Medical Establishment 10 2 2022-01-21
How Good is Hugging Face's BLOOM? Human Evaluation of Large Language Models 10 0 2022-07-21
AI Red Teams for Adversarial Training: Making ChatGPT and LLMs More Robust 9 0 2022-12-13
Writing a Super Bowl Worthy Commercial with GPT-3 9 0 2022-02-16
Inter-Annotator Agreement: An Introduction to Krippendorff’s Alpha 9 0 2022-01-06
Is Elon right? We labeled 500 Twitter users to measure the amount … 7 5 2022-05-20
Optimizing Facebook's Algorithms for Human Values Instead of Clicks 7 1 2022-07-29
Extracting text from a pdf broke ChatGPT 7 2 2025-09-16
LMArena Is a Cancer on AI 6 1 2026-01-06
Building Better Developer Search: How Neeva Measures Search Quality 5 0 2022-07-07
Wall Street Experts Tested GPT-5 and Claude. Both Struggled – Even with … 5 1 2025-11-07
Google Search and Gmail are getting overrun by Spam 4 0 2022-06-02
How We Built It: OpenAI's GSM8K Dataset of 8,500 Math Problems 4 0 2022-06-15
Unsexy AI Failures: The PDF That Broke ChatGPT 4 0 2025-10-03
RL Environments and the Hierarchy of Agentic Capabilities 4 0 2025-11-11
ChatGPT vs. Google Search 3 0 2022-12-22
Humans vs. Gary Marcus: The Complexity of Measuring Machine Intelligence 3 0 2022-06-23
LMArena Is a Plague on AI 3 1 2025-12-08
How TikTok Is Evolving the Next Generation of Search 2 1 2022-11-01
Sentiment Analysis Dataset of Social Media Stock Conversations 2 0 2022-06-10
Unsexy AI Failures: Still Confidently Hallucinating Image Text 2 0 2025-09-23
GDP.pdf: Can Frontier Models Master the Documents That Run the World? 2 None 2026-07-11
Riemann-Bench 2 None 2026-06-10
Is Sonnet 4.5 the best coding model in the world? 2 None 2025-10-15
The Obscenity List 1 None 2022-01-18
SurgeAI Blog: Human Evals vs. Academic Benchmarks 1 None 2025-09-04
Evolving Instruction Following Beyond IFEval and "Avoid the Letter C" 1 None 2026-01-24
EnterpriseBench: CoreCraft – Measuring AI Agents in Chaotic RL Environments 1 None 2026-03-07
Hemingway bench AI writing leaderboard 1 None 2026-02-04
A Product Take on Sonnet 4.5 1 None 2025-10-14
The Human/AI Frontier: A Conversation with Bogdan Grechuk 1 None 2025-09-30
AI agents still can't solve 1/3 of SWE-Bench problems. Why not? (A … 1 None 2025-09-22