Home / Companies / MongoDB / Blog / Post Details
Content Deep Dive

Building AI with MongoDB: How Patronus Automates LLM Evaluation to Boost Confidence in GenAI

Blog post from MongoDB

Post Details
Company
Date Published
Author
Mat Keep
Word Count
566
Company Posts That Month
50
Language
English
Hacker News Points
-
Post removed?
No
Summary

Patronus AI is an automated evaluation platform for large language models (LLMs) that enables engineers to score and benchmark LLM performance on real-world scenarios, generate adversarial test cases, monitor hallucinations, and detect sensitive information. The company has partnered with MongoDB Atlas to provide managed evaluation services, test suites, and adversarial data sets, helping customers verify the reliability of their RAG systems built on top of MongoDB Atlas. Patronus AI's platform has made a startling discovery that widely used state-of-the-art LLMs frequently hallucinate, incorrectly answering or refusing to answer up to 81% of financial analysts' questions. The company provides a 10-minute guide to help developers evaluate and improve the performance of their RAG systems, including exploring different indexes, modifying document chunking sizes, re-engineering prompts, and fine-tuning the embedding model itself.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 7 2,401 292 122 -7%
RAG 7 1,125 154 56 -17%
Vector Search 4 2,087 216 81 +23%
AI Guardrails 3 94 42 25 +29%
AI Model Fine-tuning 1 474 91 59 +12%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.