Home / Companies / Confident AI / Hacker News

Confident AI on HN

29 posts with 1+ points since 2022

Filters
Since:
Posts by Month (29 total)
Hacker News Posts
Title Points Comments Date
Unit Test LlamaIndex with DeepEval 35 3 2023-08-28
Tackling the Weaknesses of BertScore 9 1 2023-08-16
Best Practices for Unit Testing RAG Systems in Prod 4 0 2024-02-06
YC helped us raise our seed round in 5 days 4 0 2025-03-20
How to evaluate multi-turn LLM chatbots 3 0 2024-10-08
I used QAG to implement an LLM text summarization evals 3 0 2023-12-19
We Replaced Pinecone with PGVector 3 1 2023-11-01
DeepEval GuardRails – AI Alignment 2 0 2023-09-30
AI Agent Evaluation: The Definitive Guide to Testing AI Agents 2 0 2025-10-16
How to build your own LLM evaluation framework 2 0 2024-04-15
PDB Support for DeepEval 1 0 2023-09-07
Test for Bias After Finetuning LLMs 1 0 2023-09-02
How to measure ranking similarity for RAG systems 1 0 2023-08-21
Be confident about your LLM stack 1 0 2023-08-15
Testing for Factual Consistency in LLMs 1 0 2023-08-21
How to unit test for LLM answer relevancy 1 0 2023-08-20
Testing for Image Similarity with DeepEval 1 0 2023-10-02
How to create synthetic data to evaluate your LangChain pipelines 1 0 2023-08-16
Framework for Evaluating Rag 1 0 2023-09-11
Test LLMs for Toxicness 1 0 2023-09-02
We wrote a comprehensive guide on LLM security 1 0 2024-08-20
Overview of All Major LLM Benchmarks 1 0 2024-03-22
Best practices I learnt from helping health tech enterprise test LLMs 1 0 2024-02-27
The Complete LLM Evaluation Playbook: How To Run LLM Evals That Matter 1 0 2025-06-16
How to generate synthetic data using SOTA data evolution methods 1 0 2024-05-21
Evaluating LLMs for Lawyers 1 0 2023-09-25
LLM Evaluation Metrics: The Ultimate LLM Evaluation Guide 1 0 2024-10-02
What Is RAG? (With Examples) 1 0 2023-12-01
How to Evaluate LangChain QA Retrieval 1 0 2023-09-23