June 2024 Summaries
5 posts from Galileo
Filter
Month:
Year:
Post Summaries
Back to Blog
Multimodal models are increasingly being used across industries due to advancements in language and vision model capabilities. However, Large Language Models (LLMs) cannot understand visual information and Large Vision Models (LVMs) struggle with reasoning tasks. To address this, Multimodal Large Language Models (MLLMs) have been developed, combining the strengths of both LLMs and LVMs to handle multimodal information effectively. Despite their advanced capabilities, MLLMs are prone to hallucination, a phenomenon where they generate content that is not present or accurate based on the input data. Researchers have been actively investigating methods to detect and mitigate hallucinations in Multimodal Large Language Models (MLLMs) and Large Vision-Language Models (LVLMs). These models are driving innovation across various industries by facilitating the interpretation and generation of various modalities, such as text, images, audio, and video. To enhance the accuracy and reliability of these models, ongoing research and innovative approaches are essential.
Jun 25, 2024
3,391 words in the original blog post.
Evaluating the quality of outputs produced by large language models (LLMs) is increasingly challenging due to the complexity of generative tasks and the length of responses. The "vibe check" approach, which involves subjective human judgments through A/B testing with crowd workers, has limitations such as being expensive and time-consuming. Recent research highlights the need for more nuanced and objective approaches to assess model performance. Factors affecting LLM performance include confounders like assertiveness and complexity, subjectivity and bias in preference scores, coverage of crucial error criteria, and biases like authority, beauty, verbosity, positional, attention, sycophancy, nepotism, fallacy oversight, and others. Researchers have tried various approaches to develop reliable methods for evaluating LLM performance, including using LLM-derived metrics, prompting LLMs with designed prompts, fine-tuning LLMs with labeled evaluation data, and developing guidelines and scaling annotation processes. Galileo's ChainPoll technique combines Chain-of-Thought prompting with polling to ensure robust and nuanced assessment, while the Luna suite provides a comprehensive framework for evaluating LLM outputs in terms of accuracy, cost, speed, scalability, and overcoming issues of human vibe checks and LLM-as-a-Judge biases.
Jun 18, 2024
1,971 words in the original blog post.
Galileo Luna is a family of Evaluation Foundation Models (EFM) fine-tuned specifically for hallucination detection in RAG settings, outperforming GPT-3.5 and commercial evaluation frameworks while significantly reducing cost and latency, making it an ideal candidate for industry LLM applications. Luna excels on the RAGTruth dataset and shows excellent generalization capabilities across various industries and use cases, including finance, numerical reasoning, biomedical research, legal, and general knowledge. The model is optimized to process up to 16k input tokens in under one second on cost general-purpose GPUs, achieving a 97% reduction in cost and a 96% reduction in latency compared to GPT-3.5-based approaches. Luna's dynamic windowing technique ensures comprehensive validation and significantly improves hallucination detection accuracy, making it a highly efficient solution for industry applications.
Jun 11, 2024
1,065 words in the original blog post.
Galileo has launched GenAI Evaluation, a low-latency and cost-efficient method for evaluating generative AI models. This approach aims to reduce the reliance on human-in-the-loop evaluations and costly LLM-based evaluations, providing ultra-low-latency evaluations in milliseconds without compromising accuracy. The 5 breakthroughs of Galileo Luna include outperforming popular evaluation techniques, eliminating the need for ground truth test sets, reducing cost by up to $97% compared to GPT-3.5, achieving 18% higher accuracy than GPT-3.5 in detecting hallucinations, and enabling real-time evaluations with ultra-low latency of milliseconds. The Luna Evaluation Foundation Models are now available to all Galileo customers at no additional cost, powering various evaluation tasks such as hallucination detection, RAG analytics, security and privacy, and more.
Jun 06, 2024
1,117 words in the original blog post.
Effective evaluations are critical for developing and deploying enterprise GenAI, but many organizations still rely on outdated methods like 'vibe checks,' manual human evaluations, and asking LLMs. Evaluations that are too slow, costly, and inaccurate can hinder production needs. To productionize trustworthy AI, enterprise teams need to rethink GenAI evaluations. The goal is to gain practical strategies and insights into cutting-edge evaluation techniques, ensuring GenAI solutions are ready for production.
Jun 03, 2024
83 words in the original blog post.