June 2024 Summaries
2 posts from Voyage AI
Filter
Month:
Year:
Post Summaries
Back to Blog
Voyage-multilingual-2 is a newly released model optimized for multilingual retrieval and retrieval-augmented generation (RAG), outperforming alternatives such as OpenAI v3 large and Cohere multilingual v3 across major languages like French, German, Japanese, Spanish, and Korean, while maintaining strong performance in English. It excels with an average of 5.6% improvement over the second-best performing model and supports a large 32K context length, making it suitable for expertise-intensive domains including code, law, and finance. Evaluated using over 85 datasets covering 27 languages, voyage-multilingual-2 achieves superior retrieval accuracy, particularly in multilingual contexts, and is assessed using the normalized discounted cumulative gain (NDCG@10) metric. This model is designed to enhance Gen AI applications for global users and multilingual developers, providing a promising tool for those in need of advanced multilingual support.
Jun 10, 2024
414 words in the original blog post.
voyage-finance-2 is a newly launched finance domain-specific embedding model that excels in financial retrieval tasks, outperforming other models like OpenAI and Cohere by an average of 7% and 12%, respectively, across 11 finance retrieval datasets. This model features a 32K context length, significantly longer than its competitors, and is part of a broader portfolio that includes voyage-law-2 and voyage-code-2. It is designed to address challenging retrieval problems by focusing parameter capacity on specific domains, thereby enhancing performance in expertise-intensive areas. Evaluated on datasets like TAT-QA, FinanceBench, and others, voyage-finance-2 demonstrated superior performance using the normalized discounted cumulative gain (NDCG@10) metric, proving its efficacy in tasks involving financial news, public filings, and financial reports. This model promises to enhance Gen AI applications in the financial sector, offering users a tool optimized for high-quality retrieval and encouraging further domain-specific innovation.
Jun 03, 2024
653 words in the original blog post.