Home / Companies / Voyage AI / Blog / Post Details
Content Deep Dive

voyage-code-4: code retrieval built for coding agents

Blog post from Voyage AI

Post Details
Company
Date Published
Author
Voyage AI
Word Count
940
Company Posts That Month
1
Language
English
Hacker News Points
4
Post removed?
No
Summary

Voyage has introduced voyage-code-4, a code embedding model designed for coding agents that need accurate, low-latency semantic retrieval when investigating broad bug reports or repository-wide behaviors. The model was trained on a new corpus derived from issue-fixing pull requests, pairing natural-language problem descriptions with code changed in completed fixes to better identify code relevant to symptoms rather than named identifiers. On a 19-benchmark agentic retrieval suite, Voyage reports that voyage-code-4 exceeds voyage-code-3, Cohere Embed v4, Gemini Embedding 2, and OpenAI v3 large by 27.54%, 28.25%, 31.03%, and 48.58%, respectively, using NDCG@10; it also reports improvements across 28 traditional code-retrieval datasets. It supports embedding sizes from 256 to 2,048 dimensions through Matryoshka learning and float32, int8, and binary quantization options, while costing $0.12 per million tokens, one-third less than voyage-code-3. The model is available through the Voyage API and MongoDB Atlas Embedding and Reranking API, with the first 200 million tokens offered free.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Vector Search 13 2,358 371 127 +5%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.