voyage-code-4: code retrieval built for coding agents
Blog post from Voyage AI
Voyage has introduced voyage-code-4, a code embedding model designed for coding agents that need accurate, low-latency semantic retrieval when investigating broad bug reports or repository-wide behaviors. The model was trained on a new corpus derived from issue-fixing pull requests, pairing natural-language problem descriptions with code changed in completed fixes to better identify code relevant to symptoms rather than named identifiers. On a 19-benchmark agentic retrieval suite, Voyage reports that voyage-code-4 exceeds voyage-code-3, Cohere Embed v4, Gemini Embedding 2, and OpenAI v3 large by 27.54%, 28.25%, 31.03%, and 48.58%, respectively, using NDCG@10; it also reports improvements across 28 traditional code-retrieval datasets. It supports embedding sizes from 256 to 2,048 dimensions through Matryoshka learning and float32, int8, and binary quantization options, while costing $0.12 per million tokens, one-third less than voyage-code-3. The model is available through the Voyage API and MongoDB Atlas Embedding and Reranking API, with the first 200 million tokens offered free.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Vector Search | 13 | 2,358 | 371 | 127 | +5% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.