Currently listed through:
Model details
voyage-code-2
Voyage Code 2 is a specialized text embedding model engineered to bridge the gap between natural language queries and programming logic. By converting code snippets, function names, and technical documentation into high-dimensional numeric vectors, the model enables semantic search capabilities that move beyond simple keyword matching. It is designed to handle diverse inputs, including Python, C++, and Java, as well as documentation for major machine learning frameworks like PyTorch and TensorFlow. This architecture allows developers to perform complex tasks such as finding similar code implementations or retrieving relevant documentation, making it a versatile tool for building code-aware retrieval-augmented generation systems.
The model was developed to address the specific challenges of technical retrieval, demonstrating significant performance gains over general-purpose alternatives across multiple coding datasets such as HumanEval and MBPP. By providing a unified embedding space for both natural language and code, it simplifies the development of search indices, allowing teams to manage documentation and source code within a single workflow. Its design supports a large context window, providing the flexibility needed to process substantial code blocks or long-form technical passages. As a specialized tool, it is well-suited for integration into vector databases, where it helps maintain high retrieval relevancy in large-scale, inconsistent, or complex software repositories.
Quick Info
Powered by- Provider
- Vercel AI Gateway
- Model key
- voyage/voyage-code-2
- Release date
- Jan 1, 2024
- Last updated
- Jan 1, 2024
- Input modalities
- Output modalities
- Capabilities
Limits
- Output tokens
- 1,536 tokens
- Context window
- 8,192 tokens