Currently listed through these providers:
Model details
BGE Multilingual Gemma2
BGE Multilingual Gemma2 is a text-embedding model from BAAI that adapts Google's Gemma 2 9B into a multilingual vector encoder. According to the upstream model card, it is described as an LLM-based multilingual embedding model trained on top of google/gemma-2-9b, with training data that spans a broad set of languages including English, Chinese, Japanese, Korean, and French, and a mix of task types such as retrieval, classification, and clustering. The same card also notes that the underlying training corpus has been released openly on the Hugging Face Hub, and that the model is consumed through BAAI's open-source FlagEmbedding library rather than a custom endpoint.
The model is positioned for cross-lingual retrieval and general semantic search, reporting state-of-the-art results on multilingual benchmarks MIRACL, MTEB-pl, and MTEB-fr, alongside strong performance on MTEB, C-MTEB, and AIR-Bench. In practice, that combination makes it well suited to multilingual document and passage retrieval, semantic clustering across language boundaries, and classification pipelines where queries and documents may arrive in different languages. Teams that need an embedding model grounded in a large open base and exposed through an open-weight, community-supported toolkit will find it a natural fit for multilingual RAG and search workloads.
Quick Info
Powered by- Provider
- Infomaniak
- Model key
- bge_multilingual_gemma2
- Release date
- Jul 25, 2024
- Last updated
- Aug 1, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Input tokens
- 8,000 tokens
- Output tokens
- 3,584 tokens
- Context window
- 8,000 tokens
Latest news about BGE Multilingual Gemma2
No articles yet. Fetch the latest news to show it here.