Model Library/Qwen3 Embedding

Qwen3 Embedding

LLM
Embedding
Multilingual

Multilingual text-embedding family (0.6B-8B) serving OpenAI-compatible /v1/embeddings via vLLM

On-Demand Dedicated 1xRTX 5090

Details

Modalities

text

Recommended Hardware

1xRTX 5090

Estimated Price

Loading...

Provider

Alibaba

Family

Qwen3

Parameters

8B

Context

32768 tokens

License

apache-2.0

Qwen3 Embedding

Qwen3 Embedding is the Qwen family's text-embedding model series, purpose-built for embedding and ranking tasks such as text retrieval, code retrieval, semantic search, text classification, clustering, and bitext mining. Built on the dense Qwen3 foundation models, the series inherits their multilingual capability, long-text understanding, and reasoning skill, and ships in three sizes — 0.6B with 1024-dimension embeddings, 4B with 2560 dimensions, and 8B with 4096 dimensions — all with a 32K-token context window.

On Vast, the flagship template serves Qwen3-Embedding-8B through vLLM's pooling runner, exposing an OpenAI-compatible embeddings endpoint at /v1/embeddings. The 4B and 0.6B variants trade a few points of quality for lower cost and latency on smaller GPUs.

Key Features

  • State-of-the-art retrieval quality. The 8B embedding model ranked No.1 on the MTEB multilingual leaderboard as of June 5, 2025, with a mean task score of 70.58; the 4B model scores 69.45 and the 0.6B model 64.33 on the same benchmark.
  • Flexible embedding dimensions (MRL). All sizes support Matryoshka Representation Learning: user-defined output dimensions from 32 up to each model's native width (4096 for 8B, 2560 for 4B, 1024 for 0.6B), so you can shrink vectors to fit your index without retraining.
  • Instruction aware. Queries can carry a one-sentence task instruction (for example, "Given a web search query, retrieve relevant passages that answer the query"). Qwen's evaluations show instructions typically improve downstream task quality by 1 to 5 percent; they recommend writing instructions in English.
  • Multilingual and code-capable. The series supports over 100 languages, including programming languages, with strong multilingual, cross-lingual, and code-retrieval performance.
  • Long inputs. Up to 32K tokens per input, suitable for embedding whole documents rather than fragments.

Use Cases

  • Retrieval-augmented generation (RAG) pipelines and vector databases
  • Semantic search across multilingual corpora
  • Code search and cross-lingual retrieval
  • Text classification, clustering, and deduplication
  • Reranking pipelines, paired with the companion Qwen3-Reranker models

Architecture

Each size is a dense Qwen3 causal-language-model backbone finetuned for embedding: the 0.6B model has 28 layers, and the 4B and 8B models have 36 layers. Embeddings are produced by pooling over the final hidden states, then normalized. Serving through vLLM's pooling runner converts the causal-LM checkpoint into an embedding model at load time.

Benchmarks

On the MTEB multilingual benchmark (evaluated May 24, 2025), Qwen3-Embedding-8B leads with a mean task score of 70.58, ahead of gemini-embedding-exp-03-07 at 68.37, and well clear of text-embedding-3-large at 58.93 and Cohere-embed-multilingual-v3.0 at 61.12. Qwen3-Embedding-4B reaches 69.45 — also ahead of every compared non-Qwen model — and even the 0.6B model's 64.33 outscores text-embedding-3-large and Cohere's multilingual embedder. On MTEB English v2 the 8B model scores 75.22 mean task, with the 4B at 74.60 and the 0.6B at 70.70. Full results are on the Qwen3-Embedding-8B model card.

Quick Start Guide

Choose a model and click 'Deploy' above to find available GPUs recommended for this model.

Rent your dedicated instance preconfigured with the model you've selected.

Start sending requests to your model instance and getting responses right now.