cheon-embedding-0.6b-v1

A multilingual text embedding model for search and RAG with about 0.6B parameters. It turns queries and documents into 1024-dimensional vectors, reads up to 32,768 tokens per text and loads with AutoModel.

Highlights

  • Top-3 class on the global multilingual benchmark at 0.6B. On the same 128 MMTEB (Multilingual, v2) tasks it averages 71.73. Only two much larger models score higher (27B and 12B). It is ahead of Qwen3-Embedding-8B, a model about 13 times its size, and of Gemini Embedding.
  • Best among models of 8B or fewer parameters on these tasks: +1.76 over its 0.6B base model and +6.49 over Qwen3-Embedding-0.6B.
  • Strong where multilingual search needs it: it beats the 8B model on bitext mining, classification, multilabel classification and retrieval.
  • Long inputs and instructions: up to 32,768 tokens per text, with a one-line task instruction for queries.
  • Drop-in loading: AutoModel with trust_remote_code=True, float32 weights, 1024-dimensional unit vectors (cosine similarity = dot product).

Global benchmark: MMTEB (Multilingual, v2)

Average main score over the same 128 of the 131 tasks. This model was scored with the official mteb harness (float32, up to 32,768 tokens); the other scores are from the official MTEB results repository. The remaining three tasks are being scored.

Rank Model Parameters Average (128 tasks)
1 microsoft/harrier-oss-v1-27b 27B 75.29
2 tencent/KaLM-Embedding-Gemma3-12B-2511 12B 73.32
3 cheon-embedding-0.6b-v1 0.6B 71.73
4 Qwen/Qwen3-Embedding-8B 8B 71.56
5 Bytedance/Seed1.6-embedding-1215 undisclosed 71.23
6 Qwen/Qwen3-Embedding-4B 4B 70.40
7 nvidia/llama-embed-nemotron-8b 8B 70.38
8 microsoft/harrier-oss-v1-0.6b (base) 0.6B 69.97
10 google/gemini-embedding-001 undisclosed 69.29
21 Qwen/Qwen3-Embedding-0.6B 0.6B 65.24

Ranks are among all models in the results repository that report all 128 tasks.

By task type (same 128 tasks):

Task type Tasks cheon-embedding-0.6b-v1 harrier-oss-v1-0.6b (base) Qwen3-Embedding-8B
Bitext mining 13 83.36 82.85 80.89
Classification 43 76.25 73.88 74.00
Multilabel classification 4 43.92 31.61 34.63
Retrieval 17 72.20 71.01 70.89
Clustering 16 55.13 54.00 57.65
Pair classification 11 83.35 82.07 86.40
Reranking 5 74.20 73.25 76.40
Semantic similarity (STS) 16 77.71 77.09 81.08
Instruction reranking 3 0.86 0.81 10.06

Usage

import torch
from transformers import AutoModel, AutoTokenizer

repo = "cheonai/cheon-embedding-0.6b-v1"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModel.from_pretrained(repo, trust_remote_code=True, torch_dtype=torch.float32)
model.eval()

queries = model.encode(
    ["How do I renew my passport online?"],
    tok,
    instruction="Given a web search query, retrieve relevant passages that answer the query",
)
documents = model.encode(["Passports can be renewed online through the government portal."], tok)
scores = queries @ documents.T  # cosine similarity: the vectors are unit length

model(**tok(texts, padding=True, truncation=True, return_tensors="pt")).pooler_output gives the same vectors for texts that already carry their prefix.

Instructions

Queries take a one-line task description; documents take no prefix.

Instruct: {task description}
Query:{query text}

There is no space after Query:. encode(..., instruction=...) builds this prefix. Examples of task descriptions:

Use Task description
Web search Given a web search query, retrieve relevant passages that answer the query
Semantic similarity Retrieve semantically similar text
Parallel sentences Retrieve parallel sentences
Classification Classify the sentiment of a given review

Specifications

Parameters 596.0M
Output 1024-dimensional unit vectors
Maximum sequence length 32,768 tokens
Weights float32 (model.safetensors, 2.4 GB)
Base model microsoft/harrier-oss-v1-0.6b (MIT)

Requirements

transformers>=4.51 and torch.

License

CC-BY-NC-4.0 — research and evaluation only. Contact us for commercial use. The base model (microsoft/harrier-oss-v1-0.6b) is MIT, but this fine-tune (learned weights) is ours and ships under the license above.