Most rerankers score query-document pairs one at a time. This one instead encodes the query and all candidates together in a single context, then runs a small bidirectional refiner on top of the pooled candidate vectors so candidates can attend to each other before scoring. The refiner has no positional encoding, so shuffling the candidate order doesn't change the scores.
- Base: Qwen3-0.6B (Apache-2.0)
- Params: ~609M total (adapter already merged into the base weights)
- Refiner: 4 layers, 512 hidden dim, single attention-pooled slot per candidate
- Weights:
model.safetensors, fp32
Usage
Loads straight into AutoModel with trust_remote_code=True. The included
modeling_cheon_reranker.py is a from-scratch reimplementation of the forward
pass (no training code) — checked against the original checkpoint on random
inputs, max difference 2.4e-07, i.e. floating-point noise.
import torch
from transformers import AutoModel, AutoTokenizer
repo = "cheonai/cheon-reranker-0.6b-v1"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModel.from_pretrained(repo, trust_remote_code=True, torch_dtype=torch.float32)
model.eval()
# Query + candidates go into one sequence. Track each candidate's token span
# as you build it instead of hardcoding positions. Candidates don't need to
# share a language - the tokenizer and span logic don't care.
query = "query: 이 사건의 처리 기한은 언제까지인가요?\n"
candidates = [
"candidate (en): The appeal must be filed within 30 days of the decision.",
"candidate (ko): 이의신청은 처분을 안 날부터 30일 이내에 제기해야 합니다.",
"candidate (zh): 上诉必须在裁定后30天内提出。",
"candidate (ja): 不服申立ては、決定を知った日から30日以内に行う必要があります。",
]
ids = tok(query, add_special_tokens=False)["input_ids"]
doc_spans = []
for text in candidates:
piece = tok(text, add_special_tokens=False)["input_ids"]
start = len(ids)
ids += piece
doc_spans.append((start, len(ids)))
input_ids = torch.tensor([ids])
attention_mask = torch.ones_like(input_ids)
corpus_ids = [0] * len(candidates) # optional per-candidate corpus/collection tag
with torch.no_grad():
scores = model(
input_ids=input_ids,
attention_mask=attention_mask,
doc_spans=doc_spans,
corpus_ids=corpus_ids,
)
print(scores) # higher = more relevant
How you actually build the prompt (query/candidate layout, candidate selection, chunking long documents) is left to you — that's pipeline-specific and not part of this repo.
Benchmarks
Internal evals, merged checkpoint (score matches the pre-merge checkpoint to 8 decimal places on the two axes we re-checked):
| Benchmark | What it measures | Score |
|---|---|---|
| LawIRKo | Korean statute retrieval | 0.831 |
| KoSQA | Korean QA retrieval | 0.841 |
| AutoRAG | End-to-end RAG eval | 0.966 |
These are custom in-house harnesses, not directly comparable to public benchmarks like BEIR or MTEB.
Limitations
- Weights are fp32 — expect more memory and latency than an fp16/bf16 model this size.
- Training data, optimization, and prompt-construction details aren't included here — only the weights and the forward pass.
License
CC-BY-NC-4.0 — research and evaluation only. Contact us for commercial use. The base model (Qwen3-0.6B) is Apache-2.0, but this fine-tune (refiner, pooler, learned weights) is ours and ships under the license above.