cheon-reranker-0.6b-v1

A reranker for legal and administrative document search, built on Qwen3-0.6B.

Most rerankers score query-document pairs one at a time. This one instead encodes the query and all candidates together in a single context, then runs a small bidirectional refiner on top of the pooled candidate vectors so candidates can attend to each other before scoring. The refiner has no positional encoding, so shuffling the candidate order doesn't change the scores.

  • Base: Qwen3-0.6B (Apache-2.0)
  • Params: ~609M total (adapter already merged into the base weights)
  • Refiner: 4 layers, 512 hidden dim, single attention-pooled slot per candidate
  • Weights: model.safetensors, fp32

Usage

Loads straight into AutoModel with trust_remote_code=True. The included modeling_cheon_reranker.py is a from-scratch reimplementation of the forward pass (no training code) — checked against the original checkpoint on random inputs, max difference 2.4e-07, i.e. floating-point noise.

import torch
from transformers import AutoModel, AutoTokenizer

repo = "cheonai/cheon-reranker-0.6b-v1"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModel.from_pretrained(repo, trust_remote_code=True, torch_dtype=torch.float32)
model.eval()

# Query + candidates go into one sequence. Track each candidate's token span
# as you build it instead of hardcoding positions. Candidates don't need to
# share a language - the tokenizer and span logic don't care.
query = "query: 이 사건의 처리 기한은 언제까지인가요?\n"
candidates = [
    "candidate (en): The appeal must be filed within 30 days of the decision.",
    "candidate (ko): 이의신청은 처분을 안 날부터 30일 이내에 제기해야 합니다.",
    "candidate (zh): 上诉必须在裁定后30天内提出。",
    "candidate (ja): 不服申立ては、決定を知った日から30日以内に行う必要があります。",
]

ids = tok(query, add_special_tokens=False)["input_ids"]
doc_spans = []
for text in candidates:
    piece = tok(text, add_special_tokens=False)["input_ids"]
    start = len(ids)
    ids += piece
    doc_spans.append((start, len(ids)))

input_ids = torch.tensor([ids])
attention_mask = torch.ones_like(input_ids)
corpus_ids = [0] * len(candidates)  # optional per-candidate corpus/collection tag

with torch.no_grad():
    scores = model(
        input_ids=input_ids,
        attention_mask=attention_mask,
        doc_spans=doc_spans,
        corpus_ids=corpus_ids,
    )
print(scores)  # higher = more relevant

How you actually build the prompt (query/candidate layout, candidate selection, chunking long documents) is left to you — that's pipeline-specific and not part of this repo.

Benchmarks

Internal evals, merged checkpoint (score matches the pre-merge checkpoint to 8 decimal places on the two axes we re-checked):

Benchmark What it measures Score
LawIRKo Korean statute retrieval 0.831
KoSQA Korean QA retrieval 0.841
AutoRAG End-to-end RAG eval 0.966

These are custom in-house harnesses, not directly comparable to public benchmarks like BEIR or MTEB.

Limitations

  • Weights are fp32 — expect more memory and latency than an fp16/bf16 model this size.
  • Training data, optimization, and prompt-construction details aren't included here — only the weights and the forward pass.

License

CC-BY-NC-4.0 — research and evaluation only. Contact us for commercial use. The base model (Qwen3-0.6B) is Apache-2.0, but this fine-tune (refiner, pooler, learned weights) is ours and ships under the license above.