Model reference · open weights

Prism-Qwen3.5-Reranker

Embeddings infgrad · community Reranker 1 build Open weights 1k dl/mo

Prism-Qwen3.5-Reranker is an open-weight embedding model from infgrad. Prism-Qwen3.5-Reranker-4B (BF16) weighs 8.4 GB; the smallest configuration that runs it is RTX 3060 12 GB.

What it is

Released byinfgrad
TypeEmbedding models
TaskReranker
Parameters (lead)4.2B
Context262,144 tokens
Runs withtransformers
Released2026-04-26
Popularity1k downloads / month
Weights8.4 GB (Prism-Qwen3.5-Reranker-4B (BF16), file size)
LicenceOpen weights

What it runs on

Memory and cards for Prism-Qwen3.5-Reranker-4B (BF16)

Weights 8.4 GB (file size) · overhead about 1.1 GB.

CardRunsCounted
memory
RTX 3060 12 GBfits11.6 GB
RTX 4060 Ti 16 GBfits15.4 GB
RTX 3090 24 GBfits23.4 GB
RTX 4090 24 GBfits23.4 GB
RTX 5090 32 GBfits31.0 GB
L40S 48 GBfits44.0 GB
A100 80 GBfits78.2 GB
H100 80 GBfits78.1 GB
RTX PRO 6000 Blackwell 96 GBfits93.8 GB
DGX Spark (GB10) 128 GB unifiedfits107 GB
H200 141 GBfits138 GB
B200 180 GBfits176 GB

Estimates, not measurements: the weights are the build's file size. No cache grows with use; a batch of inputs needs working memory of its own. Counted memory is 92 % of what CUDA reports for the card.

From the model card

What infgrad says about Prism-Qwen3.5-Reranker

Beyond Relevance Scoring — Jointly Producing Contributions and Evidence for Agentic Retrieval.

A reranker family that, unlike standard rerankers that emit only a relevance score, returns three things in a single forward pass: a calibrated score, a one-sentence contribution, and a self-contained evidence passage extracted from the document.

Read the full model card

Released models

Five checkpoints are released on the Hugging Face Hub. Four are fine-tuned from the Qwen3.5 backbone; one (-4B-exp) is an experimental extension built on top of Qwen3-Reranker-4B, demonstrating that the same recipe transfers to an existing LLM-based reranker without losing ranking quality.

ModelBackboneParametersHugging Face
Prism-Qwen3.5-Reranker-0.8BQwen3.50.8Binfgrad/Prism-Qwen3.5-Reranker-0.8B
Prism-Qwen3.5-Reranker-2BQwen3.52Binfgrad/Prism-Qwen3.5-Reranker-2B
Prism-Qwen3.5-Reranker-4BQwen3.54Binfgrad/Prism-Qwen3.5-Reranker-4B
Prism-Qwen3.5-Reranker-9BQwen3.59Binfgrad/Prism-Qwen3.5-Reranker-9B
Prism-Qwen3-Reranker-4B-expQwen3-Reranker-4B4Binfgrad/Prism-Qwen3-Reranker-4B-exp

Why this model?

In agentic / RAG pipelines, a relevance score is rarely the end goal. After deciding a document is relevant, the agent still has to read it, denoise it, and decide what to do next. Prism-Reranker folds that work into the reranker itself:

  • Relevance score — s(q, d) = σ(ℓ_yes − ℓ_no) ∈ (0, 1). Calibrated, ranking-ready.
  • `` — one sentence stating every core point the document contributes to the query. Useful for the agent to plan its next step without re-reading the doc.
  • **** — a self-contained, faithfully-rephrased rewrite of the query-relevant content. Drops irrelevant background, preserves verbatim proper nouns / numbers / dates / code / URLs. You can feed directly to a downstream LLM and skip the raw document — saving context tokens and removing web-noise.

If the document is not relevant, the model outputs no and stops. No contribution/evidence is generated.

Highlights

  • Backbones: Qwen3.5 series for the four main sizes, no architectural changes; one extension variant on top of Qwen3-Reranker-4B.
  • Context length: training data capped at 10K tokens per example, covering most real-world documents.
  • Multilingual: Chinese / English primary; other languages supported but with less coverage.
  • Keyword-query robust: agents often emit keyword-style queries instead of well-formed questions. ~30% of training queries were rewritten by an LLM into keyword form, so the model handles both natural and keyword queries.
  • Real-world data distribution: in addition to open reranker datasets (MS MARCO, T2Ranking, MIRACL, …), training includes synthetic queries paired with real Tavily / Exa web-search results, matching what an actual agent sees at inference time.
  • Length × score balanced: training data was rebalanced so that document length is not a relevance shortcut.
  • Training recipe: distillation (point-wise MSE on a strong commercial reranker's scores) + SFT on yes/no + +, supervised by a 5-LLM-as-judge ensemble.

Quickstart

Two ways to call the model. Both produce the same relevance score s(q, d) = σ(ℓ_yes − ℓ_no). Use A when you also want /. Use B when you only need a score and want a drop-in replacement for any other CrossEncoder reranker.

We use one shared example throughout so you can compare the outputs side by side:

QUERY = "What is the boiling point of water at sea level?"
DOCUMENTS = [
    "Water boils at 100 C (212 F) at standard atmospheric pressure (1 atm), "
    "which corresponds to sea-level conditions.",
    "Mount Everest is the highest mountain on Earth, with a peak elevation "
    "of 8,848 meters above sea level.",
]

A. Transformers (full output: score + contribution + evidence)

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

MODEL_PATH = "infgrad/Prism-Qwen3.5-Reranker-4B"  # or any sibling repo above

SYSTEM_PROMPT = (
    "Judge whether the Document meets the requirements based on "
    "the Query and the Instruct provided. "
)

INSTRUCTION = (
    'Judge if the document is relevant to the query. Reply "yes" or "no".\n'
    'On "yes", also emit:\n'
    "One sentence covering every core point the document "
    "contributes to the query, without elaboration.\n"
    "Self-contained rewrite of the query-relevant content. Rules:\n"
    "- Faithful: rephrase only; add or infer nothing.\n"
    "- Self-contained: evidence alone must fully answer the query.\n"
    "- Concise: drop query-irrelevant background.\n"
    "- Verbatim (no translation): proper nouns, terms, abbreviations, "
    "numbers, dates, code, URLs.\n"
    "- Output language: multilingual doc → query's language; else doc's language."
    ""
)

PROMPT_TEMPLATE = (
    "system\n{system}\n"
    "user\n"
    ": {instruction}\n"
    ": {query}\n"
    ": {doc}\n"
    "assistant\n\n\n\n\n"
)

def build_prompt(query: str, doc: str) -> str:
    return PROMPT_TEMPLATE.format(
        system=SYSTEM_PROMPT, instruction=INSTRUCTION, query=query, doc=doc
    )

tokenizer = AutoTokenizer.from_pretrained(MODEL_PATH)
model = AutoModelForCausalLM.from_pretrained(
    MODEL_PATH,
    torch_dtype=torch.bfloat16,
    device_map="cuda",
    attn_implementation="sdpa",
).eval()

yes_id = tokenizer.encode("yes", add_special_tokens=False)[0]
no_id = tokenizer.encode("no", add_special_tokens=False)[0]

@torch.no_grad()
def rerank(query: str, doc: str, max_new_tokens: int = 512):
    prompt = build_prompt(query, doc)

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms