Model reference · open weights
CompassJudger-2 is an open-weight embedding model from opencompass. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.
Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.
What it is
| Released by | opencompass |
|---|---|
| Type | Embedding models |
| Task | Reranker |
| Parameters (lead) | 32.8B |
| Context | 32k tokens |
| Runs with | transformers |
| Released | 2025-07-09 |
| Popularity | 1k downloads / month |
| Licence | Open weights |
About
src="https://img.shields.io/badge/CompassJudger--2-Paper-red?logo=arxiv&logoColor=red" alt="CompassJudger-2" style="display: inline-block; vertical-align: middle;" />
We introduce CompassJudger-2, a novel series of generalist judge models designed to overcome the narrow specialization and limited robustness of existing LLM-as-judge solutions. Current judge models often struggle with comprehensive evaluation, but CompassJudger-2 addresses these limitations with a powerful new training paradigm.
Key contributions of our work include:
This repository contains the CompassJudger-2 series of models, fine-tuned on the Qwen2.5-Instruct series.
| Model Name | Size | Base Model | Download | Notes |
|---|---|---|---|---|
| 👉 CompassJudger-2-7B-Instruct | 7B | Qwen2.5-7B-Instruct | 🤗 Model | Fine-tuned for generalist judge capabilities. |
| 👉 CompassJudger-2-32B-Instruct | 32B | Qwen2.5-32B-Instruct | 🤗 Model | A larger, more powerful judge model. |
Here is a simple example demonstrating how to load the model and use it for pairwise evaluation.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_path = "opencompass/CompassJudger-2-7B-Instruct"
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype="auto",
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(model_name)
# Example: Pair-wise Comparison
prompt = """
Please act as an impartial judge to evaluate the responses provided by two AI assistants to the user question below. Your evaluation should focus on the following criteria: helpfulness, relevance, accuracy, depth, creativity, and level of detail.
- Do not let the order of presentation, response length, or assistant names influence your judgment.
- Base your decision solely on how well each response addresses the user’s question and adheres to the instructions.
Your final reply must be structured in the following format:
{
"Choice": "[Model A or Model B]"
}
User Question: {question}
Model A's Response: {answerA}
Model B's Response: {answerB}
Now it's your turn. Please provide selection result as required:
"""
messages = [
{"role": "user", "content": prompt}
]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True
)
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)
generated_ids = model.generate(
**model_inputs,
max_new_tokens=2048
)
generated_ids = [
output_ids[len(input_ids):] for input_ids, output_ids in zip(model_inputs.input_ids, generated_ids)
]
response = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]
print(response)
CompassJudger-2 sets a new state-of-the-art for judge models, outperforming general models, reward models, and other specialized judge models across a wide range of benchmarks.
| Model | JudgerBench V2 | JudgeBench | RMB | RewardBench | Average |
|---|---|---|---|---|---|
| 7B Judge Models | |||||
| CompassJudger-1-7B-Instruct | 57.96 | 46.00 | 38.18 | 80.74 | 55.72 |
| Con-J-7B-Instruct | 52.35 | 38.06 | 71.50 | 87.10 | 62.25 |
| RISE-Judge-Qwen2.5-7B | 46.12 | 40.48 | 72.64 | 88.20 | 61.61 |
| CompassJudger-2-7B-Instruct | 60.52 | 63.06 | 73.90 | 90.96 | 72.11 |
| 32B+ Judge Models | |||||
| CompassJudger-1-32B-Instruct | 60.33 | 62.29 | 77.63 | 86.17 | 71.61 |
| Skywork-Critic-Llama-3.1-70B | 52.41 | 50.65 | 65.50 | 93.30 | 65.47 |
| RISE-Judge-Qwen2.5-32B | 56.42 | 63.87 | 73.70 | 92.70 | 71.67 |
| CompassJudger-2-32B-Instruct | 62.21 | 65.48 | 72.98 | 92.62 | 73.32 |
| General Models (for reference) | |||||
| Qwen2.5-32B-Instruct | 62.97 | 59.84 | 74.99 | 85.61 | 70.85 |
| DeepSeek-V3-0324 | 64.43 | 59.68 | 78.16 | 85.17 | 71.86 |
| Qwen3-235B-A22B | 61.40 | 65.97 | 75.59 | 84.68 | 71.91 |
For detailed benchmark performance and methodology, please refer to our 📑 [Paper](https://arxiv.org/a
From the published model card. Full card on the HuggingFace links in the sidebar.
Using it via the API
Once AxForge deploys compassjudger-2 for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (compassjudger-2 below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/embeddings \
-H "Authorization: Bearer $AXFORGE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"compassjudger-2","input":"text to embed"}'
Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.