Model reference · open weights
AdvancedMathBench-AutoVerifier is an open-weight language model from internlm. AdvancedMathBench-AutoVerifier (BF16) weighs 73.0 GB; the smallest configuration that runs it is A100 80 GB.
What it is
| Released by | internlm |
|---|---|
| Type | Language models |
| Task | Vision + text |
| Parameters (lead) | 36.0B |
| Context | 262,144 tokens |
| Runs with | transformers |
| Released | 2026-09-29 |
| Popularity | 0 downloads / month |
| Weights | 73.0 GB (AdvancedMathBench-AutoVerifier (BF16), file size) |
| Licence | Licence not stated |
What it runs on
Weights 73.0 GB (file size) · KV cache 20 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · 129 MB of fixed state per request · runtime overhead from 2.5 GB on a small card · context up to 262,144 tokens.
| Card | Requests at once 8K tokens each | Requests at once 32K tokens each | Longest single request | Counted memory |
|---|---|---|---|---|
| RTX 3060 12 GB … L40S 48 GB 6 smaller cards | — | — | — | |
| A100 80 GB | 9 | 3 | 125K | 78.2 GB |
| H100 80 GB | — | — | — | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | 40 | 14 | all 256K | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | 85 | 31 | all 256K | 107 GB |
| H200 141 GB | 189 | 70 | all 256K | 138 GB |
| B200 180 GB | 311 | 115 | all 256K | 176 GB |
| 2× L40S 48 GB tensor parallel | 33 | 12 | all 256K | 44.0 GB a card |
| 4× RTX 4090 24 GB tensor parallel | 22 | 7 | 248K | 23.4 GB a card |
| 4× RTX 3090 24 GB tensor parallel | 23 | 7 | 252K | 23.4 GB a card |
| 4× RTX 5090 32 GB tensor parallel | 88 | 28 | all 256K | 31.0 GB a card |
| 2× H100 80 GB tensor parallel | 220 | 81 | all 256K | 78.1 GB a card |
| 2× A100 80 GB tensor parallel | 264 | 98 | all 256K | 78.2 GB a card |
| 2× RTX PRO 6000 Blackwell 96 GB tensor parallel | 326 | 121 | all 256K | 93.8 GB a card |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 75.7 GB | 76.2 GB |
| 5 | 76.9 GB | 79.4 GB |
| 8 | 77.8 GB | 81.8 GB |
| 16 | 80.2 GB | 88.2 GB |
| 32 | 84.9 GB | 101 GB |
| 64 | 94.4 GB | 127 GB |
On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.
Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (hybrid: linear attention with full attention every few layers); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.
From the model card
Paper · HF Paper · GitHub · Dataset · AutoVerifier
AutoVerifier evaluates natural-language mathematical proofs, explains errors, and identifies the earliest incorrect step. It serves as the automatic grader for AdvancedMathBench's ProverBench.
Qwen3_5MoeForConditionalGeneration.InternS1Tokenizer; requires sentencepiece and
trust_remote_code=True after reviewing the tokenizer code.Use proof_verifier.md with a problem, an optional reference solution, and a candidate proof split into zero-indexed steps. The following constructs the input without loading the model weights:
from pathlib import Path
from transformers import AutoTokenizer
model_dir = "." # Local model repository directory
tokenizer = AutoTokenizer.from_pretrained(model_dir, trust_remote_code=True)
steps = ["A candidate proof step.", "Another candidate proof step."]
proof = "\n\n".join(
f"\n\n{step}\n\n" for i, step in enumerate(steps)
)
template = Path(model_dir, "prompts/proof_verifier.md").read_text(encoding="utf-8")
prompt = template.format(
problem="The mathematical problem.", human_solution="", solution=proof,
)
text = tokenizer.apply_chat_template(
[{"role": "user", "content": prompt}],
tokenize=False, add_generation_prompt=True, enable_thinking=True,
)
The final response contains an assessment, identified errors, and the first error index. For example, a no-error judgment is:
-1 means no error was found; nonnegative indices identify the earliest error,
starting from 0. Parse the final answer after `` when present.
ProverBench checks each proof 8 times and accepts it only when all eight
valid judgments report -1. Missing or malformed judgments do not count as
acceptance. Sampling settings are documented in
evaluation_settings.json.
AutoVerifier is a learned grader, not a formal proof checker, and can make errors. Tested package versions and validation scope are recorded in compatibility.json. License notices are provided in LICENSE and NOTICE.md.
If you use AdvancedMathBench AutoVerifier in your research, please cite:
@misc{kong2026advancedmathbenchbenchmarksuiteadvanced,
title = {AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification},
author = {Lingkai Kong and Zijian Wu and Yuzhe Gu and Haiteng Zhao and Wenyong Huang and Shuang Sun and Zhicheng Xiong and Xiaotian Zhang and Shuya Zhao and Yan Wang and Disheng Xu and Wenwei Zhang and Kai Chen},
year = {2026},
eprint = {2607.11849},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
doi = {10.48550/arXiv.2607.11849},
url = {https://arxiv.org/abs/2607.11849}
}
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.