Model reference · open weights

WizardLM-2-8x22B

LLMs alpindale · community Text gen 1 build Open weights 9k dl/mo

WizardLM-2-8x22B is an open-weight language model from alpindale. WizardLM-2-8x22B (BF16) weighs 281 GB; the smallest configuration that runs it is 4× H100 80 GB.

What it is

Released byalpindale
TypeLanguage models
TaskText gen
Parameters (lead)140.6B
Context65,536 tokens
Runs withtransformers
Released2024-04-16
Popularity9k downloads / month
Weights281 GB (WizardLM-2-8x22B (BF16), file size)
LicenceOpen weights

What it runs on

Memory and cards for WizardLM-2-8x22B (BF16)

Weights 281 GB (file size) · KV cache 229 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 697 MB on a small card · context up to 65,536 tokens.

CardRequests at once
8K tokens each
Requests at once
32K tokens each
Longest single
request
Counted
memory
RTX 3060 12 GB … B200 180 GB
12 smaller cards
———
4× H100 80 GB
tensor parallel
82all 64K78.1 GB a card
4× A100 80 GB
tensor parallel
153all 64K78.2 GB a card
2× B200 180 GB
tensor parallel
328all 64K176 GB a card
4× RTX PRO 6000 Blackwell 96 GB
tensor parallel
4110all 64K93.8 GB a card
4× H200 141 GB
tensor parallel
13533all 64K138 GB a card
Memory needed at each load
Requests at once8K tokens each32K tokens each
1284 GB289 GB
5291 GB320 GB
8297 GB342 GB
16312 GB402 GB
32342 GB522 GB
64402 GB763 GB

On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.

Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (grouped-query attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.

From the model card

What alpindale says about WizardLM-2-8x22B

🏠 WizardLM-2 Release Blog 🤗 HF Repo •🐱 Github Repo • 🐦 Twitter • 📃 [WizardLM] • 📃 [WizardCoder] • 📃 [WizardMath]

Read the full model card

See here for the WizardLM-2-7B re-upload.

News 🔥🔥🔥 [2024/04/15]

We introduce and opensource WizardLM-2, our next generation state-of-the-art large language models, which have improved performance on complex chat, multilingual, reasoning and agent. New family includes three cutting-edge models: WizardLM-2 8x22B, WizardLM-2 70B, and WizardLM-2 7B.

  • WizardLM-2 8x22B is our most advanced model, demonstrates highly competitive performance compared to those leading proprietary works and consistently outperforms all the existing state-of-the-art opensource models.
  • WizardLM-2 70B reaches top-tier reasoning capabilities and is the first choice in the same size.
  • WizardLM-2 7B is the fastest and achieves comparable performance with existing 10x larger opensource leading models.

For more details of WizardLM-2 please read our release blog post and upcoming paper.

Model Details

Model Capacities

MT-Bench

We also adopt the automatic MT-Bench evaluation framework based on GPT-4 proposed by lmsys to assess the performance of models. The WizardLM-2 8x22B even demonstrates highly competitive performance compared to the most advanced proprietary models. Meanwhile, WizardLM-2 7B and WizardLM-2 70B are all the top-performing models among the other leading baselines at 7B to 70B model scales.

Human Preferences Evaluation

We carefully collected a complex and challenging set consisting of real-world instructions, which includes main requirements of humanity, such as writing, coding, math, reasoning, agent, and multilingual. We report the win:loss rate without tie:

  • WizardLM-2 8x22B is just slightly falling behind GPT-4-1106-preview, and significantly stronger than Command R Plus and GPT4-0314.
  • WizardLM-2 70B is better than GPT4-0613, Mistral-Large, and Qwen1.5-72B-Chat.
  • WizardLM-2 7B is comparable with Qwen1.5-32B-Chat, and surpasses Qwen1.5-14B-Chat and Starling-LM-7B-beta.

Method Overview

We built a fully AI powered synthetic training system to train WizardLM-2 models, please refer to our blog for more details of this system.

Usage

❗Note for model system prompts usage:

A chat between a curious user and an artificial intelligence assistant. The assistant gives helpful,
detailed, and polite answers to the user's questions. USER: Hi ASSISTANT: Hello.
USER: Who are you? ASSISTANT: I am WizardLM.......

We provide a WizardLM-2 inference demo code on our github.

Open LLM Leaderboard Evaluation Results

Detailed results can be found here

MetricValue
Avg.32.61
IFEval (0-Shot)52.72
BBH (3-Shot)48.58
MATH Lvl 5 (4-Shot)22.28
GPQA (0-shot)17.56
MuSR (0-shot)14.54
MMLU-PRO (5-shot)39.96

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

Benchmarks

Reported results

As published on the model card — the maker's own numbers, not measured by AxForge.

TaskDatasetMetricScore
Text GenerationIFEval (0-Shot)strict accuracy52.720
Text GenerationBBH (3-Shot)normalized accuracy48.580
Text GenerationMATH Lvl 5 (4-Shot)exact match22.280
Text GenerationGPQA (0-shot)acc_norm17.560
Text GenerationMuSR (0-shot)acc_norm14.540
Text GenerationMMLU-PRO (5-shot)accuracy39.960
© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms