Model reference · open weights
WizardLM-2-8x22B is an open-weight language model from alpindale. WizardLM-2-8x22B (BF16) weighs 281 GB; the smallest configuration that runs it is 4× H100 80 GB.
What it is
| Released by | alpindale |
|---|---|
| Type | Language models |
| Task | Text gen |
| Parameters (lead) | 140.6B |
| Context | 65,536 tokens |
| Runs with | transformers |
| Released | 2024-04-16 |
| Popularity | 9k downloads / month |
| Weights | 281 GB (WizardLM-2-8x22B (BF16), file size) |
| Licence | Open weights |
What it runs on
Weights 281 GB (file size) · KV cache 229 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 697 MB on a small card · context up to 65,536 tokens.
| Card | Requests at once 8K tokens each | Requests at once 32K tokens each | Longest single request | Counted memory |
|---|---|---|---|---|
| RTX 3060 12 GB … B200 180 GB 12 smaller cards | — | — | — | |
| 4× H100 80 GB tensor parallel | 8 | 2 | all 64K | 78.1 GB a card |
| 4× A100 80 GB tensor parallel | 15 | 3 | all 64K | 78.2 GB a card |
| 2× B200 180 GB tensor parallel | 32 | 8 | all 64K | 176 GB a card |
| 4× RTX PRO 6000 Blackwell 96 GB tensor parallel | 41 | 10 | all 64K | 93.8 GB a card |
| 4× H200 141 GB tensor parallel | 135 | 33 | all 64K | 138 GB a card |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 284 GB | 289 GB |
| 5 | 291 GB | 320 GB |
| 8 | 297 GB | 342 GB |
| 16 | 312 GB | 402 GB |
| 32 | 342 GB | 522 GB |
| 64 | 402 GB | 763 GB |
On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.
Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (grouped-query attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.
From the model card
🏠 WizardLM-2 Release Blog 🤗 HF Repo •🐱 Github Repo • 🐦 Twitter • 📃 [WizardLM] • 📃 [WizardCoder] • 📃 [WizardMath]
We introduce and opensource WizardLM-2, our next generation state-of-the-art large language models, which have improved performance on complex chat, multilingual, reasoning and agent. New family includes three cutting-edge models: WizardLM-2 8x22B, WizardLM-2 70B, and WizardLM-2 7B.
For more details of WizardLM-2 please read our release blog post and upcoming paper.
MT-Bench
We also adopt the automatic MT-Bench evaluation framework based on GPT-4 proposed by lmsys to assess the performance of models. The WizardLM-2 8x22B even demonstrates highly competitive performance compared to the most advanced proprietary models. Meanwhile, WizardLM-2 7B and WizardLM-2 70B are all the top-performing models among the other leading baselines at 7B to 70B model scales.
Human Preferences Evaluation
We carefully collected a complex and challenging set consisting of real-world instructions, which includes main requirements of humanity, such as writing, coding, math, reasoning, agent, and multilingual. We report the win:loss rate without tie:
We built a fully AI powered synthetic training system to train WizardLM-2 models, please refer to our blog for more details of this system.
❗Note for model system prompts usage:
A chat between a curious user and an artificial intelligence assistant. The assistant gives helpful,
detailed, and polite answers to the user's questions. USER: Hi ASSISTANT: Hello.
USER: Who are you? ASSISTANT: I am WizardLM.......
We provide a WizardLM-2 inference demo code on our github.
Detailed results can be found here
| Metric | Value |
|---|---|
| Avg. | 32.61 |
| IFEval (0-Shot) | 52.72 |
| BBH (3-Shot) | 48.58 |
| MATH Lvl 5 (4-Shot) | 22.28 |
| GPQA (0-shot) | 17.56 |
| MuSR (0-shot) | 14.54 |
| MMLU-PRO (5-shot) | 39.96 |
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.
Benchmarks
As published on the model card — the maker's own numbers, not measured by AxForge.
| Task | Dataset | Metric | Score |
|---|---|---|---|
| Text Generation | IFEval (0-Shot) | strict accuracy | 52.720 |
| Text Generation | BBH (3-Shot) | normalized accuracy | 48.580 |
| Text Generation | MATH Lvl 5 (4-Shot) | exact match | 22.280 |
| Text Generation | GPQA (0-shot) | acc_norm | 17.560 |
| Text Generation | MuSR (0-shot) | acc_norm | 14.540 |
| Text Generation | MMLU-PRO (5-shot) | accuracy | 39.960 |