Model reference · open weights
WebDancer is an open-weight language model from Alibaba. WebDancer-32B (BF16) weighs 65.5 GB; the smallest configuration that runs it is H100 80 GB.
Summary of the Alibaba-NLP/WebDancer-32B model card, 2026-10-04
What it is
| Released by | Alibaba |
|---|---|
| Released | 2025-06-23 |
| Parameters | 32.8B |
| VRAM | 65.5 GB for the weights |
What it runs on
| Card | Requests at once | Context max | Memory | |
|---|---|---|---|---|
| 8K each | 32K each | |||
| RTX 3060 12 GB … L40S 48 GB 6 smaller cards | — | — | — | |
| A100 80 GB | 5 | 1 | all 40K | 78.2 GB |
| H100 80 GB | 3 | — | 25K | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | 10 | 2 | all 40K | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | 16 | 4 | all 40K | 107 GB |
| H200 141 GB | 31 | 7 | all 40K | 138 GB |
| B200 180 GB | 48 | 12 | all 40K | 176 GB |
| 2× L40S 48 GB tensor parallel | 9 | 2 | all 40K | 44.0 GB a card |
| 4× RTX 4090 24 GB tensor parallel | 11 | 2 | all 40K | 23.4 GB a card |
| 4× RTX 3090 24 GB tensor parallel | 11 | 2 | all 40K | 23.4 GB a card |
| 4× RTX 5090 32 GB tensor parallel | 25 | 6 | all 40K | 31.0 GB a card |
| 2× H100 80 GB tensor parallel | 36 | 9 | all 40K | 78.1 GB a card |
| 2× A100 80 GB tensor parallel | 41 | 10 | all 40K | 78.2 GB a card |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 68.6 GB | 75.0 GB |
| 5 | 77.2 GB | 109 GB |
| 8 | 83.6 GB | 135 GB |
| 16 | 101 GB | 204 GB |
| 32 | 135 GB | 341 GB |
| 64 | 204 GB | 616 GB |
One card, with vLLM's small-card settings.
From the model card
This model was presented in the paper WebDancer: Towards Autonomous Information Seeking Agency.
You can download the model then run the inference scipts in https://github.com/Alibaba-NLP/WebAgent.
Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.