Model reference · open weights

WebDancer

LLMs Alibaba-NLP Text gen 1 build Open weights 3k dl/mo

WebDancer is an open-weight language model from Alibaba. WebDancer-32B (BF16) weighs 65.5 GB; the smallest configuration that runs it is H100 80 GB.

  • WebDancer is a 32.8B parameter text-generation model developed by Alibaba for autonomous information seeking and agentic search reasoning.
  • It utilizes a ReAct framework and was trained using a four-stage paradigm that includes supervised fine-tuning and reinforcement learning.
  • The model supports a context length of 40,960 tokens and is released under the MIT licence.

Summary of the Alibaba-NLP/WebDancer-32B model card, 2026-10-04

What it is

Released byAlibaba
Released2025-06-23
Parameters32.8B
VRAM65.5 GB for the weights

What it runs on

Memory and cards for WebDancer-32B (BF16)

65.5 GBweights, file size
262 MBcache per 1K tokens
896 MBruntime overhead, at least
40,960 tokenscontext max
CardRequests at onceContext maxMemory
8K each32K each
RTX 3060 12 GB … L40S 48 GB
6 smaller cards
———
A100 80 GB51all 40K78.2 GB
H100 80 GB3—25K78.1 GB
RTX PRO 6000 Blackwell 96 GB102all 40K93.8 GB
DGX Spark (GB10) 128 GB unified164all 40K107 GB
H200 141 GB317all 40K138 GB
B200 180 GB4812all 40K176 GB
2× L40S 48 GB
tensor parallel
92all 40K44.0 GB a card
4× RTX 4090 24 GB
tensor parallel
112all 40K23.4 GB a card
4× RTX 3090 24 GB
tensor parallel
112all 40K23.4 GB a card
4× RTX 5090 32 GB
tensor parallel
256all 40K31.0 GB a card
2× H100 80 GB
tensor parallel
369all 40K78.1 GB a card
2× A100 80 GB
tensor parallel
4110all 40K78.2 GB a card
Memory needed at each load
Requests at once8K tokens each32K tokens each
168.6 GB75.0 GB
577.2 GB109 GB
883.6 GB135 GB
16101 GB204 GB
32135 GB341 GB
64204 GB616 GB

One card, with vLLM's small-card settings.

From the model card

What Alibaba says about WebDancer

Read the model card

This model was presented in the paper WebDancer: Towards Autonomous Information Seeking Agency.

You can download the model then run the inference scipts in https://github.com/Alibaba-NLP/WebAgent.

  • Native agentic search reasoning model using ReAct framework towards autonomous information seeking agency and Deep Research-like model.
  • We introduce a four-stage training paradigm comprising browsing data construction, trajectory sampling, supervised fine-tuning for effective cold start, and reinforcement learning for improved generalization, enabling the agent to autonomously acquire autonomous search and reasoning skills.
  • Our data-centric approach integrates trajectory-level supervision fine-tuning and reinforcement learning (DAPO) to develop a scalable pipeline for training agentic systems via SFT or RL.
  • WebDancer achieves a Pass@3 score of 61.1% on GAIA and 54.6% on WebWalkerQA.

Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms