Model reference · open weights
Ling-2.6-flash is an open-weight language model from inclusionAI. Ling-2.6-flash (BF16) weighs 215 GB; the smallest configuration that runs it is 2× H200 141 GB.
Ling-2.6-flash is an open-source instruct model developed by inclusionAI for text generation and agent tasks. It features a hybrid linear architecture with 107.5B total parameters and 7.4B active parameters, supporting a 131,072 token context length. The model is optimized for inference and token efficiency, supports English, and is released under the MIT license.
Summary of the inclusionAI/Ling-2.6-flash model card, 2026-10-01
What it is
| Released by | inclusionAI |
|---|---|
| Type | Language models |
| Task | Text gen |
| Parameters (lead) | 107.5B |
| Context | 131,072 tokens |
| Released | 2026-04-28 |
| Popularity | 3k downloads / month |
| Weights | 215 GB (Ling-2.6-flash (BF16), file size) |
| Licence | Open weights |
What it runs on
Weights 215 GB (file size) · KV cache 37 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 1.6 GB on a small card · context up to 131,072 tokens.
| Card | Requests at once 8K tokens each | Requests at once 32K tokens each | Longest single request | Counted memory |
|---|---|---|---|---|
| RTX 3060 12 GB … B200 180 GB 12 smaller cards | — | — | — | |
| 2× H200 141 GB tensor parallel | 72 | 18 | all 128K | 138 GB a card |
| 4× H100 80 GB tensor parallel | 52 | 13 | all 128K | 78.1 GB a card |
| 4× A100 80 GB tensor parallel | 75 | 18 | all 128K | 78.2 GB a card |
| 2× B200 180 GB tensor parallel | 186 | 46 | all 128K | 176 GB a card |
| 4× RTX PRO 6000 Blackwell 96 GB tensor parallel | 104 | 26 | all 128K | 93.8 GB a card |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 217 GB | 218 GB |
| 5 | 218 GB | 223 GB |
| 8 | 219 GB | 226 GB |
| 16 | 221 GB | 236 GB |
| 32 | 226 GB | 255 GB |
| 64 | 236 GB | 294 GB |
On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings. Split over several cards, this model's cache is copied to every card, not divided.
Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (multi-head latent attention (MLA)); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. This model's attention (MLA) copies its cache to every card of a tensor-parallel split; vLLM can also run it data-parallel, each card with its own cache, which holds several times more requests. Assumes vLLM 0.10 or later.
From the model card
Today, we announce the official open-source release of Ling-2.6-flash, an instruct model with 104B total parameters and 7.4B active parameters.
As agent capabilities mature, skyrocketing token consumption has become a primary barrier to deployment. Unlike standard chat, agent workflows involve massive inputs and complex, multi-step execution, driving up both compute demand and user costs. While the industry is pivoting toward "long-reasoning" to push performance ceilings, a critical question remains: Are these excessive reasoning tokens truly necessary for high-frequency, everyday agent use cases?
Faced with mounting token pressure, Ling-2.6-flash takes a different path. Rather than relying on longer outputs to chase higher scores, it is systematically optimized for inference efficiency, token efficiency, and agent performance—aiming to stay highly competitive while being faster, leaner, and better suited for real production workloads.
At a high level, Ling-2.6-flash is built around three core strengths:
We have conducted a comprehensive evaluation of Ling-2.6-flash across multiple authoritative benchmarks. Ling-2.6-flash performs strongly on representative agent benchmarks such as BFCL-V4, TAU2-bench, SWE-bench Verified, and PinchBench. In practice, Ling-2.6-flash delivers a strong user experience across frameworks including Claude Code, Kilo Code, Qwen Code, Hermes Agent, and OpenClaw, etc.
Beyond agent tasks, Ling-2.6-flash also delivers strong performance across general knowledge,mathematical reasoning, instruction following, and long-context understanding, remains well aligned with SOTA models in the same size class.
- PinchBench: Comparative scores are retrieved directly from the official PinchBench leaderboard (as of April 20, 2026), adhering to their evaluation modes (potentially Reasoning Mode).
- Claw-Eval: Comparative scores are sourced from the official Claw-Eval leaderboard (version dated 2026-03-25), adhering to their evaluation modes (potentially Reasoning Mode). Official scores for GPT-OSS-120B and GPT-5.4-mini are currently unavailable and have been omitted.
- TAU2-Bench: Evaluations are conducted using official v1.0.0 code and datasets. Following the GLM-5 evaluation protocol, we applied minor prompt adjustments in the Retail and Telecom domains to ensure users express requests clearly and to prevent premature session termination. Additionally, GPT-5.2 was utilized as the User Agent across all evaluated domains.
- IFBench: Scores for GPT-OSS-120B (low) and GPT-5.4-mini (Non-Reasoning) are sourced from the AA (Artificial Analysis) Leaderboard. All other model performance data are based on internal evaluation results.
Ling-2.6-flash continues the architectural direction introduced in Ling 2.5. Building on the Ling 2.0 foundation, we incorporate a hybrid linear attention mechanism, upgrading the original GQA attention design into a 1:7 MLA + Lightning Linear hybrid architecture through incremental training.
This combination of hybrid attention and a highly sparse MoE architecture gives Ling-2.6-flash a clear advantage in inference efficiency. Compared with mainstream SOTA models in a similar size class, Ling-2.6-flash not only delivers faster time-to-first-token, but also achieves substantially higher generation throughput in long-output scenarios. At peak, both prefill throughput and decode throughput can improve by up to around 4×.
As shown in the figure below, Ling-2.6-flash’s throughput advantage becomes more pronounced as both context length and generation length increase. More importantly, this is not just a benchmark-side gain on static metrics. In real deployment settings, the model continues to unlock stronger speed benefits as task complexity grows.
Whether the workload involves long-context understanding or extended text generation, Ling-2.6-flash preserves model capability while delivering faster responses, higher throughput, and better real-world deployment efficiency.
pip install uv
uv venv ~/my_ling_env
source ~/my_ling_env/bin/activate
# uv pip "sglang-kernel>=0.4.1"
uv pip install "sglang[all]>=0.5.10.post1" --prerelease=allow
Both BF16 and FP8 models are supported by SGLang now. It depends on the dtype of the model in ${MODEL_PATH}. Here is the example to run Ling-2.6-flash with 4 GPUs, where the master node IP is ${MASTER_IP} and server p
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.