Model reference · open weights
Swift-1.5-Qwen3.8-Flash-Next is an open-weight language model from ukisai. Swift-1.5-Qwen3.8-Flash-Next-NVFP4 (NVFP4) weighs 186 GB; the smallest configuration that runs it is 2× H200 141 GB.
What it is
| Released by | ukisai |
|---|---|
| Type | Language models |
| Task | Vision + text |
| Parameters (lead) | 119.6B |
| Context | 262,144 tokens |
| Runs with | transformers |
| Based on | ukisai/Swift-Qwen3.8-Flash-Next |
| Released | 2026-09-22 |
| Popularity | 954 downloads / month |
| Weights | 186 GB (Swift-1.5-Qwen3.8-Flash-Next-NVFP4 (NVFP4), file size) |
| Licence | Its own licence terms |
What it runs on
Weights 186 GB (file size) · KV cache 25 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · 231 MB of fixed state per request · runtime overhead from 2.8 GB on a small card · context up to 262,144 tokens.
| Card | Requests at once 8K tokens each | Requests at once 32K tokens each | Longest single request | Counted memory |
|---|---|---|---|---|
| RTX 3060 12 GB … B200 180 GB 12 smaller cards | — | — | — | |
| 2× H200 141 GB tensor parallel · FP4 without its speed-up here | 160 | 67 | all 256K | 138 GB a card |
| 4× H100 80 GB tensor parallel · FP4 without its speed-up here | 135 | 46 | all 256K | 78.1 GB a card |
| 4× A100 80 GB tensor parallel · FP4 without its speed-up here | 182 | 62 | all 256K | 78.2 GB a card |
| 2× B200 180 GB tensor parallel | 323 | 134 | all 256K | 176 GB a card |
| 4× RTX PRO 6000 Blackwell 96 GB tensor parallel | 234 | 80 | all 256K | 93.8 GB a card |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 190 GB | 190 GB |
| 5 | 191 GB | 194 GB |
| 8 | 193 GB | 197 GB |
| 16 | 196 GB | 206 GB |
| 32 | 203 GB | 222 GB |
| 64 | 217 GB | 255 GB |
On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.
Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (hybrid: linear attention with full attention every few layers); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.
From the model card
NVFP4. Derived directly from Swift 1.5 Qwen3.8-Flash-Next. Native NVFP4 execution requires compatible NVIDIA Blackwell support.
Swift 1.5 Qwen3.8-Flash-Next is UkisAI's reasoning-efficient derivative of Qwen3.8-Flash-Next. It uses 63.4% fewer thinking tokens, with a 1.8x speed up while keeping the accuracy loss <1% vs base on xhigh.
We gave base Qwen3.8-Flash-Next and Swift 1.5 Qwen3.8-Flash-Next the same prompt:
Create a 3D endless runner that has the fast, playful feel of Subway Surfers, but make the world and characters your own. I want to run through a lively place, dodge things, collect rewards, and feel the pace build the longer I survive. Make it fun to control and visually memorable. Use your judgment for the setting, mechanics, and little details that make it feel like a real game. Build it so I can launch and play it locally, then tell me how to run it.
Try the game yourself here: https://ukisai.com/swift-games/flash-next
Base Qwen3.8-Flash-Next took 8 minutes 52 seconds to build its game. Swift 1.5 took 4 minutes 56 seconds.
We made Swift Flash Next efficient by figuring out which tokens were linked to pathological overthinking and penalizing them without "attacking" the reasoning length directly then regained the accuracy with RL and OPD, leading to "compressed" token usage while maintaining accuracy.
Swift 1.5 produces shorter reasoning traces. In our testing, we also observe fewer overthinking errors.
This release also features our previously mentioned post-training methods adapted specifically for coding and long-horizon agent work such as personal agents, terminal use and software engineering.
Our training data is viewable here: https://huggingface.co/datasets/ukisai/Qwen3.8-27B-multi-turn-agent-sft albeit it is not used out of the box, but rather re-sampled, turned into proper RL environments etc.
All scores compare the Qwen3.8-Flash-Next BF16 base with the Swift 1.5 BF16 checkpoint. Token columns report thinking tokens, except Terminal-Bench 2.1, which reports total generated output tokens.
.swift-table { width:100%; table-layout:fixed; border-collapse:separate; border-spacing:0; overflow:hidden; border:1px solid #27344A; border-radius:20px; background:#0D111B; font-family:-apple-system,BlinkMacSystemFont,"Segoe UI",Roboto,sans-serif; font-size:14px; color:#BFBDBD; } .swift-table th { padding:13px 8px; text-align:center; font-weight:700; color:#AEB5C7; background:#0D111B; border-right:1px solid #27344A; border-bottom:1px solid #27344A; } .swift-table td { padding:14px 8px; text-align:center; color:#BFBDBD; background:#0D111B; border-right:1px solid #27344A; border-bottom:1px solid #27344A; vertical-align:middle; overflow-wrap:break-word; } .swift-table tr > :last-child { border-right:0; } .swift-table tbody tr:last-child td { border-bottom:0; } .swift-table .benchmark-heading { color:#B7BDCD; background:#0D111B; border-bottom:3px solid #7D45B5; } .swift-table .score-heading { color:#F0C5FF; background:#52239E; border-bottom:3px solid #7D45B5; } .swift-table .tokens-heading, .swift-table .median-heading { color:#D4E8FF; background:#304FC2; border-bottom:3px solid #5687E6; } .swift-table .benchmark { padding-left:18px; text-align:left; color:#FFFFFF; font-weight:600; } .swift-table .section { padding:12px 18px; text-align:left; color:#B489FF; background:#2A2541; font-weight:700; letter-spacing:.08em; text-transform:uppercase; border-top:1px solid #3A3159; border-bottom:1px solid #3A3159; } .swift-table .swift { background:#171127; } .swift-table thead tr:nth-child(2) .swift { color:#D3A0FF; } .swift-table .reduction { color:#69BFFF; background:#101B2C; font-weight:700; } .swift-table .detail { color:#8C94A8; font-size:12px; font-weight:500; }
@media (max-width: 640px) { .swift-table { display:block !important; width:100% !important; max-width:100%; overflow-x:auto !important; -webkit-overflow-scrolling:touch; table-layout:auto !important; } .swift-table th, .swift-table td { min-width:100px; } .swift-table th:first-child, .swift-table td:first-child { min-width:160px; } }
GPQA-Diamond at each reasoning_effort setting, Swift 1.5 against the base at the same setting:
At xhigh, Swift 1.5 trails base by 0.20 percentage points while using 55.8% fewer mean and 63.4% fewer median thinking tokens. Medium and low save tokens but also lose 2.62 and 2.42 percentage points respectively.
| Format | Repository | Runtime |
|---|---|---|
| AWQ INT4 (W4A16) | Swift-1.5-Qwen3.8-Flash-Next-W4A16-AWQ | vLLM (compressed-tensors) |
| AutoRound INT4 (W4A16) | Swift-1.5-Qwen3.8-Flash-Next-W4A16-AutoRound | vLLM (auto-round) |
| NVFP4 | Swift-1.5-Qwen3.8-Flash-Next-NVFP4 | NVIDIA Blackwell |
| GGUF | Swift-1.5-Qwen3.8-Flash-Next-GGUF | llama.cpp |
| GSQ-RCO GGUF (compact 2–3 bit) | Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF | llama.cpp |
Swift 1.5 Qwen3.8-Flash-Next is a derivative of Qwen3.8-Flash-Next (Copyright (c) 2026 Qwen, Qwen Community License 1.0). UkisAI's contribution, including the adapted weights, is licensed under the Swift Open License v1.0. See [NOTICE](https://huggingface.co/ukisai/Sw
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.