Model reference · open weights

Qwen3.8-DSpark

Available as managed deployment Licence fee LLMs RadixArk Text gen 1 variants 368k dl/mo

Qwen3.8-DSpark is an open-weight language model from RadixArk. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

Released byRadixArk
TypeLanguage models
TaskText gen
Parameters (lead)1.9B
Context256k tokens
Runs withtransformers
Based onRadixArk/Qwen3.8-27B-NVFP4
Released2026-08-14
Popularity368k downloads / month
LicenceCommercial licence needed

About

What Qwen3.8-DSpark is

A DSpark speculative-decoding draft model for Qwen3.8-27B target models, trained with SpecForge and served with SGLang.

The checkpoint has been evaluated with both RadixArk/Qwen3.8-27B-NVFP4 and Qwen/Qwen3.8-27B-FP8 targets. The acceptance-length evaluation below uses the NVFP4 target. The throughput evaluation uses the FP8 target.

Read the full model card

Checkpoint

  • Draft parameters: 1,857,358,337 (1.86B)
  • Draft weight dtype: BF16
  • Hidden size: 5,120
  • Transformer layers: five full-attention layers
  • Attention: GQA with 32 query heads and eight key/value heads
  • Target auxiliary feature layers: 5, 19, 33, 47, 61
  • Markov head: VanillaMarkov, rank 256
  • Training target width: 16 future positions
  • Serving gamma: seven draft proposals
  • Target verification width: eight tokens, including the target bonus token
  • Maximum position embeddings: 262,144

The serving configuration uses block_size=7. The separate training_block_size=16 records the supervision width used during training.

Acceptance length

Results cover 64,675 completed requests across 17 workloads.

CategoryWorkloadPromptsDSpark v1DSpark v2
CodeHumanEval1643.04373.8468
CodeMBPP2573.22994.0603
CodeLiveCodeBench1,0552.59153.3462
CodeBigCodeBench1,1402.77523.4678
MathGSM8K1,3193.60304.5162
MathMATH-5005003.25594.2267
MathAIME 2025302.97983.9401
MathAMC23403.21114.1572
MathGSM-Symbolic2,0483.45544.2716
ChatMT-Bench802.60753.2860
ChatAlpaca52,0022.56593.2337
ChatArena-Hard-v27502.59103.2536
ChatIFEval5412.94573.6628
Misc.MMLU-Pro2,0482.83453.5964
Misc.GPQA-Diamond1982.76343.5109
Misc.LongBench-v25033.26023.9268
Misc.RULER-8K2,0004.95856.3009
AggregateDSpark v1DSpark v2Change
Request-count weighted, 64,675 prompts2.7211433.428567+26.00%
Workload macro, 17 workloads3.0983683.917881+26.45%

Acceptance-length protocol:

  • Runtime: SGLang v0.5.17
  • Hardware and topology: four NVIDIA GB300 GPUs, DP4 × TP1
  • Sampling: thinking enabled, temperature 1.0, top-p 0.95, top-k 20, seed 980406
  • Generation limit: 8,192 tokens; client concurrency: 128

Throughput

Throughput is total output tokens divided by end-to-end timed wall duration. Each speculative-decoding cell is output tok/s (speedup over autoregressive).

Concurrency 1

WorkloadAutoregressiveEAGLEDSpark v1DSpark v2
GSM8K94.2179.9 (1.91×)238.6 (2.53×)297.3 (3.16×)
MATH-50095.0174.0 (1.83×)214.4 (2.26×)280.0 (2.95×)
HumanEval95.8165.5 (1.73×)205.5 (2.14×)254.8 (2.66×)
MBPP93.8166.7 (1.78×)208.6 (2.22×)261.6 (2.79×)
MT-Bench95.8157.4 (1.64×)171.3 (1.79×)215.8 (2.25×)

Concurrency 8

WorkloadAutoregressiveEAGLEDSpark v1DSpark v2
GSM8K602.71,001.1 (1.66×)1,183.8 (1.96×)1,494.0 (2.48×)
MATH-500635.21,071.6 (1.69×)1,208.2 (1.90×)1,575.1 (2.48×)
HumanEval667.91,031.3 (1.54×)1,159.2 (1.74×)1,435.1 (2.15×)
MBPP635.4988.3 (1.56×)1,123.7 (1.77×)1,393.7 (2.19×)
MT-Bench647.9963.2 (1.49×)958.4 (1.48×)1,195.5 (1.85×)

Concurrency 32

WorkloadAutoregressiveEAGLEDSpark v1DSpark v2
GSM8K1,298.51,969.5 (1.52×)1,934.2 (1.49×)2,268.5 (1.75×)
MATH-5001,764.22,353.4 (1.33×)2,014.2 (1.14×)2,545.2 (1.44×)
HumanEval1,862.22,296.9 (1.23×)1,918.5 (1.03×)2,472.3 (1.33×)
MBPP1,738.42,286.3 (1.32×)1,926.3 (1.11×)2,413.1 (1.39×)
MT-Bench1,814.22,133.4 (1.18×)1,593.3 (0.88×)1,973.0 (1.09×)

Throughput protocol:

  • EAGLE uses the target-integrated MTP head loaded as Qwen3_5ForCausalLMMTP, without an external draft checkpoint
  • Hardware and topology: one NVIDIA H200 per workload, TP1 × DP1
  • 128 prompts per cell, dataset shuffle seed 42, concurrency 1/8/32, max_tokens=2048, reasoning effort xhigh, temperature 1.0, top-p 0.95, top-k 20
  • EAGLE serving: three speculative steps, top-k 1, four draft tokens, Mamba full-memory ratio 8.26, extra_buffer radix-cache strategy, float32 Mamba state
  • DSpark serving: gamma 7, target verify width 8, one speculative step, block size 7, Mamba full-memory ratio 11.93, extra_buffer radix-cache strategy, float32 Mamba state
  • Autoregressive and EAGLE serving used mem-fraction-static=0.85. DSpark used 0.80 with expandable CUDA allocation segments. Every mode retained its complete prefill and speculative-verification CUDA graph set and used max-running-requests=48.

Serving with SGLang

PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
SGLANG_RAGGED_VERIFY_MODE=static \
sglang serve \
  --trust-remote-code \
  --model-path Qwen/Qwen3.8-27B-FP8 \
  --kv-cache-dtype fp8_e4m3 \
  --mem-fraction-static 0.80 \
  --attention-backend flashinfer \
  --chunked-prefill-size 32768 \
  --max-prefill-tokens 32768 \
  --reasoning-parser qwen3 \
  --tool-call-parser qwen3_coder \
  --mamba-full-memory-ratio 11.93 \
  --mamba-radix-cache-strategy extra_buffer \
  --mamba-ssm-dtype float32 \
  --max-running-requests 48 \
  --speculative-algorithm DSPARK \
  --speculative-draft-model-path RadixArk/Qwen3.8-27B-DSpark \
  --speculative-draft-model-quantization unquant \
  --speculative-draft-attention-backend flashinfer \
  --spe

From the published model card. Full card on the HuggingFace links in the sidebar.

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys qwen3-8-dspark for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (qwen3-8-dspark below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/chat/completions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen3-8-dspark","messages":[{"role":"user","content":"Hello"}]}'

Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms