Model reference · open weights
LFM2.5-VL-DSpark is an open-weight language model from LiquidAI. LFM2.5-VL-3B-DSpark (BF16) weighs 559 MB; the smallest configuration that runs it is RTX 3060 12 GB.
What it is
| Released by | LiquidAI |
|---|---|
| Type | Language models |
| Task | Vision + text |
| Parameters (lead) | 279M |
| Context | 128,000 tokens |
| Runs with | sglang |
| Based on | LiquidAI/LFM2.5-VL-3B |
| Released | 2026-09-18 |
| Popularity | 546 downloads / month |
| Weights | 559 MB (LFM2.5-VL-3B-DSpark (BF16), file size) |
| Licence | Its own licence terms |
What it runs on
Weights 559 MB (file size) · KV cache 8 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 644 MB on a small card · context up to 128,000 tokens.
| Card | Requests at once 8K tokens each | Requests at once 32K tokens each | Longest single request | Counted memory |
|---|---|---|---|---|
| RTX 3060 12 GB | 155 | 38 | all 125K | 11.6 GB |
| RTX 4060 Ti 16 GB | 211 | 52 | all 125K | 15.4 GB |
| RTX 3090 24 GB | 330 | 82 | all 125K | 23.4 GB |
| RTX 4090 24 GB | 330 | 82 | all 125K | 23.4 GB |
| RTX 5090 32 GB | 444 | 111 | all 125K | 31.0 GB |
| L40S 48 GB | 637 | 159 | all 125K | 44.0 GB |
| A100 80 GB | 1000+ | 286 | all 125K | 78.2 GB |
| H100 80 GB | 1000+ | 272 | all 125K | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | 1000+ | 330 | all 125K | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | 1000+ | 381 | all 125K | 107 GB |
| H200 141 GB | 1000+ | 495 | all 125K | 138 GB |
| B200 180 GB | 1000+ | 636 | all 125K | 176 GB |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 1.3 GB | 1.5 GB |
| 5 | 1.5 GB | 2.5 GB |
| 8 | 1.7 GB | 3.4 GB |
| 16 | 2.3 GB | 5.5 GB |
| 32 | 3.4 GB | 9.8 GB |
| 64 | 5.5 GB | 18.4 GB |
On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.
Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (grouped-query attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). Assumes vLLM 0.10 or later.
From the model card
src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/2b08LKpev0DNEk6DlnWkY.png" alt="Liquid AI" style="width: 100%; max-width: 100%; height: auto; display: inline-block; margin-bottom: 0.5em; margin-top: 0.5em;" />
LFM2.5-VL-3B-DSpark is an experimental speculative-decoding draft model that brings DSpark to
vision-language models. It allows LiquidAI/LFM2.5-VL-3B
to decode substantially faster without changing its output, for a minimal increase in memory footprint.
This is a drafter for LiquidAI/LFM2.5-VL-3B. In SGLang on a single H100, decoding runs up to 2.66× faster. On Apple silicon, it reaches up to 3.13× with MLX-VLM on an M5 Max and up to 2.14× with llama.cpp on an M3 Ultra.
Find more information about LFM2.5-VL-3B-DSpark in our blog post.
LFM2.5-VL-3B-DSpark is a DSpark speculative-decoding draft model with the following features:
LiquidAI/LFM2.5-VL-3Bhidden_size=2048, intermediate_size=6144 with SiLU/SwiGLU, GQA with num_attention_heads=32 / num_key_value_heads=8, head_dim=64[!NOTE] On Apple silicon the drafter is run at block size 8 rather than 9.
Use each drafter checkpoint with its corresponding target model:
| Drafter | Target |
|---|---|
| LFM2.5-VL-3B-DSpark | LFM2.5-VL-3B |
| LFM2.5-VL-3B-DSpark-GGUF | LFM2.5-VL-3B-GGUF |
Speculative decoding is exact under greedy decoding: the target verifies every proposed token, so the generated text is what the target would have produced on its own. Under matched sampling settings at non-zero temperatures, speculative decoding preserves the target model's output distribution. You get the speedup, not a different model.
See LiquidAI/LFM2.5-VL-3B for quality benchmarks.
This table reports the mean number of draft tokens accepted per target verification pass at batch size 1 and temperature 0. Higher acceptance generally enables greater acceleration, but the measured speedup also depends on hardware and runtime overhead.
Following MMSpec, the evaluation covers General VQA, Text VQA, Image Captioning, Chart VQA, Complex Reasoning, and Multi-turn Conversation.
| Benchmark | 1×H100 (SGLang, block 9) | Apple M5 Max (MLX-VLM, block 8) | Apple M3 Ultra (llama.cpp, block 8) |
|---|---|---|---|
| MMMU-Pro | 4.11 | 4.07 | 4.19 |
| Multi-turn | 3.46 | 3.24 | 3.31 |
| COCO | 4.57 | 4.21 | 4.50 |
| CharXiv | 4.11 | 4.34 | 4.04 |
| TextVQA | 3.74 | 4.08 | 3.58 |
| GQA | 4.14 | 3.77 | 3.93 |
This table reports the resulting runtime performance relative to the same target model without DSpark. Each cell is formatted as decode / end-to-end.
| Dataset | 1×H100 (SGLang) | Apple M5 Max (MLX-VLM) | Apple M3 Ultra (llama.cpp) |
|---|---|---|---|
| MMMU-Pro | 2.43× / 1.97× | 2.93× / 2.62× | 2.03× / 1.74× |
| Multi-turn | 2.04× / 1.83× | 2.30× / 1.91× | 1.57× / 1.37× |
| COCO | 2.66× / 2.27× | 3.13× / 2.59× | 2.14× / 1.77× |
| CharXiv | 2.39× / 1.97× | 2.94× / 1.71× | 1.87× / 1.56× |
| TextVQA | 2.14× / 1.64× | 2.69× / 1.56× | 1.64× / 1.33× |
| GQA | 2.35× / 1.77× | 2.67× / 1.93× | 1.77× / 1.30× |
All measurements use 16-bit processing for both the vision encoder and language backbone. The H100 results use SGLang on one H100 80GB in BF16 at batch size 1, temperature 0, and block size 9. The Apple results use FP16 weights at batch size 1, temperature 0, block size 8, and up to 2,048 output tokens. Measurements were collected with Pipette.
LFM2.5-VL-3B-DSpark is supported by SGLang for NVIDIA GPUs and MLX-VLM for Apple silicon. For llama.cpp, use the GGUF checkpoint.
Requires SGLang v0.5.19 or newer. Launch the target with the draft attached:
python -m sglang.launch_server \
--model-path LiquidAI/LFM2.5-VL-3B \
--speculative-algorithm DSPARK \
--speculative-draft-model-path LiquidAI/LQuoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.