Model reference · open weights

LFM2.5-VL-DSpark

LLMs LiquidAI Vision + text 1 build Its own licence terms 546 dl/mo

LFM2.5-VL-DSpark is an open-weight language model from LiquidAI. LFM2.5-VL-3B-DSpark (BF16) weighs 559 MB; the smallest configuration that runs it is RTX 3060 12 GB.

What it is

Released byLiquidAI
TypeLanguage models
TaskVision + text
Parameters (lead)279M
Context128,000 tokens
Runs withsglang
Based onLiquidAI/LFM2.5-VL-3B
Released2026-09-18
Popularity546 downloads / month
Weights559 MB (LFM2.5-VL-3B-DSpark (BF16), file size)
LicenceIts own licence terms

What it runs on

Memory and cards for LFM2.5-VL-3B-DSpark (BF16)

Weights 559 MB (file size) · KV cache 8 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 644 MB on a small card · context up to 128,000 tokens.

CardRequests at once
8K tokens each
Requests at once
32K tokens each
Longest single
request
Counted
memory
RTX 3060 12 GB15538all 125K11.6 GB
RTX 4060 Ti 16 GB21152all 125K15.4 GB
RTX 3090 24 GB33082all 125K23.4 GB
RTX 4090 24 GB33082all 125K23.4 GB
RTX 5090 32 GB444111all 125K31.0 GB
L40S 48 GB637159all 125K44.0 GB
A100 80 GB1000+286all 125K78.2 GB
H100 80 GB1000+272all 125K78.1 GB
RTX PRO 6000 Blackwell 96 GB1000+330all 125K93.8 GB
DGX Spark (GB10) 128 GB unified1000+381all 125K107 GB
H200 141 GB1000+495all 125K138 GB
B200 180 GB1000+636all 125K176 GB
Memory needed at each load
Requests at once8K tokens each32K tokens each
11.3 GB1.5 GB
51.5 GB2.5 GB
81.7 GB3.4 GB
162.3 GB5.5 GB
323.4 GB9.8 GB
645.5 GB18.4 GB

On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.

Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (grouped-query attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). Assumes vLLM 0.10 or later.

From the model card

What LiquidAI says about LFM2.5-VL-DSpark

src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/2b08LKpev0DNEk6DlnWkY.png" alt="Liquid AI" style="width: 100%; max-width: 100%; height: auto; display: inline-block; margin-bottom: 0.5em; margin-top: 0.5em;" />

Read the full model card

LFM2.5-VL-3B-DSpark

LFM2.5-VL-3B-DSpark is an experimental speculative-decoding draft model that brings DSpark to vision-language models. It allows LiquidAI/LFM2.5-VL-3B to decode substantially faster without changing its output, for a minimal increase in memory footprint.

This is a drafter for LiquidAI/LFM2.5-VL-3B. In SGLang on a single H100, decoding runs up to 2.66× faster. On Apple silicon, it reaches up to 3.13× with MLX-VLM on an M5 Max and up to 2.14× with llama.cpp on an M3 Ultra.

Find more information about LFM2.5-VL-3B-DSpark in our blog post.

🗒️ Model Details

LFM2.5-VL-3B-DSpark is a DSpark speculative-decoding draft model with the following features:

  • Target model: LiquidAI/LFM2.5-VL-3B
  • Draft parameters: 279.5M (BF16)
  • Backbone: 4 full attention layers, hidden_size=2048, intermediate_size=6144 with SiLU/SwiGLU, GQA with num_attention_heads=32 / num_key_value_heads=8, head_dim=64
  • Extra heads: Markov head (rank 256) + confidence head
  • Block size: 9 during training; 8 or 9 at inference, depending on hardware
  • Vocabulary: 128,000

[!NOTE] On Apple silicon the drafter is run at block size 8 rather than 9.

Use each drafter checkpoint with its corresponding target model:

DrafterTarget
LFM2.5-VL-3B-DSparkLFM2.5-VL-3B
LFM2.5-VL-3B-DSpark-GGUFLFM2.5-VL-3B-GGUF

📊 Performance

Benchmarks

Speculative decoding is exact under greedy decoding: the target verifies every proposed token, so the generated text is what the target would have produced on its own. Under matched sampling settings at non-zero temperatures, speculative decoding preserves the target model's output distribution. You get the speedup, not a different model.

See LiquidAI/LFM2.5-VL-3B for quality benchmarks.

Draft acceptance

This table reports the mean number of draft tokens accepted per target verification pass at batch size 1 and temperature 0. Higher acceptance generally enables greater acceleration, but the measured speedup also depends on hardware and runtime overhead.

Following MMSpec, the evaluation covers General VQA, Text VQA, Image Captioning, Chart VQA, Complex Reasoning, and Multi-turn Conversation.

Benchmark1×H100 (SGLang, block 9)Apple M5 Max (MLX-VLM, block 8)Apple M3 Ultra (llama.cpp, block 8)
MMMU-Pro4.114.074.19
Multi-turn3.463.243.31
COCO4.574.214.50
CharXiv4.114.344.04
TextVQA3.744.083.58
GQA4.143.773.93

Measured inference speedup

This table reports the resulting runtime performance relative to the same target model without DSpark. Each cell is formatted as decode / end-to-end.

Dataset1×H100 (SGLang)Apple M5 Max (MLX-VLM)Apple M3 Ultra (llama.cpp)
MMMU-Pro2.43× / 1.97×2.93× / 2.62×2.03× / 1.74×
Multi-turn2.04× / 1.83×2.30× / 1.91×1.57× / 1.37×
COCO2.66× / 2.27×3.13× / 2.59×2.14× / 1.77×
CharXiv2.39× / 1.97×2.94× / 1.71×1.87× / 1.56×
TextVQA2.14× / 1.64×2.69× / 1.56×1.64× / 1.33×
GQA2.35× / 1.77×2.67× / 1.93×1.77× / 1.30×

All measurements use 16-bit processing for both the vision encoder and language backbone. The H100 results use SGLang on one H100 80GB in BF16 at batch size 1, temperature 0, and block size 9. The Apple results use FP16 weights at batch size 1, temperature 0, block size 8, and up to 2,048 output tokens. Measurements were collected with Pipette.

🏃 Inference

LFM2.5-VL-3B-DSpark is supported by SGLang for NVIDIA GPUs and MLX-VLM for Apple silicon. For llama.cpp, use the GGUF checkpoint.

SGLang

Requires SGLang v0.5.19 or newer. Launch the target with the draft attached:

python -m sglang.launch_server \
  --model-path LiquidAI/LFM2.5-VL-3B \
  --speculative-algorithm DSPARK \
  --speculative-draft-model-path LiquidAI/L

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms