Model reference · open weights

LFM2.5-DSpark

Available as managed deployment Licence fee LLMs LiquidAI Text gen 1 variants 105k dl/mo

LFM2.5-DSpark is an open-weight language model from LiquidAI. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

Released byLiquidAI
TypeLanguage models
TaskText gen
Runs withllama.cpp
Based onLiquidAI/LFM2.5-2.6B-DSpark
Released2026-08-19
Popularity105k downloads / month
LicenceCommercial licence needed

About

What LFM2.5-DSpark is

src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/2b08LKpev0DNEk6DlnWkY.png" alt="Liquid AI" style="width: 100%; max-width: 100%; height: auto; display: inline-block; margin-bottom: 0.5em; margin-top: 0.5em;" />

Read the full model card

LFM2.5-8B-A1B-DSpark-GGUF

GGUF build of LiquidAI/LFM2.5-8B-A1B-DSpark for llama.cpp (DSpark speculative decoding is in mainline, ggml-org/llama.cpp #25173). This is a standalone draft sidecar: it carries only the drafter (5 attention layers, rank-256 Markov head, confidence head, block size 9). Token embeddings and the LM head are shared from the target model at load time, so it must be paired with a LFM2.5-8B-A1B-GGUF target file.

Find more information about LFM2.5-DSpark in our blog post.

📦 Files

filequantsizenotes
LFM2.5-8B-A1B-DSpark-F16.ggufF16664 MBbest accept length, recommended when memory allows
LFM2.5-8B-A1B-DSpark-Q8_0.ggufQ8_0349 MBaccept length −2% vs F16
LFM2.5-8B-A1B-DSpark-Q4_K_M.ggufQ4_K_M191 MBaccept length −3% vs F16, smallest recommended — sub-4-bit draft quants measurably hurt both accept length and throughput

Draft quantization changes speed only marginally (the drafter is a small share of each cycle); choose by memory budget. The target model quant is the main speed/quality lever and is independent of this file.

🏃 How to run (llama.cpp)

llama-server -m LFM2.5-8B-A1B-F16.gguf \
  -md LFM2.5-8B-A1B-DSpark-F16.gguf \
  --spec-type draft-dspark --spec-draft-n-max 10 --spec-draft-n-min 0 \
  -fa on -ngl 99

The block size is read from the sidecar metadata (n-max is clamped to it). Speculative decoding is exact: the target verifies every proposed token, so greedy output equals the target alone; per-response timings report draft_n / draft_n_accepted.

Other models in the LFM2.5-DSpark GGUF family:

Draft (GGUF)Target (GGUF)
LFM2.5-1.2B-Instruct-DSpark-GGUFLFM2.5-1.2B-Instruct-GGUF
LFM2.5-2.6B-DSpark-GGUFLFM2.5-2.6B-GGUF
LFM2.5-8B-A1B-DSpark-GGUFLFM2.5-8B-A1B-GGUF

📊 Acceptance and benchmarks

See LiquidAI/LFM2.5-8B-A1B-DSpark for acceptance-length tables (H100 and Apple silicon) and target benchmarks.

📬 Contact

  • If you are interested in custom solutions with edge deployment, please contact our sales team.

Citation

@article{liquidAI202626B,
 author  = {Liquid AI},
 title   = {LFM2.5-2.6B: Agents Everywhere},
 journal = {Liquid AI Blog},
 year    = {2026},
 note    = {www.liquid.ai/blog/lfm2-5-2-6b},
}
@article{liquidAI2026dspark,
  author = {Liquid AI},
  title = {LFM2.5-DSpark: Up to 3.2x Faster Inference from H100 to MacBook},
  journal = {Liquid AI Blog},
  year = {2026},
  note = {www.liquid.ai/blog/lfm2.5-dspark},
}

From the published model card. Full card on the HuggingFace links in the sidebar.

How it works

How language models work

Your prompttext / messagesTransformerattention over tokensNext-token loopgenerate + streamResponsetext · tool callsA language model reads your tokens and predicts the next one, again and again, streaming the reply back.

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys lfm2-5-dspark for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (lfm2-5-dspark below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/chat/completions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"lfm2-5-dspark","messages":[{"role":"user","content":"Hello"}]}'

Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms