Model reference · open weights
LFM2.5-DSpark is an open-weight language model from LiquidAI. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.
Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.
What it is
| Released by | LiquidAI |
|---|---|
| Type | Language models |
| Task | Text gen |
| Runs with | llama.cpp |
| Based on | LiquidAI/LFM2.5-2.6B-DSpark |
| Released | 2026-08-19 |
| Popularity | 105k downloads / month |
| Licence | Commercial licence needed |
About
src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/2b08LKpev0DNEk6DlnWkY.png" alt="Liquid AI" style="width: 100%; max-width: 100%; height: auto; display: inline-block; margin-bottom: 0.5em; margin-top: 0.5em;" />
GGUF build of LiquidAI/LFM2.5-8B-A1B-DSpark for llama.cpp (DSpark speculative decoding is in mainline, ggml-org/llama.cpp #25173).
This is a standalone draft sidecar: it carries only the drafter (5 attention layers, rank-256 Markov head, confidence head, block size 9). Token embeddings and the LM head are shared from the target model at load time, so it must be paired with a LFM2.5-8B-A1B-GGUF target file.
Find more information about LFM2.5-DSpark in our blog post.
| file | quant | size | notes |
|---|---|---|---|
LFM2.5-8B-A1B-DSpark-F16.gguf | F16 | 664 MB | best accept length, recommended when memory allows |
LFM2.5-8B-A1B-DSpark-Q8_0.gguf | Q8_0 | 349 MB | accept length −2% vs F16 |
LFM2.5-8B-A1B-DSpark-Q4_K_M.gguf | Q4_K_M | 191 MB | accept length −3% vs F16, smallest recommended — sub-4-bit draft quants measurably hurt both accept length and throughput |
Draft quantization changes speed only marginally (the drafter is a small share of each cycle); choose by memory budget. The target model quant is the main speed/quality lever and is independent of this file.
llama-server -m LFM2.5-8B-A1B-F16.gguf \
-md LFM2.5-8B-A1B-DSpark-F16.gguf \
--spec-type draft-dspark --spec-draft-n-max 10 --spec-draft-n-min 0 \
-fa on -ngl 99
The block size is read from the sidecar metadata (n-max is clamped to it). Speculative decoding is exact: the target verifies every proposed token, so greedy output equals the target alone; per-response timings report draft_n / draft_n_accepted.
Other models in the LFM2.5-DSpark GGUF family:
| Draft (GGUF) | Target (GGUF) |
|---|---|
| LFM2.5-1.2B-Instruct-DSpark-GGUF | LFM2.5-1.2B-Instruct-GGUF |
| LFM2.5-2.6B-DSpark-GGUF | LFM2.5-2.6B-GGUF |
| LFM2.5-8B-A1B-DSpark-GGUF | LFM2.5-8B-A1B-GGUF |
See LiquidAI/LFM2.5-8B-A1B-DSpark for acceptance-length tables (H100 and Apple silicon) and target benchmarks.
@article{liquidAI202626B,
author = {Liquid AI},
title = {LFM2.5-2.6B: Agents Everywhere},
journal = {Liquid AI Blog},
year = {2026},
note = {www.liquid.ai/blog/lfm2-5-2-6b},
}
@article{liquidAI2026dspark,
author = {Liquid AI},
title = {LFM2.5-DSpark: Up to 3.2x Faster Inference from H100 to MacBook},
journal = {Liquid AI Blog},
year = {2026},
note = {www.liquid.ai/blog/lfm2.5-dspark},
}
From the published model card. Full card on the HuggingFace links in the sidebar.
How it works
Using it via the API
Once AxForge deploys lfm2-5-dspark for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (lfm2-5-dspark below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/chat/completions \
-H "Authorization: Bearer $AXFORGE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"lfm2-5-dspark","messages":[{"role":"user","content":"Hello"}]}'
Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.