Model reference · open weights

DeepSeek-Flash-DSpark

DeepSeek-Flash-DSpark is an open-weight language model from nvidia, listed in the AxForge catalogue. AxForge can bring it up on EU-owned hardware for you on request — with the licence handled where one is required.

LLMs nvidia 1 variants 17k downloads/mo
Request this model on EU hardware All served models Not on the shared API today — deployed on request.

About

What DeepSeek-Flash-DSpark is

Model Overview Description: The NVIDIA DeepSeek-V4-Flash-nvfp4-DSpark model is a quantized version of DeepSeek AI's DeepSeek-V4-Flash model, an autoregressive Mixture-of-Experts language model that uses an optimized Transformer architecture with hybrid attention (Compressed Sparse Attention and Heavily Compressed Attention) and Manifold-Constrained Hyper-Connections, packaged in a single checkpoint together with DeepSeek's official DSpark speculative decoding module. For more information, refer to the DeepSeek-V4-Flash model card and the DeepSeek-V4-Flash-DSpark model card. The NVIDIA DeepSeek-V4-Flash-nvfp4-DSpark model is quantized with Model Optimizer. Note: DeepSeek-V4-Flash-nvfp4-DSpark is not a new model. It is the NVFP4 backbone of nvidia/DeepSeek-V4-Flash-NVFP4 with DeepSeek's official DSpark speculative decoding module attached, so a single checkpoint serves as both target and draft model. For more details on DSpark, refer to: https://github.com/deepseek-ai/DeepSpec This model is ready for commercial/non-commercial use. <br Third-Party Community Consideration This model is not owned or developed by NVIDIA. This model has been developed and built to a third-party’s requirements for this application and use case; see link to Non-NVIDIA (DeepSeek-V4-Flash-DSpark) Model Card. References - Nvidia Model Optimizer: https://github.com/NVIDIA/Model-Optimizer - DeepSeek-V4-Flash base model card - DeepSeek-V4-Flash-DSpark model card - nvidia/DeepSeek-V4-Flash-NVFP4 model card - DeepSeek-V4 technical report - DeepSeek DSpec reference implementation: https://github.com/deepseek-ai/DeepSpec License/Terms of Use: MIT Deployment Geography: Global <br Use Case: DeepSeek V4 is well-suited for advanced reasoning, agentic AI applications, tool use scenarios, and complex problem-solving in domains such as mathematics, software engineering, and enterprise AI assistants. <br Release Date: Hugging Face 07/22/2026 via https://huggingface.co/nvidia/DeepSeek-V4-Flash-nvfp4-DSpark <br Model Architecture: Architecture Type: Transformers <br Network Architecture: Mixture-of-Experts (MoE) with Hybrid Attention (Compressed Sparse Attention + Heavily Compressed Attention) <br Number of

Summarised from the published model card. Read the full card on the HuggingFace links below.

Specifications

What it is

Makernvidia
TypeLanguage models
Parameters (lead)304.2B
Variants1
Runs withModel Optimizer
Based ondeepseek-ai/DeepSeek-V4-Flash, deepseek-ai/DeepSeek-V4-Flash-DSpark
Released2026-07-22
Popularity17k downloads / month
Likes22
LicenceOpen weights

How it works

How language models work

Your prompttext / messagesTransformerattention over tokensNext-token loopgenerate + streamResponsetext · tool callsA language model reads your tokens and predicts the next one, again and again, streaming the reply back.

Variants

Sizes & precisions

Open weights ship in several sizes and precisions. One page, all the variants — pick the one that fits your GPU. VRAM figures are estimates from model size.

VariantParamsPrecisionVRAMFits 16 GBWeights
DeepSeek-V4-Flash-nvfp4-DSpark304.2BNVFP4Weights ↗

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys nvidia-deepseek-flash-dspark for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (nvidia-deepseek-flash-dspark below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/chat/completions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"nvidia-deepseek-flash-dspark","messages":[{"role":"user","content":"Hello"}]}'

Details

Languages, data & research

Tags

Model Optimizer safetensors deepseek_v4 nvidia ModelOpt DeepSeekV4 quantized NVFP4 nvfp4 speculative-decoding DSpark text-generation conversational 8-bit

Papers

Licence

Open weights

Open weights under mit — commercial use is permitted. Deploy it on AxForge EU hardware on request. Read the licence ↗

Sources

Weights & code

Want DeepSeek-Flash-DSpark on EU-owned hardware?

Request this model on EU hardware See what’s served now

Explore

More language models

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms