Model reference · open weights

Naive-N0.5-Flash

NEW · this week LLMs NaiveAI Text gen 1 build Open weights 605 dl/mo

Naive-N0.5-Flash is an open-weight language model from NaiveAI. Naive-N0.5-Flash (BF16) weighs 618 GB; the smallest configuration that runs it is 4× B200 180 GB.

What it is

Released byNaiveAI
TypeLanguage models
TaskText gen
Context1,048,576 tokens
Released2026-09-27
Popularity605 downloads / month
Weights618 GB (Naive-N0.5-Flash (BF16), file size)
LicenceOpen weights

What it runs on

Memory and cards for Naive-N0.5-Flash (BF16)

Weights 618 GB (file size) · KV cache 147 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 1.8 GB on a small card · context up to 1,048,576 tokens.

CardRequests at once
8K tokens each
Requests at once
32K tokens each
Longest single
request
Counted
memory
RTX 3060 12 GB … 8× A100 80 GB
14 smaller cards
———
4× B200 180 GB
tensor parallel
256202K176 GB a card
8× H200 141 GB
tensor parallel
17042all 1024K138 GB a card
Memory needed at each load
Requests at once8K tokens each32K tokens each
1621 GB624 GB
5626 GB644 GB
8629 GB658 GB
16639 GB697 GB
32658 GB774 GB
64697 GB929 GB

On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.

Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (grouped-query attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.

From the model card

What NaiveAI says about Naive-N0.5-Flash

[🏠 Homepage][website] · [📰 Technical Blog][blog] · [💻 GitHub][github]

Introduction

Naive-N0.5-Flash is an open-weight 309B MoE model with 15.5B active parameters, built for coding and AI R&D. It supports a native 1M-token context window through a hybrid of Sliding-Window Attention (SWA) and lightweight DeepSeek Sparse Attention (DSA), with no full-attention layers.

Read the full model card

Key Features

  • Native 1M context, without full attention. Naive-N0.5-Flash combines Sliding-Window Attention (SWA) and lightweight DeepSeek Sparse Attention (DSA) with GQA4 at a predominantly 5:1 SWA–DSA layout. The entire network remains local or sparse, with no full-attention layers.
  • AI-optimized inference up to 2,000 tokens/s. NaiveRT, our inference system for Naive-N0.5-Flash, was built and optimized through AI-centered R&D. It combines mega-kernel fusion, Programmatic Dependent Launch (PDL), and speculative decoding, delivering 50 tokens/s per user in Standard mode and up to 2,000 tokens/s in Ultrafast mode. See the NaiveRT case study in the [technical blog][blog] for the implementation and optimization process.
  • Open weights and API. Model weights and inference code are released under the MIT license. API access will also be provided, with pricing set at $0.10 / $0.40 / $0.01 per million tokens for input, output, and cache reads, respectively.

Model Architecture

PropertySpecification
ArchitectureMixture-of-Experts (MoE)
Total parameters309B
Active parameters15.5B
Context lengthNative 1M tokens
Transformer layers48
Attention-layer composition39 SWA layers + 9 DSA layers
Attention mechanismHybrid SWA–DSA
SWA window128 tokens
DSA token selectionTop 2,048 tokens for backbone attention
DSA KV groups4 (GQA4)
Indexer query heads16

Hybrid SWA–DSA Attention

Naive-N0.5-Flash builds on the open-weight MiMo-V2.5 base model, which has a simple architecture with strong foundational capabilities in world knowledge and deep research. Most layers use Sliding-Window Attention (SWA), whose per-token decoding cost does not grow with context length, while a small number of global-attention layers preserve long-range information. At million-token context lengths, however, these global-attention layers account for much of the decoding overhead.

Naive-N0.5-Flash replaces the global-attention layers with DeepSeek Sparse Attention (DSA). A lightweight indexer scores the full history, while the backbone computes attention only over a selected subset of tokens. Although the indexer still scans the full history and the full KV cache is retained, sparse attention substantially reduces attention computation and memory access. Adapting the model to this new attention structure was one objective of continued pretraining.

The network consists of eight six-layer modules. A standard module contains five SWA layers followed by one DSA layer, with the first layer of the first module also replaced by DSA. SWA uses a 128-token window, while DSA selects the top 2,048 tokens for backbone attention. Both attention types incorporate sink bias.

Unlike the original MLA-based DSA implementation, Naive-N0.5-Flash replaces MLA with grouped-query attention (GQA) using four KV groups. For the architecture design process and indexer efficiency comparison, see model architecture in the [technical blog][blog].

Training Overview

Following the architectural changes, Naive-N0.5-Flash completed 3.25T tokens of multi-stage training with a native 1M-token context window: 50B tokens of Indexer Warmup, 3T tokens of Sparse Attention Training, and 200B tokens of Learning Rate Decay. This process adapted the model to its new sparse attention architecture while substantially improving its AI R&D and coding capabilities. See the [technical blog][blog] for training details.

Evaluation Results

[![Coding benchmarks comparing Naive-N0.5-Flash with other models across seven software engineering and agentic tasks.][coding-figure]][coding-pdf]

[![AI R&D benchmarks covering PostTrainBench, MLE-bench-30, PaperBench, SOL-ExecBench, NanoChat AutoResearch, and NanoGPT SpeedRun.][ai-rd-figure]][ai-rd-pdf]

Evaluation setup. Unless otherwise noted, our evaluations of Naive-N0.5-Flash use Claude Code 2.1.207 with a 1M-token context window, temperature 1.0, and top-p 0.95. The harness exposes only basic file I/O and Bash tools.

Sources for reported benchmark scores are as follows:

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms