Model reference · open weights
Step-5 is an open-weight language model from TypeSafeAI. Step-5-Preview-BF16 (BF16) weighs 1215 GB; the smallest configuration that runs it is 8× B200 180 GB.
What it is
| Released by | TypeSafeAI |
|---|---|
| Type | Language models |
| Task | Vision + text |
| Parameters (lead) | 604.3B |
| Context | 1,048,576 tokens |
| Runs with | transformers |
| Released | 2026-09-20 |
| Popularity | 524 downloads / month |
| Weights | 1215 GB (Step-5-Preview-BF16 (BF16), file size) |
| Licence | Its own licence terms |
What it runs on
Weights 1215 GB (file size) · KV cache 983 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · plus 18.1 GB a request for its sliding-window layers · runtime overhead from 3.0 GB on a small card · context up to 1,048,576 tokens.
| Card | Requests at once 8K tokens each | Requests at once 32K tokens each | Longest single request | Counted memory |
|---|---|---|---|---|
| RTX 3060 12 GB … 8× H200 141 GB 15 smaller cards | — | — | — | |
| 8× B200 180 GB tensor parallel | 1 | — | — | 176 GB a card |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 1244 GB | 1268 GB |
| 5 | 1349 GB | 1470 GB |
| 8 | 1427 GB | 1621 GB |
| 16 | 1637 GB | 2023 GB |
| 32 | 2055 GB | 2828 GB |
| 64 | 2893 GB | 4439 GB |
On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.
Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (attention with sliding-window layers); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.
From the model card
We are excited to release Step-5-Preview, our flagship foundation model for real-world agentic work. It is a 600B-parameter sparse Mixture-of-Experts model with 27B active parameters, a 1M-token context window, and native support for text, image, and video inputs. Try it via our API, or deploy locally with vLLM / SGLang.
Step-5-Preview is StepFun's flagship foundation model, designed from the ground up for real-world agentic tasks. It targets professional domains such as AI coding, software engineering, professional knowledge work, and financial analysis.
StepFun's core philosophy for Step 5 is the "Pareto Frontier" — achieving the optimal balance between intelligence and cost. While previous scaling efforts focused on trading more compute for stronger intelligence, the next phase requires improving the efficiency of converting compute into intelligence.
• 600B total parameters, only 27B active — near-frontier performance at a fraction of the compute. • 1M-token context window without proportional cost increases. • Competitive benchmark scores against models with 3–5× more parameters. • Built for agents — long-horizon reasoning, tool use, and autonomous execution.
Step-5-Preview represents a generational leap, with StepFun skipping the entire Step 4.x line entirely, going directly from Step-3.7-Flash to Step 5. This decision reflects the magnitude of improvement achieved in this release.
low, medium, high / xhigh.TypeSafeAI/Step-5-Preview-BF16.Step-5-Preview uses a 92-layer Transformer with a narrow-deep configuration. This design is specifically intended to create longer information propagation paths for implicit multi-hop reasoning during long prefill operations.
To handle the 1M-token context window efficiently, Step-5-Preview introduces Sparse GQA with block-wise token merging. This mechanism uses sparse indexing to select only historical information relevant to the current task, reducing the number of tokens that actually enter attention computation. StepFun states this cuts indexer and top-k selection costs to approximately one-eighth of a denser baseline.
Step 5 Preview achieves near-frontier performance with 600B total parameters but only
The model incorporates a unified multimodal encoder that processes text, images, and video frames into a shared latent space. Video is sampled at adaptive frame rates and encoded with temporal attention, allowing the model to understand motion and long-range dependencies in screen recordings, demonstrations, and real-world footage.
| Category | Specification |
|---|---|
| Model Name | Step-5-Preview |
| Developer | StepFun |
| Architecture | Sparse Mixture-of-Experts (MoE) |
| Total Parameters | 600B |
| Active Parameters | 27B per token (~4.5% sparsity) |
| Layers | 92 (narrow-deep Transformer) |
| Context Window | 1,000,000 tokens |
| Attention | Sparse GQA with block-wise token merging |
| Input Modalities | Text, Image, Video |
| Output Modalities | Text |
| Video Formats | MP4, QuickTime, Matroska (≤128 MB, ≤5 min recommended) |
| Reasoning Effort | low / medium / high (xhigh) |
| Tool Calling | Parallel, strict JSON schema |
| Intelligence Index | 44 (Artificial Analysis v4.3.2) |
| Open Weights | BF16 checkpoint available now |
| API Availability | Immediate (OpenAI-compatible) |
| License | StepFun Community License |
Step-5-Preview was trained on a massive, carefully curated corpus spanning:
The data mixture was optimized for long-horizon reasoning and tool use, with a strong emphasis on real-world professional tasks. All data was filtered for quality, safety, and li
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.