Model reference · open weights

Step-5

LLMs TypeSafeAI Vision + text 1 build Its own licence terms 524 dl/mo

Step-5 is an open-weight language model from TypeSafeAI. Step-5-Preview-BF16 (BF16) weighs 1215 GB; the smallest configuration that runs it is 8× B200 180 GB.

What it is

Released byTypeSafeAI
TypeLanguage models
TaskVision + text
Parameters (lead)604.3B
Context1,048,576 tokens
Runs withtransformers
Released2026-09-20
Popularity524 downloads / month
Weights1215 GB (Step-5-Preview-BF16 (BF16), file size)
LicenceIts own licence terms

What it runs on

Memory and cards for Step-5-Preview-BF16 (BF16)

Weights 1215 GB (file size) · KV cache 983 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · plus 18.1 GB a request for its sliding-window layers · runtime overhead from 3.0 GB on a small card · context up to 1,048,576 tokens.

CardRequests at once
8K tokens each
Requests at once
32K tokens each
Longest single
request
Counted
memory
RTX 3060 12 GB … 8× H200 141 GB
15 smaller cards
———
8× B200 180 GB
tensor parallel
1——176 GB a card
Memory needed at each load
Requests at once8K tokens each32K tokens each
11244 GB1268 GB
51349 GB1470 GB
81427 GB1621 GB
161637 GB2023 GB
322055 GB2828 GB
642893 GB4439 GB

On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.

Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (attention with sliding-window layers); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.

From the model card

What TypeSafeAI says about Step-5

We are excited to release Step-5-Preview, our flagship foundation model for real-world agentic work. It is a 600B-parameter sparse Mixture-of-Experts model with 27B active parameters, a 1M-token context window, and native support for text, image, and video inputs. Try it via our API, or deploy locally with vLLM / SGLang.


Read the full model card

📖 Table of Contents


🚀 Introduction

Step-5-Preview is StepFun's flagship foundation model, designed from the ground up for real-world agentic tasks. It targets professional domains such as AI coding, software engineering, professional knowledge work, and financial analysis.

StepFun's core philosophy for Step 5 is the "Pareto Frontier" — achieving the optimal balance between intelligence and cost. While previous scaling efforts focused on trading more compute for stronger intelligence, the next phase requires improving the efficiency of converting compute into intelligence.

• 600B total parameters, only 27B active — near-frontier performance at a fraction of the compute. • 1M-token context window without proportional cost increases. • Competitive benchmark scores against models with 3–5× more parameters. • Built for agents — long-horizon reasoning, tool use, and autonomous execution.

Step-5-Preview represents a generational leap, with StepFun skipping the entire Step 4.x line entirely, going directly from Step-3.7-Flash to Step 5. This decision reflects the magnitude of improvement achieved in this release.


✨ Key Features

  • Sparse Mixture-of-Experts (MoE): 600B total parameters, 27B active per token (~4.5% sparsity).
  • 1M-Token Context Window: Equivalent to ~1,500 A4 pages, enabled by Sparse GQA.
  • Multimodal Input: Text, image, and video (MP4, QuickTime, Matroska; ≤128 MB; ≤5 min recommended).
  • Configurable Reasoning Effort: low, medium, high / xhigh.
  • Parallel Tool Calling: Natively supported for agentic workflows.
  • Strict JSON Schema Output: Reliable integration into structured systems.
  • OpenAI-Compatible API: Available via Step API and third-party gateways.
  • Open Weights: BF16 checkpoint available now under TypeSafeAI/Step-5-Preview-BF16.

🏗️ Model Architecture

92-Layer "Narrow but Deep" Design

Step-5-Preview uses a 92-layer Transformer with a narrow-deep configuration. This design is specifically intended to create longer information propagation paths for implicit multi-hop reasoning during long prefill operations.

Sparse Grouped-Query Attention (GQA) with Block-Wise Token Merging

To handle the 1M-token context window efficiently, Step-5-Preview introduces Sparse GQA with block-wise token merging. This mechanism uses sparse indexing to select only historical information relevant to the current task, reducing the number of tokens that actually enter attention computation. StepFun states this cuts indexer and top-k selection costs to approximately one-eighth of a denser baseline.

Step 5 Preview achieves near-frontier performance with 600B total parameters but only

Multimodal Encoder

The model incorporates a unified multimodal encoder that processes text, images, and video frames into a shared latent space. Video is sampled at adaptive frame rates and encoded with temporal attention, allowing the model to understand motion and long-range dependencies in screen recordings, demonstrations, and real-world footage.


📋 Model Specifications

CategorySpecification
Model NameStep-5-Preview
DeveloperStepFun
ArchitectureSparse Mixture-of-Experts (MoE)
Total Parameters600B
Active Parameters27B per token (~4.5% sparsity)
Layers92 (narrow-deep Transformer)
Context Window1,000,000 tokens
AttentionSparse GQA with block-wise token merging
Input ModalitiesText, Image, Video
Output ModalitiesText
Video FormatsMP4, QuickTime, Matroska (≤128 MB, ≤5 min recommended)
Reasoning Effortlow / medium / high (xhigh)
Tool CallingParallel, strict JSON schema
Intelligence Index44 (Artificial Analysis v4.3.2)
Open WeightsBF16 checkpoint available now
API AvailabilityImmediate (OpenAI-compatible)
LicenseStepFun Community License

📚 Training Data

Step-5-Preview was trained on a massive, carefully curated corpus spanning:

  • Code repositories from multiple languages (Python, C++, Rust, JavaScript, Go, etc.)
  • Technical documentation, API references, and software engineering forums
  • Scientific papers in computer science, mathematics, physics, and finance
  • Financial reports, earnings calls, and market analyses
  • Multimodal data including screenshots, UI mockups, video tutorials, and screen recordings
  • Agentic trajectories from simulated and real tool-use environments

The data mixture was optimized for long-horizon reasoning and tool use, with a strong emphasis on real-world professional tasks. All data was filtered for quality, safety, and li

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms