Model reference · open weights
sapiens2-pretrain-4k is an open-weight embedding model from Meta. sapiens2-pretrain-1b-4k (BF16) weighs 6.7 GB; the smallest configuration that runs it is RTX 3060 12 GB.
What it is
| Released by | Meta |
|---|---|
| Published under | |
| Type | Embedding models |
| Task | Image embed |
| Runs with | sapiens |
| Released | 2026-04-24 |
| Popularity | 0 downloads / month |
| Weights | 6.7 GB (sapiens2-pretrain-1b-4k (BF16), file size) |
| Licence | Its own licence terms |
What it runs on
Weights 6.7 GB (file size) · overhead about 1.1 GB.
| Card | Runs | Counted memory |
|---|---|---|
| RTX 3060 12 GB | fits | 11.6 GB |
| RTX 4060 Ti 16 GB | fits | 15.4 GB |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
Estimates, not measurements: the weights are the build's file size. No cache grows with use; a batch of inputs needs working memory of its own. Counted memory is 92 % of what CUDA reports for the card.
From the model card
Sapiens2 is a family of high-resolution vision transformers pretrained on 1 billion human images — designed for human-centric tasks such as pose estimation, body-part segmentation, surface normals, and pointmaps.
This repository contains the 1B parameter pretrained backbone, trained at 4K resolution (4096 × 3072, H × W) with a window-tokenizer front-end for tractable token counts. It produces dense per-patch features at 4K input.
sapiens2_1b_4k_pretrain.safetensorsInstall the Sapiens2 repo (pip install -e .).
Important: This is the 4K variant. You must instantiate with
use_tokenizer=Trueand pass an input tensor of shape(B, 3, 4096, 3072)(H × W).
import torch
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from sapiens.backbones.standalone.sapiens2 import Sapiens2
# Build the model and load the 4K pretrained checkpoint
model = Sapiens2(
arch="sapiens2_1b",
img_size=(4096, 3072), # H × W
patch_size=16,
use_tokenizer=True, # required for the 4K variant
).eval().cuda()
ckpt_path = hf_hub_download(
repo_id="facebook/sapiens2-pretrain-1b-4k",
filename="sapiens2_1b_4k_pretrain.safetensors",
)
model.load_state_dict(load_file(ckpt_path))
# Forward pass on a single 4K image (RGB; ImageNet normalization recommended)
x = torch.randn(1, 3, 4096, 3072).cuda()
with torch.no_grad():
features = model(x)[0] # dense backbone features
| Field | Value |
|---|---|
| Architecture | Sapiens2 ViT (RoPE, GQA, SwiGLU, RMSNorm, QK-norm) + window tokenizer |
| Backbone parameters | 1.607 B |
| Embedding dim | 1536 |
| Layers | 40 |
| Attention heads | 24 |
| Pretraining resolution | 4096 × 3072 (H × W) |
| Patch size | 16 |
| Window size | 4 |
| Pretraining data | 1B human images |
| Model | Params | FLOPs | Embed dim | Layers | Heads |
|---|---|---|---|---|---|
| Sapiens2-0.1B | 0.114 B | 0.342 T | 768 | 12 | 12 |
| Sapiens2-0.4B | 0.398 B | 1.260 T | 1024 | 24 | 16 |
| Sapiens2-0.8B | 0.818 B | 2.592 T | 1280 | 32 | 16 |
| Sapiens2-1B | 1.462 B | 4.715 T | 1536 | 40 | 24 |
| Sapiens2-1B-4K (this) | 1.607 B | — | 1536 | 40 | 24 |
| Sapiens2-5B | 5.071 B | 15.722 T | 2432 | 56 | 32 |
See the Sapiens2 Collection for all variants and downstream task checkpoints (pose, segmentation, normals, pointmaps).
Released under the Sapiens2 License.
@article{khirodkarsapiens2,
title={Sapiens2},
author={Khirodkar, Rawal and Wen, He and Martinez, Julieta and Dong, Yuan and Su, Zhaoen and Saito, Shunsuke},
journal={arXiv preprint arXiv:2604.21681},
year={2026}
}
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.
How it works