Model reference · open weights

sapiens2-pretrain-4k

Embeddings facebook Image embed 1 build Its own licence terms 0 dl/mo

sapiens2-pretrain-4k is an open-weight embedding model from Meta. sapiens2-pretrain-1b-4k (BF16) weighs 6.7 GB; the smallest configuration that runs it is RTX 3060 12 GB.

What it is

Released byMeta
Published underfacebook
TypeEmbedding models
TaskImage embed
Runs withsapiens
Released2026-04-24
Popularity0 downloads / month
Weights6.7 GB (sapiens2-pretrain-1b-4k (BF16), file size)
LicenceIts own licence terms

What it runs on

Memory and cards for sapiens2-pretrain-1b-4k (BF16)

Weights 6.7 GB (file size) · overhead about 1.1 GB.

CardRunsCounted
memory
RTX 3060 12 GBfits11.6 GB
RTX 4060 Ti 16 GBfits15.4 GB
RTX 3090 24 GBfits23.4 GB
RTX 4090 24 GBfits23.4 GB
RTX 5090 32 GBfits31.0 GB
L40S 48 GBfits44.0 GB
A100 80 GBfits78.2 GB
H100 80 GBfits78.1 GB
RTX PRO 6000 Blackwell 96 GBfits93.8 GB
DGX Spark (GB10) 128 GB unifiedfits107 GB
H200 141 GBfits138 GB
B200 180 GBfits176 GB

Estimates, not measurements: the weights are the build's file size. No cache grows with use; a batch of inputs needs working memory of its own. Counted memory is 92 % of what CUDA reports for the card.

From the model card

What Meta says about sapiens2-pretrain-4k

Sapiens2 is a family of high-resolution vision transformers pretrained on 1 billion human images — designed for human-centric tasks such as pose estimation, body-part segmentation, surface normals, and pointmaps.

This repository contains the 1B parameter pretrained backbone, trained at 4K resolution (4096 × 3072, H × W) with a window-tokenizer front-end for tractable token counts. It produces dense per-patch features at 4K input.

Read the full model card

Model Details

  • Developed by: Meta
  • Model type: Vision Transformer
  • License: Sapiens2 License
  • Task: pretrain (4K resolution)
  • Format: safetensors
  • File: sapiens2_1b_4k_pretrain.safetensors

Quick Start

Install the Sapiens2 repo (pip install -e .).

Important: This is the 4K variant. You must instantiate with use_tokenizer=True and pass an input tensor of shape (B, 3, 4096, 3072) (H × W).

import torch
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from sapiens.backbones.standalone.sapiens2 import Sapiens2

# Build the model and load the 4K pretrained checkpoint
model = Sapiens2(
    arch="sapiens2_1b",
    img_size=(4096, 3072),   # H × W
    patch_size=16,
    use_tokenizer=True,      # required for the 4K variant
).eval().cuda()

ckpt_path = hf_hub_download(
    repo_id="facebook/sapiens2-pretrain-1b-4k",
    filename="sapiens2_1b_4k_pretrain.safetensors",
)
model.load_state_dict(load_file(ckpt_path))

# Forward pass on a single 4K image (RGB; ImageNet normalization recommended)
x = torch.randn(1, 3, 4096, 3072).cuda()
with torch.no_grad():
    features = model(x)[0]  # dense backbone features

Model Card

FieldValue
ArchitectureSapiens2 ViT (RoPE, GQA, SwiGLU, RMSNorm, QK-norm) + window tokenizer
Backbone parameters1.607 B
Embedding dim1536
Layers40
Attention heads24
Pretraining resolution4096 × 3072 (H × W)
Patch size16
Window size4
Pretraining data1B human images

Sapiens2 Family

ModelParamsFLOPsEmbed dimLayersHeads
Sapiens2-0.1B0.114 B0.342 T7681212
Sapiens2-0.4B0.398 B1.260 T10242416
Sapiens2-0.8B0.818 B2.592 T12803216
Sapiens2-1B1.462 B4.715 T15364024
Sapiens2-1B-4K (this)1.607 B—15364024
Sapiens2-5B5.071 B15.722 T24325632

See the Sapiens2 Collection for all variants and downstream task checkpoints (pose, segmentation, normals, pointmaps).

Intended Use

  • Feature extraction for human-centric downstream tasks at 4K resolution
  • Initialization for fine-tuning high-resolution task heads (pose, segmentation, normals, pointmap)
  • Research on human-centric vision at high resolution

License

Released under the Sapiens2 License.

Citation

@article{khirodkarsapiens2,
  title={Sapiens2},
  author={Khirodkar, Rawal and Wen, He and Martinez, Julieta and Dong, Yuan and Su, Zhaoen and Saito, Shunsuke},
  journal={arXiv preprint arXiv:2604.21681},
  year={2026}
}

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

How it works

How embedding models work

Your textsentence / documentEncodermaps meaningVectorlist of numbersAn embedding model turns text into a vector, so similar meanings sit close together — the basis of search and RAG.
© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms