Model reference · open weights
sapiens2-pretrain is an open-weight embedding model from facebook. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.
Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.
What it is
| Released by | Meta |
|---|---|
| Published under | |
| Type | Embedding models |
| Task | Image embed |
| Parameters (lead) | 395M |
| Runs with | sapiens2 |
| Released | 2026-04-23 |
| Popularity | 1k downloads / month |
| Licence | Commercial licence needed |
About
Sapiens2 is a family of high-resolution vision transformers pretrained on 1 billion human images — designed for human-centric tasks such as pose estimation, body-part segmentation, surface normals, and pointmaps.
This repository contains the 0.4B parameter pretrained backbone. It produces dense per-patch features suitable for fine-tuning downstream task heads.
sapiens2_0.4b_pretrain.safetensorsInstall the Sapiens2 repo (pip install -e .).
import torch
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from sapiens.backbones.standalone.sapiens2 import Sapiens2
# Build the model and load the pretrained checkpoint
model = Sapiens2(arch="sapiens2_0.4b", img_size=(1024, 768), patch_size=16).eval().cuda() # img_size is (H, W)
ckpt_path = hf_hub_download(repo_id="facebook/sapiens2-pretrain-0.4b", filename="sapiens2_0.4b_pretrain.safetensors")
model.load_state_dict(load_file(ckpt_path))
# Forward pass on a single image (RGB; ImageNet normalization recommended)
x = torch.randn(1, 3, 1024, 768).cuda()
with torch.no_grad():
features = model(x)[0] # dense backbone features: (B, num_tokens, embed_dim)
| Field | Value |
|---|---|
| Architecture | Sapiens2 ViT (RoPE, GQA, SwiGLU, RMSNorm, QK-norm) |
| Parameters | 0.398 B |
| FLOPs | 1.260 T |
| Embedding dim | 1024 |
| Layers | 24 |
| Attention heads | 16 |
| Pretraining resolution | 1024 × 768 (H × W) |
| Patch size | 16 |
| Pretraining data | 1B human images |
| Model | Params | FLOPs | Embed dim | Layers | Heads |
|---|---|---|---|---|---|
| Sapiens2-0.1B | 0.114 B | 0.342 T | 768 | 12 | 12 |
| Sapiens2-0.4B (this) | 0.398 B | 1.260 T | 1024 | 24 | 16 |
| Sapiens2-0.8B | 0.818 B | 2.592 T | 1280 | 32 | 16 |
| Sapiens2-1B | 1.462 B | 4.715 T | 1536 | 40 | 24 |
| Sapiens2-1B-4K | 1.607 B | — | 1536 | 40 | 24 |
| Sapiens2-5B | 5.071 B | 15.722 T | 2432 | 56 | 32 |
See the Sapiens2 Collection for all variants and downstream task checkpoints (pose, segmentation, normals, pointmaps).
Released under the Sapiens2 License.
@article{khirodkarsapiens2,
title={Sapiens2},
author={Khirodkar, Rawal and Wen, He and Martinez, Julieta and Dong, Yuan and Su, Zhaoen and Saito, Shunsuke},
journal={arXiv preprint arXiv:2604.21681},
year={2026}
}
From the published model card. Full card on the HuggingFace links in the sidebar.
Using it via the API
Once AxForge deploys sapiens2-pretrain for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (sapiens2-pretrain below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/embeddings \
-H "Authorization: Bearer $AXFORGE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"sapiens2-pretrain","input":"text to embed"}'
Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.