Model reference · open weights
ovie-ft-re10k is an open-weight image model from kyutai. ovie-ft-re10k (FP32) weighs 286 MB; the smallest configuration that runs it is RTX 3060 12 GB.
What it is
| Released by | kyutai |
|---|---|
| Type | Image models |
| Task | Image edit |
| Parameters (lead) | 143M |
| Runs with | pytorch |
| Based on | kyutai/ovie |
| Released | 2026-09-29 |
| Popularity | 0 downloads / month |
| Weights | 286 MB (ovie-ft-re10k (FP32), file size) |
| Licence | Open weights |
What it runs on
Weights 286 MB (file size) · working memory for one 1024×1024 image about 5.0 GB · overhead about 537 MB.
| Card | One 1024×1024 image | Counted memory |
|---|---|---|
| RTX 3060 12 GB | fits | 11.6 GB |
| RTX 4060 Ti 16 GB | fits | 15.4 GB |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
Estimates, not measurements: the weights are the build's file size; one 1024×1024 image needs about 5 GB of working memory (larger images more). diffusers can also place a pipeline's parts on separate cards (device_map) — not estimated here. Counted memory is 92 % of what CUDA reports for the card.
From the model card
In-the-Wild Monocular Pretraining for Novel View Generation
Part of the OVIE collection.
This is kyutai/ovie fine-tuned for 50k steps on RealEstate10K multi-view data — 2% of the 2M-step pretraining budget. The architecture is identical to the base model; only the weights differ, so it loads through the same code path.
Select this checkpoint when targeting RealEstate10K-style indoor scenes. For zero-shot, in-the-wild use, prefer the base OVIE.
Values reported in the paper for this checkpoint. RealEstate10K test split, 750 scenes, stride 3, 14 target frames (in-domain for this checkpoint):
| PSNR ↑ | SSIM ↑ | LPIPS ↓ | FID ↓ | MEt3R ↓ |
|---|---|---|---|---|
| 21.9 | 0.695 | 0.195 | 5.59 | 0.029 |
Fine-tuning selects a domain: on DL3DV, where this checkpoint is out-of-domain, the base model is stronger on pixel fidelity, perceptual similarity and FID (PSNR 14.8 vs 14.1, FID 13.6 vs 24.66), while this checkpoint keeps the better multi-view consistency (MEt3R 0.040 vs 0.078). For comparison, the same architecture trained from scratch on RealEstate10K alone reaches only 19.05 PSNR, 2.85 dB below this fine-tune — the gap comes from monocular pretraining.
import torch
from models.models import OVIEModel
from utils.pose_enc import extri_intri_to_pose_encoding
from torchvision.transforms import ToTensor
from PIL import Image
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = OVIEModel.from_pretrained("kyutai/ovie-ft-re10k").to(device)
model.eval()
image_size = model.image_size # 256
img_pil = Image.open("image.jpg").convert("RGB").resize((image_size, image_size))
img_tensor = ToTensor()(img_pil).unsqueeze(0).to(device)
extrinsics = torch.tensor([[[1.0, 0.0, 0.0, -1.25],
[0.0, 1.0, 0.0, 0.5],
[0.0, 0.0, 1.0, -2.0]]], device=device)
dummy_intrinsics = torch.zeros(1, 1, 3, 3, device=device)
camera = extri_intri_to_pose_encoding(
extrinsics=extrinsics.unsqueeze(0),
intrinsics=dummy_intrinsics,
image_size_hw=(image_size, image_size),
)
cam_token = camera[..., :7].squeeze(0)
with torch.no_grad():
pred = model(x=img_tensor, cam_params=cam_token) # (1, 3, 256, 256) in [0, 1]
To reproduce the metrics with the repository's evaluation script:
uv run python evaluate.py \
--dataset_path /PATH/TO/RE10K/TEST \
--config_path configs/config_ovie.yaml \
--from_pretrained kyutai/ovie-ft-re10k \
--stride 3 --num_target_frames 14
See the repository for installation, data preprocessing, and the full benchmark table.
@misc{ovie2026,
title={One View Is Enough: In-the-Wild Monocular Pretraining for Novel View Generation},
author={Adrien Ramanana Rahary and Nicolas Dufour and Patrick Perez and David Picard},
year={2026},
eprint={2603.23488},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2603.23488},
}
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.
How it works