Model reference · open weights
ovie-512 is an open-weight image model from kyutai. ovie-512 (FP32) weighs 290 MB; the smallest configuration that runs it is RTX 3060 12 GB.
What it is
| Released by | kyutai |
|---|---|
| Type | Image models |
| Task | Image edit |
| Parameters (lead) | 145M |
| Runs with | pytorch |
| Released | 2026-09-29 |
| Popularity | 0 downloads / month |
| Weights | 290 MB (ovie-512 (FP32), file size) |
| Licence | Open weights |
What it runs on
Weights 290 MB (file size) · working memory for one 1024×1024 image about 5.0 GB · overhead about 537 MB.
| Card | One 1024×1024 image | Counted memory |
|---|---|---|
| RTX 3060 12 GB | fits | 11.6 GB |
| RTX 4060 Ti 16 GB | fits | 15.4 GB |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
Estimates, not measurements: the weights are the build's file size; one 1024×1024 image needs about 5 GB of working memory (larger images more). diffusers can also place a pipeline's parts on separate cards (device_map) — not estimated here. Counted memory is 92 % of what CUDA reports for the card.
From the model card
In-the-Wild Monocular Pretraining for Novel View Generation
Part of the OVIE collection.
OVIE is resolution-agnostic by construction: the encoder and decoder are fully convolutional and the ViT bottleneck regenerates its positional encodings for the new grid. This checkpoint is the base model retrained at 512×512 under the identical recipe (same in-the-wild mix with MoGe-2 pseudo-pairs), for 250K steps at global batch size 256, changing only the resolution.
At a like-for-like 256×256 evaluation it matches or improves on the base model on PSNR, SSIM, LPIPS and MEt3R on both benchmarks, at a cost of roughly one FID point, which the paper attributes to the training mix — part of it sits below 512 and is upscaled, carrying little true high-frequency detail. Inference stays a single feed-forward pass with no per-scene construction: 41.6 ms per view on an H100 against 8.6 ms at 256×256.
eval@256 brings every output to 256×256 before scoring (like-for-like with the base model); eval@512 scores at native resolution and is given for reference only, since SSIM and LPIPS are resolution-dependent.
| Model | Benchmark | PSNR ↑ | SSIM ↑ | LPIPS ↓ | FID ↓ | MEt3R ↓ |
|---|---|---|---|---|---|---|
OVIE-512, eval@256 | RealEstate10K | 19.1 | 0.611 | 0.283 | 7.62 | 0.034 |
OVIE-512, eval@256 | DL3DV | 15.2 | 0.389 | 0.464 | 14.8 | 0.077 |
OVIE-512, eval@512 | RealEstate10K | 18.9 | 0.658 | 0.344 | — | — |
OVIE-512, eval@512 | DL3DV | 15.1 | 0.475 | 0.509 | — | — |
For reference, the base OVIE (256×256) scores 18.8 / 0.602 / 0.279 / 6.74 / 0.035 on RealEstate10K and 14.8 / 0.369 / 0.464 / 13.6 / 0.078 on DL3DV.
import torch
from models.models import OVIEModel
from utils.pose_enc import extri_intri_to_pose_encoding
from torchvision.transforms import ToTensor
from PIL import Image
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = OVIEModel.from_pretrained("kyutai/ovie-512").to(device)
model.eval()
image_size = model.image_size # 512, read from the saved config
img_pil = Image.open("image.jpg").convert("RGB").resize((image_size, image_size))
img_tensor = ToTensor()(img_pil).unsqueeze(0).to(device)
extrinsics = torch.tensor([[[1.0, 0.0, 0.0, -1.25],
[0.0, 1.0, 0.0, 0.5],
[0.0, 0.0, 1.0, -2.0]]], device=device)
dummy_intrinsics = torch.zeros(1, 1, 3, 3, device=device)
camera = extri_intri_to_pose_encoding(
extrinsics=extrinsics.unsqueeze(0),
intrinsics=dummy_intrinsics,
image_size_hw=(image_size, image_size),
)
cam_token = camera[..., :7].squeeze(0)
with torch.no_grad():
pred = model(x=img_tensor, cam_params=cam_token) # (1, 3, 512, 512) in [0, 1]
Note the output resolution is fixed at training time: unlike methods that render from an explicit 3D representation, OVIE-512 cannot produce arbitrary resolutions at inference.
To evaluate with the repository's script, pass --image_size 512:
uv run python evaluate.py \
--dataset_path /PATH/TO/RE10K/TEST \
--config_path configs/config_ovie.yaml \
--from_pretrained kyutai/ovie-512 \
--image_size 512 --stride 3 --num_target_frames 14
See the repository for installation, data preprocessing, and the full benchmark table.
@misc{ovie2026,
title={One View Is Enough: In-the-Wild Monocular Pretraining for Novel View Generation},
author={Adrien Ramanana Rahary and Nicolas Dufour and Patrick Perez and David Picard},
year={2026},
eprint={2603.23488},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2603.23488},
}
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.
How it works