Model reference · open weights
ovie-ft-dl3dv is an open-weight image model from kyutai. ovie-ft-dl3dv (FP32) weighs 286 MB; the smallest configuration that runs it is RTX 3060 12 GB.
What it is
| Released by | kyutai |
|---|---|
| Type | Image models |
| Task | Image edit |
| Parameters (lead) | 143M |
| Runs with | pytorch |
| Based on | kyutai/ovie |
| Released | 2026-09-29 |
| Popularity | 0 downloads / month |
| Weights | 286 MB (ovie-ft-dl3dv (FP32), file size) |
| Licence | Open weights |
What it runs on
Weights 286 MB (file size) · working memory for one 1024×1024 image about 5.0 GB · overhead about 537 MB.
| Card | One 1024×1024 image | Counted memory |
|---|---|---|
| RTX 3060 12 GB | fits | 11.6 GB |
| RTX 4060 Ti 16 GB | fits | 15.4 GB |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
Estimates, not measurements: the weights are the build's file size; one 1024×1024 image needs about 5 GB of working memory (larger images more). diffusers can also place a pipeline's parts on separate cards (device_map) — not estimated here. Counted memory is 92 % of what CUDA reports for the card.
From the model card
In-the-Wild Monocular Pretraining for Novel View Generation
Part of the OVIE collection.
This is kyutai/ovie fine-tuned for 50k steps on DL3DV, the companion to the RealEstate10K fine-tune. The architecture is identical to the base model; only the weights differ.
It is the control that shows fine-tuning selects a domain: applying the same protocol on DL3DV markedly improves in-domain reconstruction relative to the base model, but costs distributional realism and generalization.
| Checkpoint | DL3DV PSNR ↑ | DL3DV SSIM ↑ | DL3DV LPIPS ↓ | DL3DV FID ↓ | DL3DV MEt3R ↓ | RealEstate10K PSNR ↑ | RealEstate10K FID ↓ |
|---|---|---|---|---|---|---|---|
| base OVIE (out-of-domain on DL3DV) | 14.8 | 0.369 | 0.464 | 13.6 | 0.078 | 18.8 | 6.74 |
| this checkpoint (in-domain on DL3DV) | 17.07 | 0.433 | 0.368 | 17.24 | 0.059 | 19.34 | 7.14 |
Fine-tuning raises DL3DV PSNR by 2.3 dB while raising FID from 13.6 to 17.24: the model becomes a more faithful per-view reconstructor at a slight cost to global distributional realism. Notably it retains — and on RealEstate10K slightly improves — cross-domain performance even though that benchmark is now out-of-domain (PSNR 19.34 vs 18.8), so the fine-tune sharpens in-domain fidelity without eroding generalization here.
Prefer the base OVIE for in-the-wild zero-shot use and the RealEstate10K fine-tune for indoor real-estate scenes. This checkpoint is intended for DL3DV-style captures.
import torch
from models.models import OVIEModel
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = OVIEModel.from_pretrained("kyutai/ovie-ft-dl3dv").to(device)
model.eval()
image_size = model.image_size # 256
The forward pass is identical to the base model:
with torch.no_grad():
pred = model(x=img_tensor, cam_params=cam_token) # (1, 3, 256, 256) in [0, 1]
See the base model card for the complete pose-encoding snippet, and the repository for installation and the full benchmark table.
@misc{ovie2026,
title={One View Is Enough: In-the-Wild Monocular Pretraining for Novel View Generation},
author={Adrien Ramanana Rahary and Nicolas Dufour and Patrick Perez and David Picard},
year={2026},
eprint={2603.23488},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2603.23488},
}
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.
How it works