Model reference · open weights

ovie-ft-dl3dv

NEW · this week Image kyutai Image edit 1 build Open weights 0 dl/mo

ovie-ft-dl3dv is an open-weight image model from kyutai. ovie-ft-dl3dv (FP32) weighs 286 MB; the smallest configuration that runs it is RTX 3060 12 GB.

What it is

Released bykyutai
TypeImage models
TaskImage edit
Parameters (lead)143M
Runs withpytorch
Based onkyutai/ovie
Released2026-09-29
Popularity0 downloads / month
Weights286 MB (ovie-ft-dl3dv (FP32), file size)
LicenceOpen weights

What it runs on

Memory and cards for ovie-ft-dl3dv (FP32)

Weights 286 MB (file size) · working memory for one 1024×1024 image about 5.0 GB · overhead about 537 MB.

CardOne 1024×1024 imageCounted
memory
RTX 3060 12 GBfits11.6 GB
RTX 4060 Ti 16 GBfits15.4 GB
RTX 3090 24 GBfits23.4 GB
RTX 4090 24 GBfits23.4 GB
RTX 5090 32 GBfits31.0 GB
L40S 48 GBfits44.0 GB
A100 80 GBfits78.2 GB
H100 80 GBfits78.1 GB
RTX PRO 6000 Blackwell 96 GBfits93.8 GB
DGX Spark (GB10) 128 GB unifiedfits107 GB
H200 141 GBfits138 GB
B200 180 GBfits176 GB

Estimates, not measurements: the weights are the build's file size; one 1024×1024 image needs about 5 GB of working memory (larger images more). diffusers can also place a pipeline's parts on separate cards (device_map) — not estimated here. Counted memory is 92 % of what CUDA reports for the card.

From the model card

What kyutai says about ovie-ft-dl3dv

In-the-Wild Monocular Pretraining for Novel View Generation

Part of the OVIE collection.

This is kyutai/ovie fine-tuned for 50k steps on DL3DV, the companion to the RealEstate10K fine-tune. The architecture is identical to the base model; only the weights differ.

It is the control that shows fine-tuning selects a domain: applying the same protocol on DL3DV markedly improves in-domain reconstruction relative to the base model, but costs distributional realism and generalization.

Read the full model card

Metrics

CheckpointDL3DV PSNR ↑DL3DV SSIM ↑DL3DV LPIPS ↓DL3DV FID ↓DL3DV MEt3R ↓RealEstate10K PSNR ↑RealEstate10K FID ↓
base OVIE (out-of-domain on DL3DV)14.80.3690.46413.60.07818.86.74
this checkpoint (in-domain on DL3DV)17.070.4330.36817.240.05919.347.14

Fine-tuning raises DL3DV PSNR by 2.3 dB while raising FID from 13.6 to 17.24: the model becomes a more faithful per-view reconstructor at a slight cost to global distributional realism. Notably it retains — and on RealEstate10K slightly improves — cross-domain performance even though that benchmark is now out-of-domain (PSNR 19.34 vs 18.8), so the fine-tune sharpens in-domain fidelity without eroding generalization here.

Prefer the base OVIE for in-the-wild zero-shot use and the RealEstate10K fine-tune for indoor real-estate scenes. This checkpoint is intended for DL3DV-style captures.

Usage

import torch
from models.models import OVIEModel

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

model = OVIEModel.from_pretrained("kyutai/ovie-ft-dl3dv").to(device)
model.eval()
image_size = model.image_size  # 256

The forward pass is identical to the base model:

with torch.no_grad():
    pred = model(x=img_tensor, cam_params=cam_token)  # (1, 3, 256, 256) in [0, 1]

See the base model card for the complete pose-encoding snippet, and the repository for installation and the full benchmark table.

Citation

@misc{ovie2026,
      title={One View Is Enough: In-the-Wild Monocular Pretraining for Novel View Generation},
      author={Adrien Ramanana Rahary and Nicolas Dufour and Patrick Perez and David Picard},
      year={2026},
      eprint={2603.23488},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2603.23488},
}

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

How it works

How image models work

Text promptwhat to makeText encoderunderstands itDiffusion stepsdenoise to pixelsImagePNG / JPEGA diffusion model starts from noise and denoises it, guided by your prompt, into a finished image.
© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms