Model reference · open weights

ovie-512

NEW · this week Image kyutai Image edit 1 build Open weights 0 dl/mo

ovie-512 is an open-weight image model from kyutai. ovie-512 (FP32) weighs 290 MB; the smallest configuration that runs it is RTX 3060 12 GB.

What it is

Released bykyutai
TypeImage models
TaskImage edit
Parameters (lead)145M
Runs withpytorch
Released2026-09-29
Popularity0 downloads / month
Weights290 MB (ovie-512 (FP32), file size)
LicenceOpen weights

What it runs on

Memory and cards for ovie-512 (FP32)

Weights 290 MB (file size) · working memory for one 1024×1024 image about 5.0 GB · overhead about 537 MB.

CardOne 1024×1024 imageCounted
memory
RTX 3060 12 GBfits11.6 GB
RTX 4060 Ti 16 GBfits15.4 GB
RTX 3090 24 GBfits23.4 GB
RTX 4090 24 GBfits23.4 GB
RTX 5090 32 GBfits31.0 GB
L40S 48 GBfits44.0 GB
A100 80 GBfits78.2 GB
H100 80 GBfits78.1 GB
RTX PRO 6000 Blackwell 96 GBfits93.8 GB
DGX Spark (GB10) 128 GB unifiedfits107 GB
H200 141 GBfits138 GB
B200 180 GBfits176 GB

Estimates, not measurements: the weights are the build's file size; one 1024×1024 image needs about 5 GB of working memory (larger images more). diffusers can also place a pipeline's parts on separate cards (device_map) — not estimated here. Counted memory is 92 % of what CUDA reports for the card.

From the model card

What kyutai says about ovie-512

In-the-Wild Monocular Pretraining for Novel View Generation

Part of the OVIE collection.

OVIE is resolution-agnostic by construction: the encoder and decoder are fully convolutional and the ViT bottleneck regenerates its positional encodings for the new grid. This checkpoint is the base model retrained at 512×512 under the identical recipe (same in-the-wild mix with MoGe-2 pseudo-pairs), for 250K steps at global batch size 256, changing only the resolution.

At a like-for-like 256×256 evaluation it matches or improves on the base model on PSNR, SSIM, LPIPS and MEt3R on both benchmarks, at a cost of roughly one FID point, which the paper attributes to the training mix — part of it sits below 512 and is upscaled, carrying little true high-frequency detail. Inference stays a single feed-forward pass with no per-scene construction: 41.6 ms per view on an H100 against 8.6 ms at 256×256.

Read the full model card

Metrics

eval@256 brings every output to 256×256 before scoring (like-for-like with the base model); eval@512 scores at native resolution and is given for reference only, since SSIM and LPIPS are resolution-dependent.

ModelBenchmarkPSNR ↑SSIM ↑LPIPS ↓FID ↓MEt3R ↓
OVIE-512, eval@256RealEstate10K19.10.6110.2837.620.034
OVIE-512, eval@256DL3DV15.20.3890.46414.80.077
OVIE-512, eval@512RealEstate10K18.90.6580.344——
OVIE-512, eval@512DL3DV15.10.4750.509——

For reference, the base OVIE (256×256) scores 18.8 / 0.602 / 0.279 / 6.74 / 0.035 on RealEstate10K and 14.8 / 0.369 / 0.464 / 13.6 / 0.078 on DL3DV.

Usage

import torch
from models.models import OVIEModel
from utils.pose_enc import extri_intri_to_pose_encoding
from torchvision.transforms import ToTensor
from PIL import Image

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

model = OVIEModel.from_pretrained("kyutai/ovie-512").to(device)
model.eval()
image_size = model.image_size  # 512, read from the saved config

img_pil = Image.open("image.jpg").convert("RGB").resize((image_size, image_size))
img_tensor = ToTensor()(img_pil).unsqueeze(0).to(device)

extrinsics = torch.tensor([[[1.0, 0.0, 0.0, -1.25],
                            [0.0, 1.0, 0.0,  0.5],
                            [0.0, 0.0, 1.0, -2.0]]], device=device)
dummy_intrinsics = torch.zeros(1, 1, 3, 3, device=device)

camera = extri_intri_to_pose_encoding(
    extrinsics=extrinsics.unsqueeze(0),
    intrinsics=dummy_intrinsics,
    image_size_hw=(image_size, image_size),
)
cam_token = camera[..., :7].squeeze(0)

with torch.no_grad():
    pred = model(x=img_tensor, cam_params=cam_token)  # (1, 3, 512, 512) in [0, 1]

Note the output resolution is fixed at training time: unlike methods that render from an explicit 3D representation, OVIE-512 cannot produce arbitrary resolutions at inference.

To evaluate with the repository's script, pass --image_size 512:

uv run python evaluate.py \
    --dataset_path /PATH/TO/RE10K/TEST \
    --config_path configs/config_ovie.yaml \
    --from_pretrained kyutai/ovie-512 \
    --image_size 512 --stride 3 --num_target_frames 14

See the repository for installation, data preprocessing, and the full benchmark table.

Citation

@misc{ovie2026,
      title={One View Is Enough: In-the-Wild Monocular Pretraining for Novel View Generation},
      author={Adrien Ramanana Rahary and Nicolas Dufour and Patrick Perez and David Picard},
      year={2026},
      eprint={2603.23488},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2603.23488},
}

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

How it works

How image models work

Text promptwhat to makeText encoderunderstands itDiffusion stepsdenoise to pixelsImagePNG / JPEGA diffusion model starts from noise and denoises it, guided by your prompt, into a finished image.
© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms