Model reference · open weights

avtr-1

Video avaturn-live Image→video 1 build Its own licence terms 746 dl/mo

avtr-1 is an open-weight video model from avaturn-live. avtr-1 (BF16) weighs 613 MB; the smallest configuration that runs it is RTX 3060 12 GB.

  • AVTR-1 is an open-source real-time avatar model by avaturn-live that generates lip-synced speech and active listening video from a portrait image and dual-stream audio.
  • It operates at 25 fps on a single GPU, with performance benchmarks ranging from 84 ms per 5-frame chunk on an L40 to 232 ms on an RTX 4060.
  • The model weights are released under the AVTR-1 Community License, which permits commercial use for entities with annual revenue below USD 10,000,000, subject to specific use restrictions.

Summary of the avaturn-live/avtr-1 model card, 2026-10-01

What it is

Released byavaturn-live
Released2026-05-25
VRAM613 MB for the weights

What it runs on

Memory and cards for avtr-1 (BF16)

613 MBweights, file size
537 MBruntime overhead
CardWeightsMemory
RTX 3060 12 GBfits11.6 GB
RTX 4060 Ti 16 GBfits15.4 GB
RTX 3090 24 GBfits23.4 GB
RTX 4090 24 GBfits23.4 GB
RTX 5090 32 GBfits31.0 GB
L40S 48 GBfits44.0 GB
A100 80 GBfits78.2 GB
H100 80 GBfits78.1 GB
RTX PRO 6000 Blackwell 96 GBfits93.8 GB
DGX Spark (GB10) 128 GB unifiedfits107 GB
H200 141 GBfits138 GB
B200 180 GBfits176 GB

How it works

How video models work

Prompt / imagestart pointTemporal diffusionframes over timeVideoMP4 clipA video model generates a sequence of coherent frames from your prompt or a starting image.

Running it yourself

Run it on a rented GPU

Rent a machine by the hour. How to run this model is on its model card.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms