Model reference · open weights

NAVA

Available as managed deployment Video baidu Text→video 1 variants 80 dl/mo

NAVA is an open-weight video model from baidu. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

Makerbaidu
TypeVideo models
TaskText→video
Runs withcustom
Based onWan-AI/Wan2.2-TI2V-5B
Released2026-05-29
Popularity80 downloads / month
LicenceOpen weights

About

What NAVA is

⭐ If you find this model useful, please consider giving our GitHub repo a star! ⭐

📖 中文版 README


TL;DR

NAVA is a 6.3 B-parameter joint audio-video generator that synthesizes synchronized video and audio from a single prompt — including multi-speaker speech with reference-timbre control and image-conditioned continuations.

Instead of post-hoc-aligned dual towers or fully unified tri-modal stacks, NAVA uses an Align-then-Fuse MMDiT: a dedicated alignment space first establishes audio-video correspondence, then context (text, speaker embeddings) is fused via cross-attention. On Verse-Bench it sets new SOTA on Sync-C / Sync-D / video quality / audio WER while using 2× to 5× fewer parameters than open-source baselines.

Highlights

  • 720p 1-min Fast Generation — 720p synchronized audio-video in ~1 minute via 8-GPU Ulysses sequence parallel.
  • Dual-Channel Audio — stereo audio (scene + speech) jointly denoised with video, no post-hoc vocoder alignment.
  • Precise Multi-Timbre Control — reference WAVs bound to ... speech spans for per-speaker voice identity.
  • Language-Described Camera Control — shot composition, motion, and pacing directly from the prompt.
  • Multi-Resolution — landscape / portrait / square aspect ratios from the same checkpoint.

Model Details

Quick Facts

ArchitectureAlign-then-Fuse MMDiT (Wan2.2 backbone)
Parameters6.3 B (backbone, joint AV)
ModalityJoint audio + video, text-conditioned
Resolution1280×704 (recommended) · 960×960 also supported
Frames / FPS37 frames @ 24 fps ≈ 6 s · 55–61 frames ≈ 9–10 s
Audio25 latent tokens / sec, ≤ 10 s
SamplingFlow matching · UniPC scheduler · 50 default steps
Precisionbf16
ParallelismSingle-GPU or Ulysses sequence parallel (up to 8 GPUs)
Base modelWan-AI/Wan2.2-TI2V-5B

Architecture

NAVA instantiates Native Audio-Visual Alignment as an Align-then-Fuse MMDiT stack:

  • Hierarchical Alignment Layers — 10 double-stream blocks. Video and audio keep separate QKV projections and FFNs but share a joint self-attention over concatenated [video_tokens; audio_tokens], plus dedicated cross-attention to text. This builds an alignment space where AV correspondence is learned without semantic context interference.
  • Unified Fusion Layers — 20 single-stream blocks. Video and audio share QKV/FFN; a unified joint attention treats all tokens as one stream, with a single text cross-attention path. This is where context-conditioned denoising happens.
  • Backbone hyperparameters. dim=3072, ffn_dim=14336, 24 attention heads, 30 layers (10 double + 20 single), text_len=512, patch size (1, 2, 2). RMSNorm on QK; cross-attention norm; ε = 1e-6.
  • Positional encoding. 3D RoPE for video (temporal + height + width), 1D RoPE for audio, applied jointly inside the joint-attention path.
  • Timbre-in-Context Conditioning. Reference-WAV speaker embeddings (ReDimNet, 192-d) are injected through the context pathway and bound to ... speech spans, enabling per-speaker timbre control in multi-speaker scenes.
  • 3D cross-modal CFG. Independent classifier-free guidance scales for video, audio, and the cross-modal alignment direction (video_align_guidance_scale, audio_align_guidance_scale) keep AV synchronization tight at inference.

What's Different from Existing Open-Source AV Models

Design axisTypical baselinesNAVA
Stream layoutDual-tower (post-hoc align) or fully unified tri-modalAlign-then-Fuse — alignment space first, context fused after
Speech controlCaption-only, no per-speaker timbreTimbre-in-Context via reference WAVs
Param budget10 B – 32 B6.3 B

Components Shipped Alongside the Backbone

ComponentDescriptionSize
WanAVModel (backbone)MMDiT, joint AV attention6.3 B
Wan2.2 Video VAECausal 3D ConvNet · 16×16×4 spatial-temporal compression · 48 latent channels2.7 GB
LTX Audio VAE + Vocoder128 latent channels · 25 tokens/sec · built-in waveform decoder348 MB
umt5-xxl Text EncoderT5 · 4096-d embeddings11 GB
ReDimNetSpeaker embedding · 192-d~50 MB

Evaluation

Table 1 — VerseBench (general AV capability)

NAVA achieves the best AV synchronization (Sync-C / Sync-D), video quality, and audio WER, with the smallest parameter budget.

ModelParamsResolutionSync-C ↑Sync-D ↓IB ↑Video Quality ↑WER ↓PQ ↑FD ↓
Ovi 1.110 B720p7.48397.97910.1990.6360.1025.84320.9418
MOVAA18B (32 B)720p7.28887.8080.2690.6030.1267.23310.9222
Davinci15 B540p7.14877.81580.2690.6000.1515.95590.9307
LTX 2.319 B512p7.24767.69020.3370.5760.1066.94590.8287
NAVA (ours)6.3 B720p7.79147.56550.3130.6590.0996.86090.8328

Table 2 — Seed-TTS-eval (speech quality)

Among joint AV models, NAVA delivers speech quality close to dedicated audio-only systems. Audio-only rows are listed for reference; they are not directly comparable.

CategoryModelWER ↓Speaker Similarity ↑
Audio-Only (reference)CosyVoice4.2960.9
Audio-Only (reference)Qwen2.5-Omni2.7263.2
Audio-Video JointDreamID-Omni33.4434.1
Audio-Video JointNAVA (ours)5.8162.4

How to Use

TL;DR command. After §1 setup is complete:

bash scripts/inference.sh           # General T2AV
bash scripts/inference_timbre.sh    # I2AV + timbre control

Outputs land under eval_results/.

1 · Setu

From the published model card. Full card on the HuggingFace links in the sidebar.

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys nava for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (nava below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/videos/generations \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"nava","prompt":"a drone shot over a forest"}'

Create an account — your API key is available in the console. 5M tokens/month currently included with every new account at launch.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms