Model reference · open weights
UniVR-Planning is an open-weight language model from ByteDance. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.
Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.
What it is
| Maker | ByteDance |
|---|---|
| Type | Language models |
| Task | Vision + text |
| Runs with | transformers |
| Based on | BAAI/Emu3.5 |
| Released | 2026-07-13 |
| Popularity | 3 downloads / month |
| Licence | Open weights |
About
UniVR is the first framework that simultaneously learns complex reasoning, fine-grained physical dynamics, and long-term planning from pure visual demonstrations — without relying on dense image-text pairs or task-specific heuristics.
Built on Emu3.5 (34B), UniVR uses a unified next-token prediction objective to directly generate visual reasoning traces given an image and instruction. Training employs a two-stage pipeline: supervised cold initialization on the VR-X dataset, followed by VR-GRPO reinforcement learning with complementary global and step-focal rewards.
| Feature | Detail |
|---|---|
| Architecture | Emu3.5 34B (VQ-VAE unified generative model) |
| Training | SFT (310k samples) → VR-GRPO RL (3k samples) |
| Visual Thinking | Native visual-space reasoning, no intermediate text chain |
| Benchmark | VR-X: 16 sources, 6 task categories, 1.8k evaluation samples |
| Model | Description | Path |
|---|---|---|
| UniVR-34B-Planning | Optimized for long-horizon planning tasks (robotic manipulation, tool use, multi-step control) | Planning/ |
| UniVR-34B-General | Full UniVR recipe with interleaved image-text data; suitable for general visual reasoning | General/ |
Both checkpoints live as subdirectories of this repo: ByteDance/UniVR-34B-Planning.
UniVR achieves up to 25% improvement over the Emu3.5 baseline and approaches Gemini 3 Pro + Nano Banana 2 with only 34B parameters.
| Method | Visual Thinking | Guidance | Robot | Editing | Spatial | Puzzle | Search | Overall↑ |
|---|---|---|---|---|---|---|---|---|
| Gemini-3-pro + Nano Banana 2 | ✗ | 66.2 | 67.1 | 63.7 | 55.1 | 65.5 | 79.0 | 66.1 |
| GPT-5 + GPT-image-1.5 | ✗ | 68.2 | 64.1 | 58.0 | 49.3 | 64.0 | 77.4 | 63.5 |
| Emu3.5 34B | ✗ | 38.6 | 42.8 | 32.7 | 35.3 | 43.4 | 46.2 | 39.8 |
| UniVR 34B | ✓ | 59.5 | 68.0 | 48.5 | 46.5 | 62.2 | 64.3 | 58.2 |
| Δ v.s. Emu3.5 | ↑20.9 | ↑25.2 | ↑15.8 | ↑11.2 | ↑18.8 | ↑18.1 | ↑18.4 |
Enhanced visual reasoning also boosts standard multimodal benchmarks — no degradation of the base model's capabilities.
| Method | MMMU | MME(P) | MME(C) | MMBench | MathVista | MM-Vet |
|---|---|---|---|---|---|---|
| Emu 3.5 | 0.292 | 781.1 | 324.6 | 0.183 | 41.7 | 28.0 |
| UniVR | 0.337 | 799.3 | 338.5 | 0.198 | 44.0 | 35.6 |
| Δ v.s. Emu3.5 | ↑0.045 | ↑18.2 | ↑13.9 | ↑0.015 | ↑2.3 | ↑7.6 |
git clone https://github.com/bytedance/UniVR.git
cd UniVR
bash install.sh
cd UniVR_SFT
# Download checkpoint (choose Planning or General subdir)
huggingface-cli download ByteDance/UniVR-34B-Planning --include "Planning/*" --local-dir weights/UniVR-34B-Planning
# Download VisionTokenizer
huggingface-cli download BAAI/Emu3.5-VisionTokenizer --local-dir weights/Emu3.5-VisionTokenizer
# Run inference
bash scripts/inference.sh
Configure configs/config.py to set model paths and prompts:
{
"prompt": "Tie the red rope around the white gift box. Finish this task in 3 steps.",
"reference_image": "path/to/first_frame.jpg",
}
SFT (Cold Initialization):
cd UniVR_SFT
# LoRA (2 nodes × 8 GPUs)
bash scripts/train_sft_lora.sh
# Full parameter (4 nodes × 8 GPUs)
bash scripts/train_sft_full.sh
RL (VR-GRPO):
cd UniVR_RL
bash examples/emu3_grpo_lora.sh
UniVR proposes VR-GRPO (Visual Reasoning GRPO), a reinforcement learning paradigm that combines:
R_reason = R_g − λ|R_g − R_s|, enforcing both terminal correctness and procedural integrity.This design prevents reward hacking in long-horizon tasks where global-only rewards overlook intermediate physical violations and logical gaps.
UniVR is trained on VR-X, a large-scale benchmark curated from 1.5M raw samples across 16 diverse sources:
| Category | Sources | Examples |
|---|---|---|
| Visual Guidance | EgoDex, Action100M, Epic-Kitchen, VideoCraftBench | Cooking, handcrafting, daily activities |
| Robot Manipulation | AgiBot, Droid, Bridge, ZebraCoT-Robot | Robotic grasping, tool use, multi-step control |
| Editing | ZebraCoT-Multiobject | Object manipulation, scene editing |
| Spatial Perception | ThinkMorph-Navigation, ZebraCoT-Embodiment | Navigation, spatial reasoning |
| Visual Search | VisualCoT, ThinkMorph-Search | Object localization, attention |
| Puzzle & Game | VRBench, Zebra-Jigsaw, ThinkMorph-VisPuzzle | Mazes, jigsaw, visual puzzles |
Download: ByteDance/VR-X-SFT-RL
@article{ren2026univr,
title={UniVR: Thinking in Visual Space for Unified Visual Reasoning},
author={Zhongwei Ren and Yunchao Wei and Zhao Yao and Guixun Luo and Yao Zhao and Weibo Gong and Xiao Liu and Anran Wang and Xiangtai Li and Xiaojie Jin},
url={https://arxiv.org/abs/2607.12800},
year={2026},
}
This project is released under the CC BY 4.0 License.
UniVR is built upon Emu3.5 and verl. We thank the authors for their excellent open-source contributions.
From the published model card. Full card on the HuggingFace links in the sidebar.
How it works
Using it via the API
Once AxForge deploys univr-planning for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (univr-planning below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/chat/completions \
-H "Authorization: Bearer $AXFORGE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"univr-planning","messages":[{"role":"user","content":"Hello"}]}'
Create an account — your API key is available in the console. 5M tokens/month currently included with every new account at launch.