Model reference · open weights
OpenSora-STDiT-HQ-16x256x256 is an open-weight embedding model from hpcai-tech. OpenSora-STDiT-v1-HQ-16x256x256 (FP32) weighs 1.5 GB; the smallest configuration that runs it is RTX 3060 12 GB.
What it is
| Released by | hpcai-tech |
|---|---|
| Type | Embedding models |
| Task | Embeddings |
| Parameters (lead) | 760M |
| Runs with | transformers |
| Released | 2024-03-20 |
| Popularity | 15 downloads / month |
| Weights | 1.5 GB (OpenSora-STDiT-v1-HQ-16x256x256 (FP32), file size) |
| Licence | Open weights |
What it runs on
Weights 1.5 GB (file size) · overhead about 1.1 GB.
| Card | Runs | Counted memory |
|---|---|---|
| RTX 3060 12 GB | fits | 11.6 GB |
| RTX 4060 Ti 16 GB | fits | 15.4 GB |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
Estimates, not measurements: the weights are the build's file size. No cache grows with use; a batch of inputs needs working memory of its own. Counted memory is 92 % of what CUDA reports for the card.
From the model card
We present Open-Sora, an initiative dedicated to efficiently produce high-quality video and make the model, tools and contents accessible to all. By embracing open-source principles, Open-Sora not only democratizes access to advanced video generation techniques, but also offers a streamlined and user-friendly platform that simplifies the complexities of video production. With Open-Sora, we aim to inspire innovation, creativity, and inclusivity in the realm of content creation.
More details can be founded at Open-Sora GitHub.
You can launch this video generation with this model in a Gradio application.
# git clone Open-Sora
git clone https://github.com/hpcaitech/Open-Sora.git
cd Open-Sora
# launch gradio
python scripts/demo.py --model-type v1-HQ-16x256x256
If you want to use this STDiT model in code,
from transformers import AutoModel
stdit = AutoModel.from_pretrained("hpcai-tech/OpenSora-STDiT-v1-HQ-16x256x256")
Do note that this model alone cannot generate video, it should work alongside a vae model and a text encoder model like how we did in the demo.
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.