Model reference · open weights

Supra2-IMG

Image SupraLabs Text→image 1 build Open weights 503 dl/mo

Supra2-IMG is an open-weight image model from SupraLabs. Supra2-IMG (BF16) weighs 417 MB; the smallest configuration that runs it is RTX 3060 12 GB.

What it is

Released bySupraLabs
TypeImage models
TaskText→image
Runs withtransformers
Released2026-09-21
Popularity503 downloads / month
Weights417 MB (Supra2-IMG (BF16), file size)
LicenceOpen weights

What it runs on

Memory and cards for Supra2-IMG (BF16)

Weights 417 MB (file size) · working memory for one 1024×1024 image about 5.0 GB · overhead about 537 MB.

CardOne 1024×1024 imageCounted
memory
RTX 3060 12 GBfits11.6 GB
RTX 4060 Ti 16 GBfits15.4 GB
RTX 3090 24 GBfits23.4 GB
RTX 4090 24 GBfits23.4 GB
RTX 5090 32 GBfits31.0 GB
L40S 48 GBfits44.0 GB
A100 80 GBfits78.2 GB
H100 80 GBfits78.1 GB
RTX PRO 6000 Blackwell 96 GBfits93.8 GB
DGX Spark (GB10) 128 GB unifiedfits107 GB
H200 141 GBfits138 GB
B200 180 GBfits176 GB

Estimates, not measurements: the weights are the build's file size; one 1024×1024 image needs about 5 GB of working memory (larger images more). diffusers can also place a pipeline's parts on separate cards (device_map) — not estimated here. Counted memory is 92 % of what CUDA reports for the card.

From the model card

What SupraLabs says about Supra2-IMG

Text-To-Image • 100M Parameters • SOTA quality

Supra2-IMG is a tiny 100M parameters text-to-image (T2I) model that has been trained from scratch on high-quality synthetic data and delivers state-of-the-art image quality for its size.


Read the full model card

Samples


Model

About the model

The model is a tiny diffusion transformer (DiT) with ~105M parameters.

  • Pipeline: text-to-image
  • Parameter Count: 104.1M
  • Encoder: frozen Flan-T5-Base
  • VAE: SD-VAE-FT-MSE
  • Image resolution: 256²
  • Latent size: 32²
  • Patch: 2
  • Context length (Flan-T5): 128 tokens

Model config

  • D_MODEL: 576
  • DEPTH: 14
  • N_HEADS: 9
  • HEAD_DIM: 64
  • MLP_RATIO: 4.0
  • D_CTX: 768
  • VAE_SCALE: 0.18215

Training

Dataset

The model was trained for 10 epochs on the full LucasFang/FLUX-Reason-6M dataset.

Data preparation

All train data images were downloaded as parquets + metadata and prepared by first chosing the prompt. This was done in the following order (each next prompt is a fallback for the previous prompt): caption_composition &arrowright; caption_entity &arrowright; caption_text &arrowright; caption_style &arrowright; caption_imaginative. That way, we ensured only using the highest quality data for pretraining the model.

Exact image count

5.6M images

Epochs count

10 epochs

Hardware

The training ran on a single Nvidia H100 SXM 80GB Runpod Pod for 9 hours (incl. data preparation) with a 2.5TB disk.


How to run the model

First, run:

# Create project directory
mkdir Supra2-IMG
cd Supra2-IMG

# Download the inference script
wget https://huggingface.co/SupraLabs/Supra2-IMG/resolve/main/inference.py

Then, you can generate images by running:

python inference.py --prompt "a sea jellyfish floating in the pitch-black ocean depths"  --seed 0  --cfg 3.0  --steps 50  --n 1  --out jellyfish.png

Recommended settings for sampling

  • --seed: 0
  • --cfg: 3.0
  • --steps: 50

Expected Output

The script will output something like:

=== Supra2-IMG inference ===
[device] ...
[ckpt] found ./model_final_ema.pt
[model] building SupraDiT ...
[model] 104.1M parameters
[model] loading weights from ./model_final_ema.pt ...
[model] weights loaded in 0.7s
[text] ctx_len=128
[text] loading tokenizer + google/flan-t5-base ...
...
[text] prompt tokens=15  n=1  seed=0  cfg=3.0  steps=50
[cfg] using stored unconditional embeddings
[sample] Euler flow, 50 steps ...
Generating: ...
[sample] denoising done in ...s
[vae] decoding latents ...
[done] saved 1 image(s) -> ...png

Possible output

Acknowledgments

  • Thank you, FLUX-Reason-6M team, for proving the LucasFang/FLUX-Reason-6M dataset. Without your work, this would never be possible!
  • Thanks to the creator of HobbyLM-Image for inspiration
  • Thanks to Runpod for providing the compute!
  • Also a big Thank You To Stability AI and Google for providing SD-VAE and Flan-T5
  • Thanks to all the other people inspiring us to make great open-source models for the community!

What's next

We will keep improving Supra2-IMG, maybe for a next-gen like Supra2.5-IMG, and we will share our progress and findings on the way to the best open-source T2I model 🤗 Please give us a like and a follow on Hugging Face if you want to support our work!

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

Running it yourself

Run it on a rented GPU

Rent a machine by the hour — how to run this model is on its model card.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms