Model reference · open weights
vit-roberta-fa-image-captioning-flickr30k is an open-weight language model from hezarai. vit-roberta-fa-image-captioning-flickr30k (BF16) weighs 933 MB; the smallest configuration that runs it is RTX 3060 12 GB.
Summary of the hezarai/vit-roberta-fa-image-captioning-flickr30k model card, 2026-10-01
What it is
| Released by | hezarai |
|---|---|
| Released | 2023-09-29 |
| VRAM | 933 MB for the weights |
What it runs on
How much memory each request adds isn't estimated yet for this architecture. The weights need at least the cards below, plus room for the context.
| Card | Weights alone |
|---|---|
| RTX 3060 12 GB | fits |
| RTX 4060 Ti 16 GB | fits |
| RTX 3090 24 GB | fits |
| RTX 4090 24 GB | fits |
| RTX 5090 32 GB | fits |
| L40S 48 GB | fits |
| A100 80 GB | fits |
| H100 80 GB | fits |
| RTX PRO 6000 Blackwell 96 GB | fits |
| DGX Spark (GB10) 128 GB unified | fits |
| H200 141 GB | fits |
| B200 180 GB | fits |
From the model card
The encoder (ViT) was initialized from https://huggingface.co/google/vit-base-patch16-224 and the decoder (RoBERTa) was initialized from https://huggingface.co/HooshvareLab/roberta-fa-zwnj-base .
pip install hezar
from hezar.models import Model
model = Model.load("hezarai/vit-roberta-fa-image-captioning-flickr30k")
captions = model.predict("example_image.jpg")
print(captions)
Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.
How it works