Model reference · open weights

vit-roberta-fa-image-captioning-flickr30k

LLMs hezarai Image→text 1 build Licence not stated 2k dl/mo

vit-roberta-fa-image-captioning-flickr30k is an open-weight language model from hezarai. vit-roberta-fa-image-captioning-flickr30k (BF16) weighs 933 MB; the smallest configuration that runs it is RTX 3060 12 GB.

  • vit-roberta-fa-image-captioning-flickr30k is an image-to-text model developed by hezarai for generating Persian captions.
  • It utilizes a ViT encoder initialized from google/vit-base-patch16-224 and a RoBERTa decoder initialized from HooshvareLab/roberta-fa-zwnj-base.
  • The model is designed to process images and output text in the Persian language.

Summary of the hezarai/vit-roberta-fa-image-captioning-flickr30k model card, 2026-10-01

What it is

Released byhezarai
Released2023-09-29
VRAM933 MB for the weights

What it runs on

Memory and cards for vit-roberta-fa-image-captioning-flickr30k (BF16)

933 MBweights, file size
762 MBruntime overhead, at least

How much memory each request adds isn't estimated yet for this architecture. The weights need at least the cards below, plus room for the context.

CardWeights alone
RTX 3060 12 GBfits
RTX 4060 Ti 16 GBfits
RTX 3090 24 GBfits
RTX 4090 24 GBfits
RTX 5090 32 GBfits
L40S 48 GBfits
A100 80 GBfits
H100 80 GBfits
RTX PRO 6000 Blackwell 96 GBfits
DGX Spark (GB10) 128 GB unifiedfits
H200 141 GBfits
B200 180 GBfits

From the model card

What hezarai says about vit-roberta-fa-image-captioning-flickr30k

Read the model card

The encoder (ViT) was initialized from https://huggingface.co/google/vit-base-patch16-224 and the decoder (RoBERTa) was initialized from https://huggingface.co/HooshvareLab/roberta-fa-zwnj-base .

Usage

pip install hezar
from hezar.models import Model

model = Model.load("hezarai/vit-roberta-fa-image-captioning-flickr30k")
captions = model.predict("example_image.jpg")
print(captions)

Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.

How it works

How language models work

Your prompttext / messagesTransformerattention over tokensNext-token loopgenerate + streamResponsetext · tool callsA language model reads your tokens and predicts the next one, again and again, streaming the reply back.
© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms