Model reference · open weights

vit-gpt2-image-captioning

LLMs nlpconnect Image→text 1 build Open weights 102k dl/mo

vit-gpt2-image-captioning is an open-weight language model from nlpconnect. vit-gpt2-image-captioning (BF16) weighs 982 MB; the smallest configuration that runs it is RTX 3060 12 GB.

vit-gpt2-image-captioning is an image-to-text model developed by nlpconnect for generating captions from images. It utilizes a ViT encoder and GPT-2 decoder architecture, serving as a PyTorch version of a Flax checkpoint originally trained by ydshieh. The model is released under the Apache-2.0 license.

Summary of the nlpconnect/vit-gpt2-image-captioning model card, 2026-10-01

What it is

Released bynlpconnect
TypeLanguage models
TaskImage→text
Runs withtransformers
Released2022-03-02
Popularity102k downloads / month
Weights982 MB (vit-gpt2-image-captioning (BF16), file size)
LicenceOpen weights

What it runs on

Memory and cards for vit-gpt2-image-captioning (BF16)

Weights 982 MB (file size) · runtime overhead from 762 MB on a small card.

How much memory each request adds is not estimated yet for this architecture — only the weights are. They need the cards below at the least, plus room for the context.

CardThe weights alone
RTX 3060 12 GBfits
RTX 4060 Ti 16 GBfits
RTX 3090 24 GBfits
RTX 4090 24 GBfits
RTX 5090 32 GBfits
L40S 48 GBfits
A100 80 GBfits
H100 80 GBfits
RTX PRO 6000 Blackwell 96 GBfits
DGX Spark (GB10) 128 GB unifiedfits
H200 141 GBfits
B200 180 GBfits

From the model card

What nlpconnect says about vit-gpt2-image-captioning

Read the model card

This is an image captioning model trained by @ydshieh in flax this is pytorch version of this.

The Illustrated Image Captioning using transformers

  • https://ankur3107.github.io/blogs/the-illustrated-image-captioning-using-transformers/

Sample running code


from transformers import VisionEncoderDecoderModel, ViTImageProcessor, AutoTokenizer
import torch
from PIL import Image

model = VisionEncoderDecoderModel.from_pretrained("nlpconnect/vit-gpt2-image-captioning")
feature_extractor = ViTImageProcessor.from_pretrained("nlpconnect/vit-gpt2-image-captioning")
tokenizer = AutoTokenizer.from_pretrained("nlpconnect/vit-gpt2-image-captioning")

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model.to(device)

max_length = 16
num_beams = 4
gen_kwargs = {"max_length": max_length, "num_beams": num_beams}
def predict_step(image_paths):
  images = []
  for image_path in image_paths:
    i_image = Image.open(image_path)
    if i_image.mode != "RGB":
      i_image = i_image.convert(mode="RGB")

    images.append(i_image)

  pixel_values = feature_extractor(images=images, return_tensors="pt").pixel_values
  pixel_values = pixel_values.to(device)

  output_ids = model.generate(pixel_values, **gen_kwargs)

  preds = tokenizer.batch_decode(output_ids, skip_special_tokens=True)
  preds = [pred.strip() for pred in preds]
  return preds

predict_step(['doctor.e16ba4e4.jpg']) # ['a woman in a hospital bed with a woman in a hospital bed']

Sample running code using transformers pipeline


from transformers import pipeline

image_to_text = pipeline("image-to-text", model="nlpconnect/vit-gpt2-image-captioning")

image_to_text("https://ankur3107.github.io/assets/images/image-captioning-example.png")

# [{'generated_text': 'a soccer game with a player jumping to catch the ball '}]

Contact for any help

  • https://huggingface.co/ankur310794
  • https://twitter.com/ankur310794
  • http://github.com/ankur3107
  • https://www.linkedin.com/in/ankur310794

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms