Model reference · open weights

git-large

LLMs microsoft Image→text 1 build Open weights 654 dl/mo

git-large is an open-weight language model from microsoft. git-large (BF16) weighs 1.6 GB; the smallest configuration that runs it is RTX 3060 12 GB.

What it is

Released bymicrosoft
TypeLanguage models
TaskImage→text
Runs withtransformers
Released2023-01-02
Popularity654 downloads / month
Weights1.6 GB (git-large (BF16), file size)
LicenceOpen weights

What it runs on

Memory and cards for git-large (BF16)

Weights 1.6 GB (file size) · runtime overhead from 762 MB on a small card.

How much memory each request adds is not estimated yet for this architecture — only the weights are. They need the cards below at the least, plus room for the context.

CardThe weights alone
RTX 3060 12 GBfits
RTX 4060 Ti 16 GBfits
RTX 3090 24 GBfits
RTX 4090 24 GBfits
RTX 5090 32 GBfits
L40S 48 GBfits
A100 80 GBfits
H100 80 GBfits
RTX PRO 6000 Blackwell 96 GBfits
DGX Spark (GB10) 128 GB unifiedfits
H200 141 GBfits
B200 180 GBfits

From the model card

What microsoft says about git-large

GIT (short for GenerativeImage2Text) model, large-sized version. It was introduced in the paper GIT: A Generative Image-to-text Transformer for Vision and Language by Wang et al. and first released in this repository.

Disclaimer: The team releasing GIT did not write a model card for this model so this model card has been written by the Hugging Face team.

Read the full model card

Model description

GIT is a Transformer decoder conditioned on both CLIP image tokens and text tokens. The model is trained using "teacher forcing" on a lot of (image, text) pairs.

The goal for the model is simply to predict the next text token, giving the image tokens and previous text tokens.

The model has full access to (i.e. a bidirectional attention mask is used for) the image patch tokens, but only has access to the previous text tokens (i.e. a causal attention mask is used for the text tokens) when predicting the next text token.

This allows the model to be used for tasks like:

  • image and video captioning
  • visual question answering (VQA) on images and videos
  • even image classification (by simply conditioning the model on the image and asking it to generate a class for it in text).

Intended uses & limitations

You can use the raw model for image captioning. See the model hub to look for fine-tuned versions on a task that interests you.

How to use

For code examples, we refer to the documentation.

Training data

From the paper:

We collect 0.8B image-text pairs for pre-training, which include COCO (Lin et al., 2014), Conceptual Captions (CC3M) (Sharma et al., 2018), SBU (Ordonez et al., 2011), Visual Genome (VG) (Krishna et al., 2016), Conceptual Captions (CC12M) (Changpinyo et al., 2021), ALT200M (Hu et al., 2021a), and an extra 0.6B data following a similar collection procedure in Hu et al. (2021a).

=> however this is for the model referred to as "GIT" in the paper, which is not open-sourced.

This checkpoint is "GIT-large", which is a smaller variant of GIT trained on 20 million image-text pairs.

See table 11 in the paper for more details.

Preprocessing

We refer to the original repo regarding details for preprocessing during training.

During validation, one resizes the shorter edge of each image, after which center cropping is performed to a fixed-size resolution. Next, frames are normalized across the RGB channels with the ImageNet mean and standard deviation.

Evaluation results

For evaluation results, we refer readers to the paper.

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

How it works

How language models work

Your prompttext / messagesTransformerattention over tokensNext-token loopgenerate + streamResponsetext · tool callsA language model reads your tokens and predicts the next one, again and again, streaming the reply back.
© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms