Model reference · open weights

git-large-coco

LLMs microsoft Image→text 1 build Open weights 2k dl/mo

git-large-coco is an open-weight language model from microsoft. git-large-coco (FP32) weighs 788 MB; the smallest configuration that runs it is RTX 3060 12 GB.

git-large-coco is a 394M parameter image-to-text model developed by Microsoft. It is a Transformer decoder that uses CLIP image tokens and text tokens to generate captions, answer visual questions, or classify images. The model supports a context length of 1024 tokens, operates in English, and is released under the MIT licence.

Summary of the microsoft/git-large-coco model card, 2026-10-01

What it is

Released bymicrosoft
TypeLanguage models
TaskImage→text
Parameters (lead)394M
Context1,024 tokens
Runs withtransformers
Released2023-01-02
Popularity2k downloads / month
Weights788 MB (git-large-coco (FP32), file size)
LicenceOpen weights

What it runs on

Memory and cards for git-large-coco (FP32)

Weights 788 MB (file size) · KV cache 18 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 1.8 GB on a small card · context up to 1,024 tokens.

CardRequests at once
1K, its whole window tokens each
Requests at once
32K tokens each
Longest single
request
Counted
memory
RTX 3060 12 GB479—all 1K11.6 GB
RTX 4060 Ti 16 GB680—all 1K15.4 GB
RTX 3090 24 GB1000+—all 1K23.4 GB
RTX 4090 24 GB1000+—all 1K23.4 GB
RTX 5090 32 GB1000+—all 1K31.0 GB
L40S 48 GB1000+—all 1K44.0 GB
A100 80 GB1000+—all 1K78.2 GB
H100 80 GB1000+—all 1K78.1 GB
RTX PRO 6000 Blackwell 96 GB1000+—all 1K93.8 GB
DGX Spark (GB10) 128 GB unified1000+—all 1K107 GB
H200 141 GB1000+—all 1K138 GB
B200 180 GB1000+—all 1K176 GB
Memory needed at each load
Requests at once1K, its whole window tokens each32K tokens each
12.6 GB—
52.7 GB—
82.7 GB—
162.9 GB—
323.2 GB—
643.8 GB—

On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.

Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (multi-head attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). Assumes vLLM 0.10 or later.

From the model card

What microsoft says about git-large-coco

Read the model card

GIT (short for GenerativeImage2Text) model, large-sized version, fine-tuned on COCO. It was introduced in the paper GIT: A Generative Image-to-text Transformer for Vision and Language by Wang et al. and first released in this repository.

Disclaimer: The team releasing GIT did not write a model card for this model so this model card has been written by the Hugging Face team.

Model description

GIT is a Transformer decoder conditioned on both CLIP image tokens and text tokens. The model is trained using "teacher forcing" on a lot of (image, text) pairs.

The goal for the model is simply to predict the next text token, giving the image tokens and previous text tokens.

The model has full access to (i.e. a bidirectional attention mask is used for) the image patch tokens, but only has access to the previous text tokens (i.e. a causal attention mask is used for the text tokens) when predicting the next text token.

This allows the model to be used for tasks like:

  • image and video captioning
  • visual question answering (VQA) on images and videos
  • even image classification (by simply conditioning the model on the image and asking it to generate a class for it in text).

Intended uses & limitations

You can use the raw model for image captioning. See the model hub to look for fine-tuned versions on a task that interests you.

How to use

For code examples, we refer to the documentation.

Training data

From the paper:

We collect 0.8B image-text pairs for pre-training, which include COCO (Lin et al., 2014), Conceptual Captions (CC3M) (Sharma et al., 2018), SBU (Ordonez et al., 2011), Visual Genome (VG) (Krishna et al., 2016), Conceptual Captions (CC12M) (Changpinyo et al., 2021), ALT200M (Hu et al., 2021a), and an extra 0.6B data following a similar collection procedure in Hu et al. (2021a).

=> however this is for the model referred to as "GIT" in the paper, which is not open-sourced.

This checkpoint is "GIT-large", which is a smaller variant of GIT trained on 20 million image-text pairs.

Next, the model was fine-tuned on COCO.

See table 11 in the paper for more details.

Preprocessing

We refer to the original repo regarding details for preprocessing during training.

During validation, one resizes the shorter edge of each image, after which center cropping is performed to a fixed-size resolution. Next, frames are normalized across the RGB channels with the ImageNet mean and standard deviation.

Evaluation results

For evaluation results, we refer readers to the paper.

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms