Model reference · open weights

t5gemma-2

LLMs google Vision + text 3 builds Open, with conditions 50k dl/mo

t5gemma-2 is an open-weight language model from Google. t5gemma-2-4b-4b (BF16) weighs 15.0 GB; the smallest configuration that runs it is RTX 4090 24 GB.

  • T5Gemma 2 is a lightweight encoder-decoder model from Google designed for image-text-to-text tasks such as question answering and summarization.
  • The 1B-1B variant contains 2.1B parameters and supports a 128K context window across over 140 languages.
  • It is released under the gemma license and accepts text and image inputs to generate text outputs.

Summary of the google/t5gemma-2-1b-1b model card, 2026-10-01. The estimate below is for another build of the family.

What it is

Released byGoogle
Released2025-10-25
Parameters2.1B
VRAM15.0 GB for the weights

What it runs on

Memory and cards for t5gemma-2-4b-4b (BF16)

15.0 GBweights, file size
762 MBruntime overhead, at least

How much memory each request adds isn't estimated yet for this architecture. The weights need at least the cards below, plus room for the context.

CardWeights alone
RTX 3060 12 GB … RTX 4060 Ti 16 GB
2 smaller cards
does not fit
RTX 3090 24 GBfits
RTX 4090 24 GBfits
RTX 5090 32 GBfits
L40S 48 GBfits
A100 80 GBfits
H100 80 GBfits
RTX PRO 6000 Blackwell 96 GBfits
DGX Spark (GB10) 128 GB unifiedfits
H200 141 GBfits
B200 180 GBfits

Builds

Sizes, precisions & builds

BuildParamsPrecisionWeightsSmallest setup
t5gemma-2-1b-1b ↗ 2.1BBF16 4.2 GBRTX 3060 12 GB
t5gemma-2-270m-270m ↗ 786MBF16 1.6 GBRTX 3060 12 GB
t5gemma-2-4b-4b (above) ↗ 8.9BBF16 15.0 GBRTX 4090 24 GB

How it works

How language models work

Your prompttext / messagesTransformerattention over tokensNext-token loopgenerate + streamResponsetext · tool callsA language model reads your tokens and predicts the next one, again and again, streaming the reply back.
© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms