Model reference · open weights

translategemma

LLMs google Vision + text 3 builds Open, with conditions 31k dl/mo

translategemma is an open-weight language model from Google. translategemma-27b-it (BF16) weighs 54.9 GB; the smallest configuration that runs it is H100 80 GB.

  • TranslateGemma is a family of lightweight open translation models developed by Google, based on the Gemma 3 architecture.
  • The 5.0B parameter variant is designed for image-text-to-text tasks, handling translation across 55 languages with a total input context of 2K tokens.
  • It is released under the gemma licence and supports both direct text translation and text extraction from images.

Summary of the google/translategemma-4b-it model card, 2026-10-01. The estimate below is for another build of the family.

What it is

Released byGoogle
Released2026-01-12
Parameters5.0B
VRAM54.9 GB for the weights

What it runs on

Memory and cards for translategemma-27b-it (BF16)

54.9 GBweights, file size
762 MBruntime overhead, at least

How much memory each request adds isn't estimated yet for this architecture. The weights need at least the cards below, plus room for the context.

CardWeights alone
RTX 3060 12 GB … L40S 48 GB
6 smaller cards
does not fit
A100 80 GBfits
H100 80 GBfits
RTX PRO 6000 Blackwell 96 GBfits
DGX Spark (GB10) 128 GB unifiedfits
H200 141 GBfits
B200 180 GBfits

Builds

Sizes, precisions & builds

BuildParamsPrecisionWeightsSmallest setup
translategemma-4b-it ↗ 5.0BBF16 8.6 GBRTX 3060 12 GB
translategemma-27b-it (above) ↗ 28.8BBF16 54.9 GBH100 80 GB
translategemma-12b-it ↗ 13.2BBF16 24.4 GBRTX 5090 32 GB

How it works

How language models work

Your prompttext / messagesTransformerattention over tokensNext-token loopgenerate + streamResponsetext · tool callsA language model reads your tokens and predicts the next one, again and again, streaming the reply back.
© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms