Model reference · open weights
gemma is an open-weight language model from Google. gemma-7b (BF16) weighs 34.2 GB; the smallest configuration that runs it is L40S 48 GB.
Summary of the google/gemma-2b model card, 2026-10-01. The estimate below is for another build of the family.
What it is
| Released by | |
|---|---|
| Released | 2024-02-08 |
| Parameters | 2.5B |
| VRAM | 34.2 GB for the weights |
What it runs on
How much memory each request adds isn't estimated yet for this architecture. The weights need at least the cards below, plus room for the context.
| Card | Weights alone |
|---|---|
| RTX 3060 12 GB … RTX 5090 32 GB 5 smaller cards | does not fit |
| L40S 48 GB | fits |
| A100 80 GB | fits |
| H100 80 GB | fits |
| RTX PRO 6000 Blackwell 96 GB | fits |
| DGX Spark (GB10) 128 GB unified | fits |
| H200 141 GB | fits |
| B200 180 GB | fits |
| 2× RTX 4090 24 GB split by layers (llama.cpp) | fits |
| 2× RTX 3090 24 GB split by layers (llama.cpp) | fits |
| 2× RTX 5090 32 GB split by layers (llama.cpp) | fits |
Builds
How it works