Model reference · open weights
gemma-3n is an open-weight language model from Google. gemma-3n-E4B-it (BF16) weighs 15.7 GB; the smallest configuration that runs it is 2× RTX 3060 12 GB.
Gemma 3n E2B IT is an instruction-tuned multimodal model by Google that accepts text, image, video, and audio inputs to generate text. It features a raw parameter count of 6B but operates with an effective size of 2B parameters to support efficient execution on low-resource devices. The model supports a 32K token context window, was trained on data in over 140 languages, and is released under the gemma license.
Summary of the google/gemma-3n-E2B-it model card, 2026-10-01 — the estimate below is for another build of the family
What it is
| Released by | |
|---|---|
| Type | Language models |
| Task | Vision + text |
| Parameters (lead) | 5.4B |
| Runs with | transformers |
| Based on | google/gemma-3n-E4B-it |
| Released | 2025-06-12 |
| Popularity | 210k downloads / month |
| Weights | 15.7 GB (gemma-3n-E4B-it (BF16), file size) |
| Licence | Open, with conditions |
What it runs on
Weights 15.7 GB (file size) · KV cache 14 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · plus 264 MB a request for its sliding-window layers · runtime overhead from 2.1 GB on a small card · context up to 32,768 tokens.
| Card | Requests at once 8K tokens each | Requests at once 32K tokens each | Longest single request | Counted memory |
|---|---|---|---|---|
| RTX 3060 12 GB … RTX 4060 Ti 16 GB 2 smaller cards | — | — | — | |
| RTX 3090 24 GB | 14 | 7 | all 32K | 23.4 GB |
| RTX 4090 24 GB | 14 | 7 | all 32K | 23.4 GB |
| RTX 5090 32 GB | 34 | 18 | all 32K | 31.0 GB |
| L40S 48 GB | 68 | 35 | all 32K | 44.0 GB |
| A100 80 GB | 158 | 82 | all 32K | 78.2 GB |
| H100 80 GB | 93 | 38 | all 32K | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | 120 | 49 | all 32K | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | 143 | 58 | all 32K | 107 GB |
| H200 141 GB | 195 | 79 | all 32K | 138 GB |
| B200 180 GB | 259 | 64 | all 32K | 176 GB |
| 2× RTX 3060 12 GB tensor parallel | 8 | 4 | all 32K | 11.6 GB a card |
| 2× RTX 4060 Ti 16 GB tensor parallel | 28 | 14 | all 32K | 15.4 GB a card |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 18.2 GB | 18.5 GB |
| 5 | 19.7 GB | 21.5 GB |
| 8 | 20.8 GB | 23.7 GB |
| 16 | 23.9 GB | 29.5 GB |
| 32 | 30.0 GB | 41.3 GB |
| 64 | 42.2 GB | 64.8 GB |
On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.
Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (attention with sliding-window layers); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.
Builds
| Build | Parameters | Precision | Weights | Smallest configuration (1 request, 8K) |
|---|---|---|---|---|
| gemma-3n-E2B-it ↗ | 5.4B | BF16 | 10.9 GB | RTX 4060 Ti 16 GB |
| gemma-3n-E4B-it ↗ | 7.8B | BF16 | 15.7 GB | RTX 4090 24 GB |
| gemma-3n-E4B-it (above) ↗ packaged by unsloth |
8.4B | BF16 | 15.7 GB | 2× RTX 3060 12 GB |
| gemma-3n-E2B-it-GGUF ↗ packaged by unsloth 24 builds: IQ2_XXS 2.1 GB … F16 8.9 GB |
— | GGUF | 3.0 GB | RTX 3060 12 GB |
Weights from each build's files as published; ≈ = calculated from the parameter count where the files have not been read. A build's name shows what it runs on; ↗ opens it on Hugging Face.
From the model card
ollama run hf.co/unsloth/gemma-3n-E4B-it:Q4_K_XL - auto-sets correct chat template and settingsModel Page: Gemma 3n
Resources and Technical Documentation:
Terms of Use: Terms
Authors: Google DeepMind
Summary description and brief definition of inputs and outputs.
Gemma is a family of lightweight, state-of-the-art open models from Google, built from the same research and technology used to create the Gemini models. Gemma 3n models are designed for efficient execution on low-resource devices. They are capable of multimodal input, handling text, image, video, and audio input, and generating text outputs, with open weights for pre-trained and instruction-tuned variants. These models were trained with data in over 140 spoken languages.
Gemma 3n models use selective parameter activation technology to reduce resource requirements. This technique allows the models to operate at an effective size of 2B and 4B parameters, which is lower than the total number of parameters they contain. For more information on Gemma 3n's efficient parameter management technology, see the Gemma 3n page.
Below, there are some code snippets on how to get quickly started with running the model. First, install the Transformers library. Gemma 3n is supported starting from transformers 4.53.0.
$ pip install -U transformers
Then, copy the snippet from the section that is relevant for your use case.
pipeline APIYou can initialize the model and processor for inference with pipeline as
follows.
from transformers import pipeline
import torch
pipe = pipeline(
"image-text-to-text",
model="google/gemma-3n-e4b-it",
device="cuda",
torch_dtype=torch.bfloat16,
)
With instruction-tuned models, you need to use chat templates to process our inputs first. Then, you can pass it to the pipeline.
messages = [
{
"role": "system",
"content": [{"type": "text", "text": "You are a helpful assistant."}]
},
{
"role": "user",
"content": [
{"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"},
{"type": "text", "text": "What animal is on the candy?"}
]
}
]
output = pipe(text=messages, max_new_tokens=200)
print(output[0]["generated_text"][-1]["content"])
# Okay, let's take a look!
# Based on the image, the animal on the candy is a **turtle**.
# You can see the shell shape and the head and legs.
from transformers import AutoProcessor, Gemma3nForConditionalGeneration
from PIL import Image
import requests
import torch
model_id = "google/gemma-3n-e4b-it"
model = Gemma3nForConditionalGeneration.from_pretrained(model_id, device_map="auto", torch_dtype=torch.bfloat16,).eval()
processor = AutoProcessor.from_pretrained(model_id)
messages = [
{
"role": "system",
"content": [{"type": "text", "text": "You are a helpful assistant."}]
},
{
"role": "user",
"content": [
{"type": "image", "image": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/bee.jpg"},
{"type": "text", "text": "Describe this image in detail."}
]
}
]
inputs = processor.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
input_len = inputs["input_ids"].shape[-1]
with torch.inference_mode():
generation = model.generate(**inputs, max_new_tokens=100, do_sample=False)
generation = generation[0][input_len:]
decoded = processor.decode(generation, skip_special_tokens=True)
print(decoded)
# **Overall Impression:** The image is a close-up shot of a vibrant garden scene,
# focusing on a cluster of pink cosmos flowers and a busy bumblebee.
# It has a slightly soft, natural feel, likely captured in daylight.
@article{gemma_3n_2025,
title={Gemma 3n},
url={https://ai.google.dev/gemma/docs/gemma-3n},
publisher={Google DeepMind},
author={Gemma Team},
year={2025}
}
Data used for model training and how the data was processed.
These models were trained on a dataset that includes a wide variety of sources totalling approximately 11 trillion tokens. The knowledge cutoff date for the training data was June 2024. Here are the key components:
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.
How it works