Model reference · open weights
gemma-3n is an open-weight language model from Google. gemma-3n-E2B-it (BF16) weighs 10.9 GB; the smallest configuration that runs it is RTX 4060 Ti 16 GB.
Gemma 3n E2B IT is an instruction-tuned multimodal model by Google that accepts text, image, video, and audio inputs to generate text. It features a raw parameter count of 6B but operates with an effective size of 2B parameters to support efficient execution on low-resource devices. The model supports a 32K token context window, was trained on data in over 140 languages, and is released under the gemma license.
Summary of the google/gemma-3n-E2B-it model card, 2026-10-01
What it is
| Released by | |
|---|---|
| Type | Language models |
| Task | Vision + text |
| Parameters (lead) | 5.4B |
| Runs with | transformers |
| Based on | google/gemma-3n-E4B-it |
| Released | 2025-06-12 |
| Popularity | 210k downloads / month |
| Weights | 10.9 GB (gemma-3n-E2B-it (BF16), file size) |
| Licence | Open, with conditions |
What it runs on
Weights 10.9 GB (file size) · runtime overhead from 762 MB on a small card.
How much memory each request adds is not estimated yet for this architecture — only the weights are. They need the cards below at the least, plus room for the context.
| Card | The weights alone |
|---|---|
| RTX 3060 12 GB | does not fit |
| RTX 4060 Ti 16 GB | fits |
| RTX 3090 24 GB | fits |
| RTX 4090 24 GB | fits |
| RTX 5090 32 GB | fits |
| L40S 48 GB | fits |
| A100 80 GB | fits |
| H100 80 GB | fits |
| RTX PRO 6000 Blackwell 96 GB | fits |
| DGX Spark (GB10) 128 GB unified | fits |
| H200 141 GB | fits |
| B200 180 GB | fits |
Builds
| Build | Parameters | Precision | Weights | Smallest configuration (1 request, 8K) |
|---|---|---|---|---|
| gemma-3n-E2B-it (above) ↗ | 5.4B | BF16 | 10.9 GB | RTX 4060 Ti 16 GB |
| gemma-3n-E4B-it ↗ | 7.8B | BF16 | 15.7 GB | RTX 4090 24 GB |
| gemma-3n-E4B-it ↗ packaged by unsloth |
8.4B | BF16 | 15.7 GB | 2× RTX 3060 12 GB |
| gemma-3n-E2B-it-GGUF ↗ packaged by unsloth 24 builds: IQ2_XXS 2.1 GB … F16 8.9 GB |
— | GGUF | 3.0 GB | RTX 3060 12 GB |
Weights from each build's files as published; ≈ = calculated from the parameter count where the files have not been read. A build's name shows what it runs on; ↗ opens it on Hugging Face.
From the model card
ollama run hf.co/unsloth/gemma-3n-E4B-it:Q4_K_XL - auto-sets correct chat template and settingsModel Page: Gemma 3n
Resources and Technical Documentation:
Terms of Use: Terms
Authors: Google DeepMind
Summary description and brief definition of inputs and outputs.
Gemma is a family of lightweight, state-of-the-art open models from Google, built from the same research and technology used to create the Gemini models. Gemma 3n models are designed for efficient execution on low-resource devices. They are capable of multimodal input, handling text, image, video, and audio input, and generating text outputs, with open weights for pre-trained and instruction-tuned variants. These models were trained with data in over 140 spoken languages.
Gemma 3n models use selective parameter activation technology to reduce resource requirements. This technique allows the models to operate at an effective size of 2B and 4B parameters, which is lower than the total number of parameters they contain. For more information on Gemma 3n's efficient parameter management technology, see the Gemma 3n page.
Below, there are some code snippets on how to get quickly started with running the model. First, install the Transformers library. Gemma 3n is supported starting from transformers 4.53.0.
$ pip install -U transformers
Then, copy the snippet from the section that is relevant for your use case.
pipeline APIYou can initialize the model and processor for inference with pipeline as
follows.
from transformers import pipeline
import torch
pipe = pipeline(
"image-text-to-text",
model="google/gemma-3n-e4b-it",
device="cuda",
torch_dtype=torch.bfloat16,
)
With instruction-tuned models, you need to use chat templates to process our inputs first. Then, you can pass it to the pipeline.
messages = [
{
"role": "system",
"content": [{"type": "text", "text": "You are a helpful assistant."}]
},
{
"role": "user",
"content": [
{"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"},
{"type": "text", "text": "What animal is on the candy?"}
]
}
]
output = pipe(text=messages, max_new_tokens=200)
print(output[0]["generated_text"][-1]["content"])
# Okay, let's take a look!
# Based on the image, the animal on the candy is a **turtle**.
# You can see the shell shape and the head and legs.
from transformers import AutoProcessor, Gemma3nForConditionalGeneration
from PIL import Image
import requests
import torch
model_id = "google/gemma-3n-e4b-it"
model = Gemma3nForConditionalGeneration.from_pretrained(model_id, device_map="auto", torch_dtype=torch.bfloat16,).eval()
processor = AutoProcessor.from_pretrained(model_id)
messages = [
{
"role": "system",
"content": [{"type": "text", "text": "You are a helpful assistant."}]
},
{
"role": "user",
"content": [
{"type": "image", "image": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/bee.jpg"},
{"type": "text", "text": "Describe this image in detail."}
]
}
]
inputs = processor.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
input_len = inputs["input_ids"].shape[-1]
with torch.inference_mode():
generation = model.generate(**inputs, max_new_tokens=100, do_sample=False)
generation = generation[0][input_len:]
decoded = processor.decode(generation, skip_special_tokens=True)
print(decoded)
# **Overall Impression:** The image is a close-up shot of a vibrant garden scene,
# focusing on a cluster of pink cosmos flowers and a busy bumblebee.
# It has a slightly soft, natural feel, likely captured in daylight.
@article{gemma_3n_2025,
title={Gemma 3n},
url={https://ai.google.dev/gemma/docs/gemma-3n},
publisher={Google DeepMind},
author={Gemma Team},
year={2025}
}
Data used for model training and how the data was processed.
These models were trained on a dataset that includes a wide variety of sources totalling approximately 11 trillion tokens. The knowledge cutoff date for the training data was June 2024. Here are the key components:
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.
How it works