Model reference · open weights

gemma-3n

LLMs google Vision + text 4 builds Open, with conditions 210k dl/mo

gemma-3n is an open-weight language model from Google. gemma-3n-E4B-it (BF16) weighs 15.7 GB; the smallest configuration that runs it is 2× RTX 3060 12 GB.

Gemma 3n E2B IT is an instruction-tuned multimodal model by Google that accepts text, image, video, and audio inputs to generate text. It features a raw parameter count of 6B but operates with an effective size of 2B parameters to support efficient execution on low-resource devices. The model supports a 32K token context window, was trained on data in over 140 languages, and is released under the gemma license.

Summary of the google/gemma-3n-E2B-it model card, 2026-10-01 — the estimate below is for another build of the family

What it is

Released byGoogle
TypeLanguage models
TaskVision + text
Parameters (lead)5.4B
Runs withtransformers
Based ongoogle/gemma-3n-E4B-it
Released2025-06-12
Popularity210k downloads / month
Weights15.7 GB (gemma-3n-E4B-it (BF16), file size)
LicenceOpen, with conditions

What it runs on

Memory and cards for gemma-3n-E4B-it (BF16)

Weights 15.7 GB (file size) · KV cache 14 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · plus 264 MB a request for its sliding-window layers · runtime overhead from 2.1 GB on a small card · context up to 32,768 tokens.

CardRequests at once
8K tokens each
Requests at once
32K tokens each
Longest single
request
Counted
memory
RTX 3060 12 GB … RTX 4060 Ti 16 GB
2 smaller cards
———
RTX 3090 24 GB147all 32K23.4 GB
RTX 4090 24 GB147all 32K23.4 GB
RTX 5090 32 GB3418all 32K31.0 GB
L40S 48 GB6835all 32K44.0 GB
A100 80 GB15882all 32K78.2 GB
H100 80 GB9338all 32K78.1 GB
RTX PRO 6000 Blackwell 96 GB12049all 32K93.8 GB
DGX Spark (GB10) 128 GB unified14358all 32K107 GB
H200 141 GB19579all 32K138 GB
B200 180 GB25964all 32K176 GB
2× RTX 3060 12 GB
tensor parallel
84all 32K11.6 GB a card
2× RTX 4060 Ti 16 GB
tensor parallel
2814all 32K15.4 GB a card
Memory needed at each load
Requests at once8K tokens each32K tokens each
118.2 GB18.5 GB
519.7 GB21.5 GB
820.8 GB23.7 GB
1623.9 GB29.5 GB
3230.0 GB41.3 GB
6442.2 GB64.8 GB

On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.

Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (attention with sliding-window layers); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.

Builds

Sizes, precisions & builds

BuildParametersPrecisionWeightsSmallest configuration (1 request, 8K)
gemma-3n-E2B-it ↗ 5.4BBF16 10.9 GBRTX 4060 Ti 16 GB
gemma-3n-E4B-it ↗ 7.8BBF16 15.7 GBRTX 4090 24 GB
gemma-3n-E4B-it (above) ↗
packaged by unsloth
8.4BBF16 15.7 GB2× RTX 3060 12 GB
gemma-3n-E2B-it-GGUF ↗
packaged by unsloth
24 builds: IQ2_XXS 2.1 GB … F16 8.9 GB
—GGUF 3.0 GBRTX 3060 12 GB

Weights from each build's files as published; ≈ = calculated from the parameter count where the files have not been read. A build's name shows what it runs on; ↗ opens it on Hugging Face.

From the model card

What Google says about gemma-3n

Read the model card
  • Currently only text is supported.
  • Ollama: ollama run hf.co/unsloth/gemma-3n-E4B-it:Q4_K_XL - auto-sets correct chat template and settings
  • Set temperature = 1.0, top_k = 64, top_p = 0.95, min_p = 0.0
  • Gemma 3n max tokens (context length): 32K. Gemma 3n chat template:

🦥 Fine-tune Gemma 3n with Unsloth

Gemma-3n-E4B model card

Model Page: Gemma 3n

Resources and Technical Documentation:

Terms of Use: Terms
Authors: Google DeepMind

Model Information

Summary description and brief definition of inputs and outputs.

Description

Gemma is a family of lightweight, state-of-the-art open models from Google, built from the same research and technology used to create the Gemini models. Gemma 3n models are designed for efficient execution on low-resource devices. They are capable of multimodal input, handling text, image, video, and audio input, and generating text outputs, with open weights for pre-trained and instruction-tuned variants. These models were trained with data in over 140 spoken languages.

Gemma 3n models use selective parameter activation technology to reduce resource requirements. This technique allows the models to operate at an effective size of 2B and 4B parameters, which is lower than the total number of parameters they contain. For more information on Gemma 3n's efficient parameter management technology, see the Gemma 3n page.

Inputs and outputs

  • Input:
    • Text string, such as a question, a prompt, or a document to be summarized
    • Images, normalized to 256x256, 512x512, or 768x768 resolution and encoded to 256 tokens each
    • Audio data encoded to 6.25 tokens per second from a single channel
    • Total input context of 32K tokens
  • Output:
    • Generated text in response to the input, such as an answer to a question, analysis of image content, or a summary of a document
    • Total output length up to 32K tokens, subtracting the request input tokens

Usage

Below, there are some code snippets on how to get quickly started with running the model. First, install the Transformers library. Gemma 3n is supported starting from transformers 4.53.0.

$ pip install -U transformers

Then, copy the snippet from the section that is relevant for your use case.

Running with the pipeline API

You can initialize the model and processor for inference with pipeline as follows.

from transformers import pipeline
import torch
pipe = pipeline(
    "image-text-to-text",
    model="google/gemma-3n-e4b-it",
    device="cuda",
    torch_dtype=torch.bfloat16,
)

With instruction-tuned models, you need to use chat templates to process our inputs first. Then, you can pass it to the pipeline.

messages = [
    {
        "role": "system",
        "content": [{"type": "text", "text": "You are a helpful assistant."}]
    },
    {
        "role": "user",
        "content": [
            {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"},
            {"type": "text", "text": "What animal is on the candy?"}
        ]
    }
]
output = pipe(text=messages, max_new_tokens=200)
print(output[0]["generated_text"][-1]["content"])
# Okay, let's take a look!
# Based on the image, the animal on the candy is a **turtle**.
# You can see the shell shape and the head and legs.
Running the model on a single GPU
from transformers import AutoProcessor, Gemma3nForConditionalGeneration
from PIL import Image
import requests
import torch
model_id = "google/gemma-3n-e4b-it"
model = Gemma3nForConditionalGeneration.from_pretrained(model_id, device_map="auto", torch_dtype=torch.bfloat16,).eval()
processor = AutoProcessor.from_pretrained(model_id)
messages = [
    {
        "role": "system",
        "content": [{"type": "text", "text": "You are a helpful assistant."}]
    },
    {
        "role": "user",
        "content": [
            {"type": "image", "image": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/bee.jpg"},
            {"type": "text", "text": "Describe this image in detail."}
        ]
    }
]
inputs = processor.apply_chat_template(
    messages,
    add_generation_prompt=True,
    tokenize=True,
    return_dict=True,
    return_tensors="pt",
).to(model.device)
input_len = inputs["input_ids"].shape[-1]
with torch.inference_mode():
    generation = model.generate(**inputs, max_new_tokens=100, do_sample=False)
    generation = generation[0][input_len:]
decoded = processor.decode(generation, skip_special_tokens=True)
print(decoded)
# **Overall Impression:** The image is a close-up shot of a vibrant garden scene,
# focusing on a cluster of pink cosmos flowers and a busy bumblebee.
# It has a slightly soft, natural feel, likely captured in daylight.

Citation

@article{gemma_3n_2025,
    title={Gemma 3n},
    url={https://ai.google.dev/gemma/docs/gemma-3n},
    publisher={Google DeepMind},
    author={Gemma Team},
    year={2025}
}

Model Data

Data used for model training and how the data was processed.

Training Dataset

These models were trained on a dataset that includes a wide variety of sources totalling approximately 11 trillion tokens. The knowledge cutoff date for the training data was June 2024. Here are the key components:

  • Web Documents: A diverse collection of web text ensures the model is exposed to a broad range of linguistic styles, topics, and vocabulary.

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

How it works

How language models work

Your prompttext / messagesTransformerattention over tokensNext-token loopgenerate + streamResponsetext · tool callsA language model reads your tokens and predicts the next one, again and again, streaming the reply back.
© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms