Model reference · open weights

gemma-3n

LLMs google Vision + text 4 builds Open, with conditions 210k dl/mo

gemma-3n is an open-weight language model from Google. gemma-3n-E2B-it (BF16) weighs 10.9 GB; the smallest configuration that runs it is RTX 4060 Ti 16 GB.

Gemma 3n E2B IT is an instruction-tuned multimodal model by Google that accepts text, image, video, and audio inputs to generate text. It features a raw parameter count of 6B but operates with an effective size of 2B parameters to support efficient execution on low-resource devices. The model supports a 32K token context window, was trained on data in over 140 languages, and is released under the gemma license.

Summary of the google/gemma-3n-E2B-it model card, 2026-10-01

What it is

Released byGoogle
TypeLanguage models
TaskVision + text
Parameters (lead)5.4B
Runs withtransformers
Based ongoogle/gemma-3n-E4B-it
Released2025-06-12
Popularity210k downloads / month
Weights10.9 GB (gemma-3n-E2B-it (BF16), file size)
LicenceOpen, with conditions

What it runs on

Memory and cards for gemma-3n-E2B-it (BF16)

Weights 10.9 GB (file size) · runtime overhead from 762 MB on a small card.

How much memory each request adds is not estimated yet for this architecture — only the weights are. They need the cards below at the least, plus room for the context.

CardThe weights alone
RTX 3060 12 GBdoes not fit
RTX 4060 Ti 16 GBfits
RTX 3090 24 GBfits
RTX 4090 24 GBfits
RTX 5090 32 GBfits
L40S 48 GBfits
A100 80 GBfits
H100 80 GBfits
RTX PRO 6000 Blackwell 96 GBfits
DGX Spark (GB10) 128 GB unifiedfits
H200 141 GBfits
B200 180 GBfits

Builds

Sizes, precisions & builds

BuildParametersPrecisionWeightsSmallest configuration (1 request, 8K)
gemma-3n-E2B-it (above) ↗ 5.4BBF16 10.9 GBRTX 4060 Ti 16 GB
gemma-3n-E4B-it ↗ 7.8BBF16 15.7 GBRTX 4090 24 GB
gemma-3n-E4B-it ↗
packaged by unsloth
8.4BBF16 15.7 GB2× RTX 3060 12 GB
gemma-3n-E2B-it-GGUF ↗
packaged by unsloth
24 builds: IQ2_XXS 2.1 GB … F16 8.9 GB
—GGUF 3.0 GBRTX 3060 12 GB

Weights from each build's files as published; ≈ = calculated from the parameter count where the files have not been read. A build's name shows what it runs on; ↗ opens it on Hugging Face.

From the model card

What Google says about gemma-3n

Read the model card
  • Currently only text is supported.
  • Ollama: ollama run hf.co/unsloth/gemma-3n-E4B-it:Q4_K_XL - auto-sets correct chat template and settings
  • Set temperature = 1.0, top_k = 64, top_p = 0.95, min_p = 0.0
  • Gemma 3n max tokens (context length): 32K. Gemma 3n chat template:

🦥 Fine-tune Gemma 3n with Unsloth

Gemma-3n-E4B model card

Model Page: Gemma 3n

Resources and Technical Documentation:

Terms of Use: Terms
Authors: Google DeepMind

Model Information

Summary description and brief definition of inputs and outputs.

Description

Gemma is a family of lightweight, state-of-the-art open models from Google, built from the same research and technology used to create the Gemini models. Gemma 3n models are designed for efficient execution on low-resource devices. They are capable of multimodal input, handling text, image, video, and audio input, and generating text outputs, with open weights for pre-trained and instruction-tuned variants. These models were trained with data in over 140 spoken languages.

Gemma 3n models use selective parameter activation technology to reduce resource requirements. This technique allows the models to operate at an effective size of 2B and 4B parameters, which is lower than the total number of parameters they contain. For more information on Gemma 3n's efficient parameter management technology, see the Gemma 3n page.

Inputs and outputs

  • Input:
    • Text string, such as a question, a prompt, or a document to be summarized
    • Images, normalized to 256x256, 512x512, or 768x768 resolution and encoded to 256 tokens each
    • Audio data encoded to 6.25 tokens per second from a single channel
    • Total input context of 32K tokens
  • Output:
    • Generated text in response to the input, such as an answer to a question, analysis of image content, or a summary of a document
    • Total output length up to 32K tokens, subtracting the request input tokens

Usage

Below, there are some code snippets on how to get quickly started with running the model. First, install the Transformers library. Gemma 3n is supported starting from transformers 4.53.0.

$ pip install -U transformers

Then, copy the snippet from the section that is relevant for your use case.

Running with the pipeline API

You can initialize the model and processor for inference with pipeline as follows.

from transformers import pipeline
import torch
pipe = pipeline(
    "image-text-to-text",
    model="google/gemma-3n-e4b-it",
    device="cuda",
    torch_dtype=torch.bfloat16,
)

With instruction-tuned models, you need to use chat templates to process our inputs first. Then, you can pass it to the pipeline.

messages = [
    {
        "role": "system",
        "content": [{"type": "text", "text": "You are a helpful assistant."}]
    },
    {
        "role": "user",
        "content": [
            {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"},
            {"type": "text", "text": "What animal is on the candy?"}
        ]
    }
]
output = pipe(text=messages, max_new_tokens=200)
print(output[0]["generated_text"][-1]["content"])
# Okay, let's take a look!
# Based on the image, the animal on the candy is a **turtle**.
# You can see the shell shape and the head and legs.
Running the model on a single GPU
from transformers import AutoProcessor, Gemma3nForConditionalGeneration
from PIL import Image
import requests
import torch
model_id = "google/gemma-3n-e4b-it"
model = Gemma3nForConditionalGeneration.from_pretrained(model_id, device_map="auto", torch_dtype=torch.bfloat16,).eval()
processor = AutoProcessor.from_pretrained(model_id)
messages = [
    {
        "role": "system",
        "content": [{"type": "text", "text": "You are a helpful assistant."}]
    },
    {
        "role": "user",
        "content": [
            {"type": "image", "image": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/bee.jpg"},
            {"type": "text", "text": "Describe this image in detail."}
        ]
    }
]
inputs = processor.apply_chat_template(
    messages,
    add_generation_prompt=True,
    tokenize=True,
    return_dict=True,
    return_tensors="pt",
).to(model.device)
input_len = inputs["input_ids"].shape[-1]
with torch.inference_mode():
    generation = model.generate(**inputs, max_new_tokens=100, do_sample=False)
    generation = generation[0][input_len:]
decoded = processor.decode(generation, skip_special_tokens=True)
print(decoded)
# **Overall Impression:** The image is a close-up shot of a vibrant garden scene,
# focusing on a cluster of pink cosmos flowers and a busy bumblebee.
# It has a slightly soft, natural feel, likely captured in daylight.

Citation

@article{gemma_3n_2025,
    title={Gemma 3n},
    url={https://ai.google.dev/gemma/docs/gemma-3n},
    publisher={Google DeepMind},
    author={Gemma Team},
    year={2025}
}

Model Data

Data used for model training and how the data was processed.

Training Dataset

These models were trained on a dataset that includes a wide variety of sources totalling approximately 11 trillion tokens. The knowledge cutoff date for the training data was June 2024. Here are the key components:

  • Web Documents: A diverse collection of web text ensures the model is exposed to a broad range of linguistic styles, topics, and vocabulary.

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

How it works

How language models work

Your prompttext / messagesTransformerattention over tokensNext-token loopgenerate + streamResponsetext · tool callsA language model reads your tokens and predicts the next one, again and again, streaming the reply back.
© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms