Model reference · open weights

gemma-4-text-only

Embeddings principled-intelligence Embeddings 1 build Open weights 856 dl/mo

gemma-4-text-only is an open-weight embedding model from principled-intelligence. gemma-4-E4B-it-text-only (BF16) weighs 15.0 GB; the smallest configuration that runs it is RTX 4090 24 GB.

What it is

Released byprincipled-intelligence
TypeEmbedding models
TaskEmbeddings
Parameters (lead)7.5B
Context131,072 tokens
Runs withtransformers
Released2026-04-02
Popularity856 downloads / month
Weights15.0 GB (gemma-4-E4B-it-text-only (BF16), file size)
LicenceOpen weights

What it runs on

Memory and cards for gemma-4-E4B-it-text-only (BF16)

Weights 15.0 GB (file size) · overhead about 1.1 GB.

CardRunsCounted
memory
RTX 3060 12 GB … RTX 4060 Ti 16 GBdoes not fit
RTX 3090 24 GBfits23.4 GB
RTX 4090 24 GBfits23.4 GB
RTX 5090 32 GBfits31.0 GB
L40S 48 GBfits44.0 GB
A100 80 GBfits78.2 GB
H100 80 GBfits78.1 GB
RTX PRO 6000 Blackwell 96 GBfits93.8 GB
DGX Spark (GB10) 128 GB unifiedfits107 GB
H200 141 GBfits138 GB
B200 180 GBfits176 GB

Estimates, not measurements: the weights are the build's file size. No cache grows with use; a batch of inputs needs working memory of its own. Counted memory is 92 % of what CUDA reports for the card.

From the model card

What principled-intelligence says about gemma-4-text-only

If all you need is text, these are the Gemma 4 models for you.

Trimmed checkpoints of the Gemma 4 model family with vision and audio encoder weights removed — smaller files, lower VRAM, drop-in text-only replacement.

⚠️ Disclaimer: These models were tested exclusively with HuggingFace Transformers. vLLM, SGLang, llama.cpp, Ollama, and other inference engines are not supported yet — partly because Transformers support for Gemma 4 is still cooking in those projects, and partly because we just threw these checkpoints on the Hub while messing around in the lab. If you get any of these running on other engines, we'd love to hear about it — open a discussion or drop a community post. We didn't set out to build a production-ready model zoo; we just left the oven door open. Use accordingly.

For official details on the Gemma 4 model family — architecture, benchmarks, training data, and intended use — see the original Gemma 4 E4B-it model card.

Read the full model card

How It Works

The Gemma 4 E4B architecture consists of a vision encoder, an audio encoder, and a language model sharing a single checkpoint. During text-only inference the vision and audio encoders are never called, but their weights are still loaded into memory. By loading the checkpoint with the causal LM class instead of the full conditional generation class, HuggingFace Transformers instantiates only the language model component. Re-saving that model produces a checkpoint with no vision or audio weights, which can subsequently be loaded with the standard AutoModelForCausalLM interface.

Why bother?

  • Lower VRAM — vision and audio encoder weights are freed, reducing peak memory usage
  • Smaller checkpoints — faster downloads and storage savings
  • Simpler loading — standard AutoModelForCausalLM, no multimodal dependencies
  • Drop-in replacement — identical tokenizer, same chat template, same text generation behavior as the original Gemma 4 models

Available Models

ModelHuggingFace Hub
Gemma-4-E2B-it-text-onlyprincipled-intelligence/gemma-4-E2B-it-text-only
Gemma-4-E4B-it-text-onlyprincipled-intelligence/gemma-4-E4B-it-text-only

Size Reduction

We compared the text-only checkpoint against the original Gemma 4 E4B-it across two metrics: peak VRAM usage when loaded in bfloat16 with device_map="auto", and total parameter count.

Gemma 4 E4B-it vs. Gemma 4 E4B-it-text-only

MetricGemma 4 E4B-itText-OnlyReduction
VRAM (GB)15.915.0~6%
Parameters (B)8.007.52~6%
File size (GB)16.0015.00~6%

Note: The "E" in E4B stands for "effective" parameters. The Gemma 4 E4B architecture uses Per-Layer Embeddings (PLE) to maximize parameter efficiency on-device — the total parameter count is higher than the effective size. The text-only variant removes the vision and audio encoder weights while preserving the full language model, including all PLE parameters.

Quickstart

The latest transformers is required:

uv pip install transformers>=5.5.0

Load and run inference exactly like any causal LM:

from transformers import pipeline

pipe = pipeline(
    "text-generation",
    model="principled-intelligence/gemma-4-E4B-it-text-only",
    device_map="auto",
)

messages = [{"role": "user", "content": "What is the capital of Italy?"}]
print(pipe(messages, max_new_tokens=512))
from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "principled-intelligence/gemma-4-E4B-it-text-only"

tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, device_map="auto")

messages = [
    {"role": "user", "content": "What is the capital of Italy?"},
]

text = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)

output_ids = model.generate(**inputs, max_new_tokens=512)
response = tokenizer.decode(output_ids[0][inputs.input_ids.shape[-1]:], skip_special_tokens=True)
print(response)

Gemma 4 thinks by default, generating internal reasoning content before the final response. Thinking is enabled by including the `` token at the start of the system prompt. To disable thinking, remove the token. Many libraries like Transformers handle this via the chat template for you.

Contributing

Contributions are welcome! Whether it's getting these checkpoints running on vLLM, SGLang, llama.cpp, Ollama, or something else entirely — we'd love your help. Bug reports, compatibility notes, and PRs are all appreciated. Open a discussion or community post and let us know what you find.

License

These checkpoints are released under the Apache 2.0 License, consistent with the original Gemma 4 models.


Made with love from Principled Intelligence ❤️

Learn more about what we build in Principled Intelligence on our website.

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms