Model reference · open weights

gemma-4-qat-UD-japanese-imatrix

gemma-4-qat-UD-japanese-imatrix is an open-weight language model from dahara1, listed in the AxForge catalogue. AxForge can bring it up on EU-owned hardware for you on request — with the licence handled where one is required.

LLMs dahara1 1 variants 5k downloads/mo
Request this model on EU hardware All served models Not on the shared API today — deployed on request.

About

What gemma-4-qat-UD-japanese-imatrix is

gemma-4-12B-it-qat-UD-japanese-imatrix developed by dahara1@webbigdata google/gemma-4-12B-it-qat-q40-unquantized を日本語能力保持を念頭に量子化してサイズを圧縮したモデル google/gemma-4-12B-it-qat-q40-unquantized is a model that compresses the size by quantizing it with the aim of preserving Japanese language proficiency. 💼 導入をご検討の企業様へ / For Enterprise Decision Makers - データは社外に出ません: 本モデルは自社サーバー・PC上で完結して動作します。クラウドへの送信は一切不要です。 - 商用利用可能: Apache 2.0 ライセンス(Google Gemma 4 と同一)。法務確認が容易です。 - 開発元によるサポート: 動作検証済み構成のご提供、本番導入支援、年間サポート契約を有償でご用意しています。→ お問い合わせ - Your data stays on-premise: This model runs entirely on your own servers or PCs. No cloud transmission required. - Commercial use permitted: Apache 2.0 license (same as Google Gemma 4). - Developer support available: Verified configurations, production deployment assistance, and annual support contracts. → Contact us 特徴 / Features 一言で言えば沢山の細かい改善をしてサイズを1/4に圧縮しつつ日本語能力を強化した強力なモデルです。CPUのみでも動かす事ができます。 In short, it's a powerful model that has undergone numerous minor improvements, compressing its size to one-quarter while enhancing its Japanese language capabilities. It can even run on a CPU alone. このモデルの特徴 - 日本語性能を重点的に保持するように独自の動的量子化をしています - ベンチマークをしっかりと行って堅牢性・優位性を確かめています Features of this gguf - We use dynamic quantization to prioritize and maintain the performance of the Japanese language. - We conduct thorough benchmarks to confirm its robustness and superiority. クイックスタート GPUがなくても動きますが、推奨サイズであるgemma-4-12B-it-qat-ja-UD-Q4KXL.ggufではシステムメモリは16GB以上、ディスク容量が7GB以上必要です。 It will run without a GPU, but the recommended model, gemma-4-12B-it-qat-ja-UD-Q4KXL.gguf, requires at least 12GB of system memory and at least 7GB of disk space. 非常に非力なマシンの場合はColab-CLIの力を借りてGemma4を動かす方法もあるので参考にしてください。 If you have a very underpowered machine, you can also refer to How to run Gemma4 with the help of Colab-CLI. llama.cppを使います。直近でGemma 4対応のアップデートがいくつかありました。常に最新版を使う事をおすすめします。(本件の動作確認はversion: 9556 (19bba67c1) で行っています) llama.cpp以外のツールでも動く可能性はありますが、他のツールは製作者が意図していない設定で誤動作をする場合があるので留意してください We will be using llama.cpp. There have been several recent updates to support Gemma 4. So it is recommended to always use the latest version. (This issue was confirmed to work with version: 9556 (

Summarised from the published model card. Read the full card on the HuggingFace links below.

Specifications

What it is

Makerdahara1
TypeLanguage models
Variants1
Based ongoogle/gemma-4-12B-it-qat-q4_0-unquantized
Released2026-06-09
Popularity5k downloads / month
Likes14
LicenceOpen weights

How it works

How language models work

Your prompttext / messagesTransformerattention over tokensNext-token loopgenerate + streamResponsetext · tool callsA language model reads your tokens and predicts the next one, again and again, streaming the reply back.

Variants

Sizes & precisions

Open weights ship in several sizes and precisions. One page, all the variants — pick the one that fits your GPU. VRAM figures are estimates from model size.

VariantParamsPrecisionVRAMFits 16 GBWeights
gemma-4-12B-it-qat-UD-japanese-imatrixBF16Weights ↗

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys gemma-4-qat-ud-japanese-imatrix for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (gemma-4-qat-ud-japanese-imatrix below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/chat/completions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"gemma-4-qat-ud-japanese-imatrix","messages":[{"role":"user","content":"Hello"}]}'

Details

Languages, data & research

Languages

ja en

Tags

gguf japanese llama.cpp gemma4 MTP unsloth any-to-any ja en endpoints_compatible imatrix conversational

Licence

Open weights

Open weights under apache-2.0 — commercial use is permitted. Deploy it on AxForge EU hardware on request. Read the licence ↗

Sources

Weights & code

Want gemma-4-qat-UD-japanese-imatrix on EU-owned hardware?

Request this model on EU hardware See what’s served now

Explore

More language models

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms