Model reference · open weights
japanese-clip-vit-b-16 is an open-weight embedding model from rinna. japanese-clip-vit-b-16 (FP32) weighs 393 MB; the smallest configuration that runs it is RTX 3060 12 GB.
japanese-clip-vit-b-16 is a feature-extraction model developed by rinna for processing Japanese text and images. It utilizes a ViT-B/16 Transformer for image encoding and a 12-layer BERT for text encoding, with 197M parameters and a context length of 512 tokens. The model was trained on the CC12M dataset with Japanese-translated captions and is released under the Apache 2.0 license.
Summary of the rinna/japanese-clip-vit-b-16 model card, 2026-10-01
What it is
| Released by | rinna |
|---|---|
| Type | Embedding models |
| Task | Embeddings |
| Parameters (lead) | 197M |
| Context | 512 tokens |
| Runs with | transformers |
| Released | 2022-04-27 |
| Popularity | 37k downloads / month |
| Weights | 393 MB (japanese-clip-vit-b-16 (FP32), file size) |
| Licence | Open weights |
What it runs on
Weights 393 MB (file size) · overhead about 1.1 GB.
| Card | Runs | Counted memory |
|---|---|---|
| RTX 3060 12 GB | fits | 11.6 GB |
| RTX 4060 Ti 16 GB | fits | 15.4 GB |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
Estimates, not measurements: the weights are the build's file size. No cache grows with use; a batch of inputs needs working memory of its own. Counted memory is 92 % of what CUDA reports for the card.
From the model card
This is a Japanese CLIP (Contrastive Language-Image Pre-Training) model trained by rinna Co., Ltd..
Please see japanese-clip for the other available models.
$ pip install git+https://github.com/rinnakk/japanese-clip.git
import io
import requests
from PIL import Image
import torch
import japanese_clip as ja_clip
device = "cuda" if torch.cuda.is_available() else "cpu"
model, preprocess = ja_clip.load("rinna/japanese-clip-vit-b-16", cache_dir="/tmp/japanese_clip", device=device)
tokenizer = ja_clip.load_tokenizer()
img = Image.open(io.BytesIO(requests.get('https://images.pexels.com/photos/2253275/pexels-photo-2253275.jpeg?auto=compress&cs=tinysrgb&dpr=3&h=750&w=1260').content))
image = preprocess(img).unsqueeze(0).to(device)
encodings = ja_clip.tokenize(
texts=["犬", "猫", "象"],
max_seq_len=77,
device=device,
tokenizer=tokenizer, # this is optional. if you don't pass, load tokenizer each time
)
with torch.no_grad():
image_features = model.get_image_features(image)
text_features = model.get_text_features(**encodings)
text_probs = (100.0 * image_features @ text_features.T).softmax(dim=-1)
print("Label probs:", text_probs) # prints: [[1.0, 0.0, 0.0]]
The model was trained a ViT-B/16 Transformer architecture as an image encoder and uses a 12-layer BERT as a text encoder. The image encoder was initialized from the AugReg vit-base-patch16-224 model.
The model was trained on CC12M translated the captions to Japanese.
May 12, 2022
@misc{rinna-japanese-clip-vit-b-16,
title = {rinna/japanese-clip-vit-b-16},
author = {Shing, Makoto and Zhao, Tianyu and Sawada, Kei},
url = {https://huggingface.co/rinna/japanese-clip-vit-b-16}
}
@inproceedings{sawada2024release,
title = {Release of Pre-Trained Models for the {J}apanese Language},
author = {Sawada, Kei and Zhao, Tianyu and Shing, Makoto and Mitsui, Kentaro and Kaga, Akio and Hono, Yukiya and Wakatsuki, Toshiaki and Mitsuda, Koh},
booktitle = {Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)},
month = {5},
year = {2024},
pages = {13898--13905},
url = {https://aclanthology.org/2024.lrec-main.1213},
note = {\url{https://arxiv.org/abs/2404.01657}}
}
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.