Model reference · open weights

sup-simcse-ja

Available as managed deployment Embeddings cl-nagoya Embeddings 1 variants 1k dl/mo

sup-simcse-ja is an open-weight embedding model from cl-nagoya. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

Released bycl-nagoya
TypeEmbedding models
TaskEmbeddings
Context512 tokens
Runs withsentence-transformers
Released2023-10-02
Popularity1k downloads / month
LicenceOpen weights

About

What sup-simcse-ja is

Usage (Sentence-Transformers)

Using this model becomes easy when you have sentence-transformers installed:

pip install -U fugashi[unidic-lite] sentence-transformers

Then you can use the model like this:

Read the full model card
from sentence_transformers import SentenceTransformer
sentences = ["こんにちは、世界!", "文埋め込み最高!文埋め込み最高と叫びなさい", "極度乾燥しなさい"]

model = SentenceTransformer("cl-nagoya/sup-simcse-ja-base")
embeddings = model.encode(sentences)
print(embeddings)

Usage (HuggingFace Transformers)

Without sentence-transformers, you can use the model like this: First, you pass your input through the transformer model, then you have to apply the right pooling-operation on-top of the contextualized word embeddings.

from transformers import AutoTokenizer, AutoModel
import torch

def cls_pooling(model_output, attention_mask):
    return model_output[0][:,0]

# Sentences we want sentence embeddings for
sentences = ['This is an example sentence', 'Each sentence is converted']

# Load model from HuggingFace Hub
tokenizer = AutoTokenizer.from_pretrained("cl-nagoya/sup-simcse-ja-base")
model = AutoModel.from_pretrained("cl-nagoya/sup-simcse-ja-base")

# Tokenize sentences
encoded_input = tokenizer(sentences, padding=True, truncation=True, return_tensors='pt')

# Compute token embeddings
with torch.no_grad():
    model_output = model(**encoded_input)

# Perform pooling. In this case, cls pooling.
sentence_embeddings = cls_pooling(model_output, encoded_input['attention_mask'])

print("Sentence embeddings:")
print(sentence_embeddings)

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'max_seq_length': 512, 'do_lower_case': False}) with Transformer model: BertModel
  (1): Pooling({'word_embedding_dimension': 768, 'pooling_mode_cls_token': True, 'pooling_mode_mean_tokens': False, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False})
)

Model Summary

  • Fine-tuning method: Supervised SimCSE
  • Base model: cl-tohoku/bert-base-japanese-v3
  • Training dataset: JSNLI
  • Pooling strategy: cls (with an extra MLP layer only during training)
  • Hidden size: 768
  • Learning rate: 5e-5
  • Batch size: 512
  • Temperature: 0.05
  • Max sequence length: 64
  • Number of training examples: 2^20
  • Validation interval (steps): 2^6
  • Warmup ratio: 0.1
  • Dtype: BFloat16

See the GitHub repository for a detailed experimental setup.

Citing & Authors

@misc{
  hayato-tsukagoshi-2023-simple-simcse-ja,
  author = {Hayato Tsukagoshi},
  title = {Japanese Simple-SimCSE},
  year = {2023},
  publisher = {GitHub},
  journal = {GitHub repository},
  howpublished = {\url{https://github.com/hppRC/simple-simcse-ja}}
}

From the published model card. Full card on the HuggingFace links in the sidebar.

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys sup-simcse-ja for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (sup-simcse-ja below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/embeddings \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"sup-simcse-ja","input":"text to embed"}'

Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms