Model reference · open weights

llm-jp-4.1-thinking

LLMs llm-jp Text gen · MoE 3 builds Open weights 921 dl/mo

llm-jp-4.1-thinking is an open-weight language model from llm-jp. llm-jp-4.1-32b-a3b-thinking (BF16) weighs 64.3 GB; the smallest configuration that runs it is H100 80 GB.

  • llm-jp-4.1-8b-thinking is a text-generation model developed by the Research and Development Center for Large Language Models at the National Institute of Informatics.
  • It is a dense transformer-based language model with 8.6 billion parameters and a context length of 65,536 tokens.
  • The model supports English and Japanese and is distributed under the Apache 2.0 license.

Summary of the llm-jp/llm-jp-4.1-8b-thinking model card, 2026-10-04. The estimate below is for another build of the family.

What it is

Released byllm-jp
Released2026-09-15
Parameters8.6B
VRAM64.3 GB for the weights

What it runs on

Memory and cards for llm-jp-4.1-32b-a3b-thinking (BF16)

64.3 GBweights, file size
66 MBcache per 1K tokens
1.3 GBruntime overhead, at least
65,536 tokenscontext max
CardRequests at onceContext maxMemory
8K each32K each
RTX 3060 12 GB … L40S 48 GB
6 smaller cards
———
A100 80 GB235all 64K78.2 GB
H100 80 GB112all 64K78.1 GB
RTX PRO 6000 Blackwell 96 GB4010all 64K93.8 GB
DGX Spark (GB10) 128 GB unified6516all 64K107 GB
H200 141 GB12230all 64K138 GB
B200 180 GB18847all 64K176 GB
2× L40S 48 GB
tensor parallel
399all 64K44.0 GB a card
4× RTX 4090 24 GB
tensor parallel
4411all 64K23.4 GB a card
4× RTX 3090 24 GB
tensor parallel
4411all 64K23.4 GB a card
4× RTX 5090 32 GB
tensor parallel
10125all 64K31.0 GB a card
2× H100 80 GB
tensor parallel
14235all 64K78.1 GB a card
Memory needed at each load
Requests at once8K tokens each32K tokens each
166.1 GB67.7 GB
568.3 GB76.3 GB
869.9 GB82.8 GB
1674.2 GB100.0 GB
3282.8 GB134 GB
64100.0 GB203 GB

One card, with vLLM's small-card settings.

Builds

Sizes, precisions & builds

BuildParamsPrecisionWeightsSmallest setup
llm-jp-4.1-8b-thinking ↗ 8.6BBF16 17.2 GB2× RTX 3060 12 GB
llm-jp-4.1-33b-thinking ↗ 33.2BBF16 66.4 GBH100 80 GB
llm-jp-4.1-32b-a3b-thinking (above) ↗ 32.1BBF16 64.3 GBH100 80 GB

From the model card

What llm-jp says about llm-jp-4.1-thinking

Read the model card

LLM-jp-4.1 is a series of large language models developed by the Research and Development Center for Large Language Models at the National Institute of Informatics.

This repository provides the llm-jp-4.1-8b-thinking model. For an overview of the LLM-jp-4.1 models across different parameter sizes, please refer to:

Base models are trained with pre-training and mid-training only. Post-trained models are aligned using supervised fine-tuning (SFT) and direct preference optimization (DPO), without reinforcement learning.

For more details on the training procedures and evaluation results, please refer to our technical blog (in Japanese).

For practical usage examples and detailed instructions on how to use the models, please also refer to our cookbook.

To support the continued development of LLM-jp, we would greatly appreciate it if you could share how you utilize LLM-jp outcomes via the survey form.

Usage

Please refer to our cookbook for practical usage examples and detailed instructions on how to use the models.

Model Details

  • Model type: Transformer-based Language Model
  • Architectures:

Dense model:

ParamsLayersHidden sizeHeadsContext lengthEmbedding parametersNon-embedding parametersTotal parameters
8B324,0963265,536805,306,3687,784,894,4648,590,200,832
33B645,1204065,5361,006,632,96032,212,915,20033,219,548,160

MoE model:

ParamsLayersHidden sizeHeadsRouted ExpertsActivated ExpertsContext lengthEmbedding parametersNon-embedding parametersActivated parametersTotal parameters
32B-A3B322,56040128865,536503,316,48031,635,712,5123,827,476,99232,139,028,992

Tokenizer

The tokenizer of this model is based on a Unigram byte-fallback model implemented with huggingface/tokenizers. The vocabulary entries were converted from llm-jp-tokenizer v4.0. Please refer to README.md of llm-jp-tokenizer for details on the vocabulary construction procedure (the pure SentencePiece training does not reproduce our vocabulary).

[!NOTE] The chat template of this model is designed to be compatible with the OpenAI Harmony response format. However, the tokenizer differs from the one assumed by the openai-harmony library, and therefore direct tokenization with openai-harmony is not supported. For correct behavior, please use the tokenizer provided with this model. For detailed usage, please refer to our cookbook.

Training

Pre-training

This model was trained through a multi-stage pipeline consisting of pre-training and mid-training phases, using a total of 11.7T tokens.

The corpora used for pre-training and mid-training are publicly available at the following links:

[!NOTE] Although most of the corpora have been released, some portions are excluded from public release due to licensing constraints.

Post-training

We have fine-tuned the pre-trained checkpoint using SFT and further aligned it with DPO.

The datasets used for post-training are also publicly available at the following links:

Evaluation

We evaluated llm-jp-4.1 on a variety of benchmarks covering general capabilities, safety, and tool calling.

For more detailed evaluation results and analysis, please refer to our technical blog.

swallow-evaluation-instruct

We evaluated the models on a range of benchmarks covering the following six categories:

  • Math
    • Math 500
    • AIME 2024 (pass@1, pass@32)
    • AIME 2025 (pass@1, pass@32)
    • AIME 2026 (pass@1, pass@32)
    • MCLM Math 100 (pass@1, pass@4)
    • PolyMath JA High
    • PolyMath JA Top
  • Science
    • GPQA Diamond (pass@1, pass@4)
    • JGPQA Diamond
  • Knowledge & QA
    • JAM-CQA
    • JEMHopQA
    • JMMLU
    • MMLU-ProX JA
    • MMLU-ProX EN
  • Code
    • LiveCodeBench v6 (pass@1, pass@10)
    • JHumanEval (pass@1, pass@10)
    • HumanEval+ (pass@1, pass@10)
  • Instruction Following (IF)
    • MIFEval JA
    • IFBench
  • Machine Translation (MT)
    • WMT20 EN-JA
    • WMT20 JA-EN

For LLM-jp and gpt-oss models, reasoning_effort was set to high. For Olmo-3-7B-Think, Olmo-3.1-32B-Think, Qwen3, Qwen3.5, Qwen3.6, and Gemma 4, enable_thinking was set to True. For Qwen3.8-27B and Muse-Glimmer-30B, reasoning_effort was set to xhigh.

The figure below shows the average score across the benchmarks in each category.

llm-jp-judge

We evaluated the models using an LLM-as-a-Judge framework on the following benchmarks:

  • MT-Bench (JA/EN): A benchmark for measuring multi-turn conversational task-solving ability.
  • AnswerCarefully: A benchmark for evaluating safety in Japanese. We used 3

Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms