Model reference · open weights
K2-Type is an open-weight language model from IFM. K2-Type-0.9B (BF16) weighs 2.2 GB; the smallest configuration that runs it is RTX 3060 12 GB.
Summary of the IFM/K2-Type-0.9B model card, 2026-10-04
What it is
| Released by | IFM |
|---|---|
| Released | 2026-09-26 |
| Parameters | 1.1B |
| VRAM | 2.2 GB for the weights |
What it runs on
| Card | Requests at once | Context max | Memory | |
|---|---|---|---|---|
| 8K each | 32K each | |||
| RTX 3060 12 GB | 18 | 4 | all 128K | 11.6 GB |
| RTX 4060 Ti 16 GB | 27 | 6 | all 128K | 15.4 GB |
| RTX 3090 24 GB | 44 | 11 | all 128K | 23.4 GB |
| RTX 4090 24 GB | 43 | 10 | all 128K | 23.4 GB |
| RTX 5090 32 GB | 60 | 15 | all 128K | 31.0 GB |
| L40S 48 GB | 87 | 21 | all 128K | 44.0 GB |
| A100 80 GB | 160 | 40 | all 128K | 78.2 GB |
| H100 80 GB | 154 | 38 | all 128K | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | 187 | 46 | all 128K | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | 216 | 54 | all 128K | 107 GB |
| H200 141 GB | 281 | 70 | all 128K | 138 GB |
| B200 180 GB | 362 | 90 | all 128K | 176 GB |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 3.2 GB | 4.6 GB |
| 5 | 5.1 GB | 12.1 GB |
| 8 | 6.5 GB | 17.8 GB |
| 16 | 10.2 GB | 32.8 GB |
| 32 | 17.8 GB | 62.9 GB |
| 64 | 32.8 GB | 123 GB |
One card, with vLLM's small-card settings.
From the model card
A 0.9B decision model in the style of TypeSafe's Jev (K2-Horizon-0.9B backbone; 1.08B parameters stored, including the base model's unused language-model head). You send one state (text or JSON) and any number of typed questions; it returns a probability for every option of every question from one forward pass. It never generates text.
| Question type | You give | You get |
|---|---|---|
noul | a statement, optional definitions of false / true | P(true) |
choice | 1-255 named options with optional descriptions | the best option, its confidence, all probabilities |
score | 2-255 ordered levels | expected level, probability per level |
Questions share the state but cannot see each other (block-causal attention mask), so adding a question never changes another's answer. Built on IFM/K2-Horizon-0.9B.
JevBench public set (231 items, jevbench commit 26eb72d,
typesafe adapter against this repo's server, one H200, no network):
| Tier | Correct |
|---|---|
standard (72, original.jsonl) | 66 |
easy (48, easy.jsonl) | 47 |
hard (111, hard.jsonl) | 63 |
| Total | 176 / 231 = 0.762 |
Brier 0.328, ECE 0.065; latency p50 27 ms, p95 60 ms per decision; mean 590 input tokens per decision. For reference, public accuracy on the JevBench v1.4 board: Gemma 4 E2B + LoRA (system-one-open) 0.732, decider-2b 0.710, kev 0.6B 0.667, kev 4B 0.662, Qwen3.5-4B entries 0.74-0.82, Jev 1.13.0 0.866. The official JevBench score also uses sealed items; every listed system scores well below its public accuracy there.
Snake: the same weights play Snake from a text board (one choice + four yes/no questions per move): 66 food per game on 12x12 on average (max 101).
pip install -U huggingface_hub
hf download IFM/K2-Type-0.9B --local-dir K2-Type-0.9B && cd K2-Type-0.9B
pip install -r requirements.txt
python -m jev.serve --run . --port 8000 # from this repo's root; needs one CUDA GPU
The first request after start-up takes a few seconds (CUDA warm-up); later requests take about 20-60 ms.
curl -s localhost:8000/v1/systemone -H 'Content-Type: application/json' -d '{
"state": {"subject": "Charged twice", "body": "You billed my card twice for March. Refund one or I cancel."},
"questions": {
"queue": {"type": "choice", "instructions": "Which queue handles this?",
"criteria": {"billing": "Payments and refunds", "technical": "Bugs and login", "general": "Anything else"}},
"angry": {"type": "noul", "instructions": "The customer sounds angry."},
"urgency": {"type": "score", "instructions": "How urgent is it?", "criteria": ["Low", "Normal", "High", "Critical"]}
}}'
The request and response follow TypeSafe's /v1/systemone wire format, so clients written for Jev or Kev work
unchanged. GET /health reports the model name and temperature.
This repository contains the weights and the minimal code needed to serve them (jev/: input encoding, the pointer
head, and the HTTP server). Training code and data are not released at this time.
Requires transformers >= 5.17 (remote code); tested on torch 2.8. Do not use the backbone for text generation: its weights were trained
for the decision head, and the language-model head is left from the base model. Use the jev/ server.
... | question option ... | ..., using five reserved tokens of
the base tokenizer (decision_config.json).pointer_head.safetensors) scores each option's hidden state against the question's hidden state; a softmax at temperature 1.478 gives the probabilities.Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.