Model reference · open weights

K2-Type

LLMs IFM Text gen 1 build Open weights 1k dl/mo

K2-Type is an open-weight language model from IFM. K2-Type-0.9B (BF16) weighs 2.2 GB; the smallest configuration that runs it is RTX 3060 12 GB.

  • K2-Type is a 1.1B parameter decision model developed by IFM that returns probabilities for typed questions from a single forward pass without generating text.
  • It supports a 131072 token context length, operates in English, and is licensed under apache-2.0.
  • The model is designed for tasks such as classification, choice selection, and scoring, and it achieves a public accuracy of 0.762 on the JevBench dataset.

Summary of the IFM/K2-Type-0.9B model card, 2026-10-04

What it is

Released byIFM
Released2026-09-26
Parameters1.1B
VRAM2.2 GB for the weights

What it runs on

Memory and cards for K2-Type-0.9B (BF16)

2.2 GBweights, file size
57 MBcache per 1K tokens
562 MBruntime overhead, at least
131,072 tokenscontext max
CardRequests at onceContext maxMemory
8K each32K each
RTX 3060 12 GB184all 128K11.6 GB
RTX 4060 Ti 16 GB276all 128K15.4 GB
RTX 3090 24 GB4411all 128K23.4 GB
RTX 4090 24 GB4310all 128K23.4 GB
RTX 5090 32 GB6015all 128K31.0 GB
L40S 48 GB8721all 128K44.0 GB
A100 80 GB16040all 128K78.2 GB
H100 80 GB15438all 128K78.1 GB
RTX PRO 6000 Blackwell 96 GB18746all 128K93.8 GB
DGX Spark (GB10) 128 GB unified21654all 128K107 GB
H200 141 GB28170all 128K138 GB
B200 180 GB36290all 128K176 GB
Memory needed at each load
Requests at once8K tokens each32K tokens each
13.2 GB4.6 GB
55.1 GB12.1 GB
86.5 GB17.8 GB
1610.2 GB32.8 GB
3217.8 GB62.9 GB
6432.8 GB123 GB

One card, with vLLM's small-card settings.

From the model card

What IFM says about K2-Type

Read the model card

A 0.9B decision model in the style of TypeSafe's Jev (K2-Horizon-0.9B backbone; 1.08B parameters stored, including the base model's unused language-model head). You send one state (text or JSON) and any number of typed questions; it returns a probability for every option of every question from one forward pass. It never generates text.

Question typeYou giveYou get
noula statement, optional definitions of false / trueP(true)
choice1-255 named options with optional descriptionsthe best option, its confidence, all probabilities
score2-255 ordered levelsexpected level, probability per level

Questions share the state but cannot see each other (block-causal attention mask), so adding a question never changes another's answer. Built on IFM/K2-Horizon-0.9B.

Results

JevBench public set (231 items, jevbench commit 26eb72d, typesafe adapter against this repo's server, one H200, no network):

TierCorrect
standard (72, original.jsonl)66
easy (48, easy.jsonl)47
hard (111, hard.jsonl)63
Total176 / 231 = 0.762

Brier 0.328, ECE 0.065; latency p50 27 ms, p95 60 ms per decision; mean 590 input tokens per decision. For reference, public accuracy on the JevBench v1.4 board: Gemma 4 E2B + LoRA (system-one-open) 0.732, decider-2b 0.710, kev 0.6B 0.667, kev 4B 0.662, Qwen3.5-4B entries 0.74-0.82, Jev 1.13.0 0.866. The official JevBench score also uses sealed items; every listed system scores well below its public accuracy there.

Snake: the same weights play Snake from a text board (one choice + four yes/no questions per move): 66 food per game on 12x12 on average (max 101).

Quickstart

pip install -U huggingface_hub
hf download IFM/K2-Type-0.9B --local-dir K2-Type-0.9B && cd K2-Type-0.9B
pip install -r requirements.txt
python -m jev.serve --run . --port 8000          # from this repo's root; needs one CUDA GPU

The first request after start-up takes a few seconds (CUDA warm-up); later requests take about 20-60 ms.

curl -s localhost:8000/v1/systemone -H 'Content-Type: application/json' -d '{
  "state": {"subject": "Charged twice", "body": "You billed my card twice for March. Refund one or I cancel."},
  "questions": {
    "queue":  {"type": "choice", "instructions": "Which queue handles this?",
               "criteria": {"billing": "Payments and refunds", "technical": "Bugs and login", "general": "Anything else"}},
    "angry":  {"type": "noul", "instructions": "The customer sounds angry."},
    "urgency": {"type": "score", "instructions": "How urgent is it?", "criteria": ["Low", "Normal", "High", "Critical"]}
  }}'

The request and response follow TypeSafe's /v1/systemone wire format, so clients written for Jev or Kev work unchanged. GET /health reports the model name and temperature.

This repository contains the weights and the minimal code needed to serve them (jev/: input encoding, the pointer head, and the HTTP server). Training code and data are not released at this time.

Requires transformers >= 5.17 (remote code); tested on torch 2.8. Do not use the backbone for text generation: its weights were trained for the decision head, and the language-model head is left from the base model. Use the jev/ server.

How it works

  • Input layout: ... | question option ... | ..., using five reserved tokens of the base tokenizer (decision_config.json).
  • A pointer head (pointer_head.safetensors) scores each option's hidden state against the question's hidden state; a softmax at temperature 1.478 gives the probabilities.
  • Training, in short: full fine-tune on ~354k decision records (public classification/NLI/QA sets, Kev's decision data, game positions with exact or search labels, and synthetic decision items written and blind-verified by a large model), targets mixed with soft labels from a 7B teacher decision model; then 50 iterations of PPO on Snake; then a temperature refit on held-out calibration data. Training data were checked against every Kev evaluation suite and the 231 JevBench public items: no exact or containment overlap.

Limits

  • One pass, no reasoning: multi-step arithmetic, date differences and very long documents (> 8192 tokens) are weak.
  • Calibration is fitted on Kev's calibration suite; on other distributions probabilities can be over- or under-confident (JevBench public ECE 0.065).
  • Answers are bounded by the options you give; it cannot say "none of these" unless you offer that option.

Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms