Model reference · open weights
ForgePlex-M2 is an open-weight language model from ForgeWorks. ForgePlex-M2-9M (FP32) weighs 20 MB; the smallest configuration that runs it is RTX 3060 12 GB.
Summary of the ForgeWorks/ForgePlex-M2-9M model card, 2026-10-04
What it is
| Released by | ForgeWorks |
|---|---|
| Released | 2026-09-29 |
| Parameters | 10M |
| VRAM | 20 MB for the weights |
What it runs on
| Card | Requests at once | Context max | Memory |
|---|---|---|---|
| 1K each | |||
| RTX 3060 12 GB | 1000+ | all 1K | 11.6 GB |
| RTX 4060 Ti 16 GB | 1000+ | all 1K | 15.4 GB |
| RTX 3090 24 GB | 1000+ | all 1K | 23.4 GB |
| RTX 4090 24 GB | 1000+ | all 1K | 23.4 GB |
| RTX 5090 32 GB | 1000+ | all 1K | 31.0 GB |
| L40S 48 GB | 1000+ | all 1K | 44.0 GB |
| A100 80 GB | 1000+ | all 1K | 78.2 GB |
| H100 80 GB | 1000+ | all 1K | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | 1000+ | all 1K | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | 1000+ | all 1K | 107 GB |
| H200 141 GB | 1000+ | all 1K | 138 GB |
| B200 180 GB | 1000+ | all 1K | 176 GB |
| Requests at once | 1K tokens each |
|---|---|
| 1 | 466 MB |
| 5 | 478 MB |
| 8 | 487 MB |
| 16 | 510 MB |
| 32 | 556 MB |
| 64 | 648 MB |
One card, with vLLM's small-card settings.
From the model card
We would like to thank Axiomic Labs for allowing us to use their TrainWork framework to train this model.
ForgePlex-M2-9M is a ~9.95M-parameter decoder-only language model from ForgeWorks. Trained on 30B tokens. Our second attempt at creating a <10m parameter model. We're proud of this product, while M1 was a promising start, M2 shows what we can do.
What was the motivation behind M2?
"That is an excellent question. To be completely honest. George Mallory was asked why he wanted to climb Everest. He said, “Because it’s there.” I just see M2 as a mountain to climb"
| Metric | Value |
|---|---|
| Unique parameters | 9,949,698 |
| Intelligence Index | 9.14 |
| HellaSwag | 28.02% |
| ARC (combined) | 30.00% |
| PIQA | 57.18% |
| ArithMark-3 | 34.60% |
| License | Apache-2.0 |
Trained on an all new dataset
| Data | Percentage |
|---|---|
| Finephrase | 50% |
| DCLM - Baseline | 30% |
| Proprietary Axiomic Labs dataset releasing soon | 10% |
| Finemath | 10% |
GQA + NeoX-style RoPE + RMSNorm + SwiGLU, with Qwen3.5-style attention output gates and Axiomic Labs TX4 style refresh gates on inject layers [5, 10] (kernel 9). XSA is off. Weights keep training key layout (no Llama remapping).
| Component | Details |
|---|---|
| Position encoding | RoPE (theta=5,000, NeoX even/odd) |
| Normalization | RMSNorm (eps=1e-6) |
| Feed-forward | SwiGLU (intermediate 707) |
| Attention | GQA — 8Q / 2KV, head_dim=32 + attn output gate |
| Refresh | Layers 5, 10, kernel 9 |
| Bias | None |
| Embedding | Weight tying |
| Depth × width | 11 layers × 256 hidden |
| Context | 1024 tokens |
| Vocab | 4,096 custom BPE |
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_path = r"C:\slm\ForgePlex-M2-9M"
tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_path,
trust_remote_code=True,
torch_dtype=torch.float32,
device_map="auto",
)
prompt = "Once upon a time"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.inference_mode():
out = model.generate(**inputs, max_new_tokens=80, do_sample=False)
print(tokenizer.decode(out[0], skip_special_tokens=True))
Or run python usage.py from this folder.
Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.