Model reference · open weights

ForgePlex-M2

NEW · this week LLMs ForgeWorks Text gen 1 build Open weights 886 dl/mo

ForgePlex-M2 is an open-weight language model from ForgeWorks. ForgePlex-M2-9M (FP32) weighs 20 MB; the smallest configuration that runs it is RTX 3060 12 GB.

  • ForgePlex-M2-9M is a 9.95M-parameter decoder-only language model developed by ForgeWorks for text generation.
  • It supports English with a context length of 1024 tokens and is released under the Apache-2.0 license.
  • The model features an 11-layer architecture with 256 hidden dimensions and a custom 4,096-token vocabulary.

Summary of the ForgeWorks/ForgePlex-M2-9M model card, 2026-10-04

What it is

Released byForgeWorks
Released2026-09-29
Parameters10M
VRAM20 MB for the weights

What it runs on

Memory and cards for ForgePlex-M2-9M (FP32)

20 MBweights, file size
3 MBcache per 1K tokens
444 MBruntime overhead, at least
1,024 tokenscontext max
CardRequests at onceContext maxMemory
1K each
RTX 3060 12 GB1000+all 1K11.6 GB
RTX 4060 Ti 16 GB1000+all 1K15.4 GB
RTX 3090 24 GB1000+all 1K23.4 GB
RTX 4090 24 GB1000+all 1K23.4 GB
RTX 5090 32 GB1000+all 1K31.0 GB
L40S 48 GB1000+all 1K44.0 GB
A100 80 GB1000+all 1K78.2 GB
H100 80 GB1000+all 1K78.1 GB
RTX PRO 6000 Blackwell 96 GB1000+all 1K93.8 GB
DGX Spark (GB10) 128 GB unified1000+all 1K107 GB
H200 141 GB1000+all 1K138 GB
B200 180 GB1000+all 1K176 GB
Memory needed at each load
Requests at once1K tokens each
1466 MB
5478 MB
8487 MB
16510 MB
32556 MB
64648 MB

One card, with vLLM's small-card settings.

From the model card

What ForgeWorks says about ForgePlex-M2

Read the model card

We would like to thank Axiomic Labs for allowing us to use their TrainWork framework to train this model.

ForgePlex-M2-9M

ForgePlex-M2-9M is a ~9.95M-parameter decoder-only language model from ForgeWorks. Trained on 30B tokens. Our second attempt at creating a <10m parameter model. We're proud of this product, while M1 was a promising start, M2 shows what we can do.

Q&A

What was the motivation behind M2?

"That is an excellent question. To be completely honest. George Mallory was asked why he wanted to climb Everest. He said, “Because it’s there.” I just see M2 as a mountain to climb"

MetricValue
Unique parameters9,949,698
Intelligence Index9.14
HellaSwag28.02%
ARC (combined)30.00%
PIQA57.18%
ArithMark-334.60%
LicenseApache-2.0

Training

Trained on an all new dataset

DataPercentage
Finephrase50%
DCLM - Baseline30%
Proprietary Axiomic Labs dataset releasing soon10%
Finemath10%

Architecture

GQA + NeoX-style RoPE + RMSNorm + SwiGLU, with Qwen3.5-style attention output gates and Axiomic Labs TX4 style refresh gates on inject layers [5, 10] (kernel 9). XSA is off. Weights keep training key layout (no Llama remapping).

ComponentDetails
Position encodingRoPE (theta=5,000, NeoX even/odd)
NormalizationRMSNorm (eps=1e-6)
Feed-forwardSwiGLU (intermediate 707)
AttentionGQA — 8Q / 2KV, head_dim=32 + attn output gate
RefreshLayers 5, 10, kernel 9
BiasNone
EmbeddingWeight tying
Depth × width11 layers × 256 hidden
Context1024 tokens
Vocab4,096 custom BPE

Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_path = r"C:\slm\ForgePlex-M2-9M"
tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_path,
    trust_remote_code=True,
    torch_dtype=torch.float32,
    device_map="auto",
)

prompt = "Once upon a time"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.inference_mode():
    out = model.generate(**inputs, max_new_tokens=80, do_sample=False)
print(tokenizer.decode(out[0], skip_special_tokens=True))

Or run python usage.py from this folder.

Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms