Model reference · open weights
RWKV7-G1j-20260831 is an open-weight language model from RWKV. RWKV7-G1j-13.3B-20260831 (BF16) weighs 26.5 GB; the smallest configuration that runs it is RTX 5090 32 GB.
What it is
| Released by | RWKV |
|---|---|
| Type | Language models |
| Task | Text gen |
| Parameters (lead) | 13.3B |
| Runs with | transformers |
| Released | 2026-09-02 |
| Popularity | 1k downloads / month |
| Weights | 26.5 GB (RWKV7-G1j-13.3B-20260831 (BF16), file size) |
| Licence | Open weights |
What it runs on
Weights 26.5 GB (file size) · runtime overhead from 698 MB on a small card.
How much memory each request adds is not estimated yet for this architecture — only the weights are. They need the cards below at the least, plus room for the context.
| Card | The weights alone |
|---|---|
| RTX 3060 12 GB … RTX 4090 24 GB 4 smaller cards | does not fit |
| RTX 5090 32 GB | tight |
| L40S 48 GB | fits |
| A100 80 GB | fits |
| H100 80 GB | fits |
| RTX PRO 6000 Blackwell 96 GB | fits |
| DGX Spark (GB10) 128 GB unified | fits |
| H200 141 GB | fits |
| B200 180 GB | fits |
From the model card
This is an official BlinkDL release of RWKV-7 Goose in Hugging Face Transformers format. RWKV-7 is an attention-free recurrent architecture with a constant-size recurrent state and constant inference work per generated token. Training remains parallelizable.
This checkpoint is a base model pretrained with web, code, synthetic, instruction, chat, and reasoning data. It is suitable for evaluation, post-training, and fine-tuning; the included chat template is a prompt interface, not a claim that the checkpoint is a safety-aligned assistant.
The Transformers integration, conversion, release packaging, linear-time RWKV tokenizer, and optional TileLang inference implementation are distributed with this release.
tokenizer.json generated from the canonical RWKV World byte vocabulary.chat_template.jinja supports system, multi-turn, thinking, and
strict model-generated tool-call prompts.inference/ bundle
provides PyTorch fallback and TileLang acceleration without changing the
standard model root.| Field | Value |
|---|---|
| Repository | RWKV/RWKV7-G1j-13.3B-20260831 |
| Architecture class | Rwkv7ForCausalLM |
| Public size label | 13.3B |
| Source parameters | 13,270,298,624 |
| Serialized parameters | 13,270,298,624 |
| Synthesized compatibility tensors | 0 |
| Layers | 61 |
| Hidden / FFN size | 4096 / 16384 |
| Heads / head size | 64 / 64 |
| Vocabulary | 65536 |
| Training context | 16384 tokens |
| Weight dtype | bfloat16 |
| Numerical conversion | source dtype preserved |
| Metadata profile | g1j |
| Metadata provenance | locked-profile |
| Source checkpoint | BlinkDL/rwkv7-g1/rwkv7-g1j-13.3b-20260831-ctx16384.pth |
| Source SHA-256 | 559371f5b9aef13189ae54b345ac096af4ad2b689996c05d89de687612b3ae65 |
The repository includes configuration_rwkv7.py, modeling_rwkv7.py, and the
exact linear-time tokenization_rwkv7.py. The model modules are adapted
from the Transformers RWKV-7 integration at commit
4ad9ed0.
Review those files and pin a model-repository revision in production. Passing
trust_remote_code=True selects this bundled implementation even when the local
Transformers installation also provides native RWKV-7 support.
import torch
from transformers import (
AutoModelForCausalLM,
AutoTokenizer,
)
model_id = "RWKV/RWKV7-G1j-13.3B-20260831"
tokenizer = AutoTokenizer.from_pretrained(
model_id,
trust_remote_code=True,
)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
dtype=torch.bfloat16,
)
The recurrent cache returned by the model can be passed back for incremental
decoding. Use an attention_mask for padded batches.
The model defaults to the chunk-parallel WKV path for efficient multi-token
prefill. To reproduce the portable token-order reference path, set
model.config.wkv_implementation = "eager" before the first forward pass.
Chunked execution changes floating-point operation order, so small numerical
differences from eager execution are expected.
import re
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
THINK_RE = re.compile(r"\A?\s*(.*?)\s*?", re.DOTALL)
def assistant_content(completion, thinking, *, close_incomplete=False):
prefix = "\n"
reply = prefix + completion
thinking_block = THINK_RE.match(reply)
if thinking:
if thinking_block is not None or not close_incomplete:
return reply.strip()
return f"{reply.rstrip()}\n".strip()
return "" if thinking_block is None else reply[thinking_block.end():].strip()
model_id = "RWKV/RWKV7-G1j-13.3B-20260831"
tokenizer = AutoTokenizer.from_pretrained(
model_id,
trust_remote_code=True,
)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
dtype=torch.bfloat16,
).to("cuda")
messages = [{"role": "user", "content": "Explain why RWKV uses constant state."}]
thinking = False
max_new_tokens = 256
inputs = tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
thinking=thinking,
return_dict=True,
return_tensors="pt",
).to(model.device)
output = model.generate(
**inputs,
max_new_tokens=max_new_tokens,
do_sample=True,
temperature=1.0,
top_p=0.5,
eos_token_id=0,
pad_token_id=0,
stop_strings=["\n\nUser:"],
tokenizer=tokenizer,
)
completion = tokenizer.decode(
output[0, inputs["input_ids"].shape[1]:],
skip_special_tokens=True,
)
completion = completion.split("\n\nUser:", 1)[0]
reached_token_limit = output.shape[1] - inputs["input_ids"].shape[1] >= max_new_tokens
print(
assistant_content(
completion,
thinking,
close_incomplete=reached_token_limit,
)
)
Set thinking=True for the RWKV thinking prefix. The intentional generation
prefixes are Assistant: followed by a newline and
Assistant: <think. Only the enabled thinking prefix intentionally leaves its opening
tag incomplete. The post-processing above reconstructs that prefix before removing an
empty thinking block or preserving an enabled one. If generation hits the token limit
inside thinking, it closes the displayed block be
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.