Model reference · open weights

circuit-sparsity

LLMs openai Text gen 1 build Open weights 465 dl/mo

circuit-sparsity is an open-weight language model from OpenAI. circuit-sparsity (FP32) weighs 838 MB; the smallest configuration that runs it is RTX 3060 12 GB.

What it is

Released byOpenAI
TypeLanguage models
TaskText gen
Parameters (lead)419M
Runs withtransformers
Released2025-12-11
Popularity465 downloads / month
Weights838 MB (circuit-sparsity (FP32), file size)
LicenceOpen weights

What it runs on

Memory and cards for circuit-sparsity (FP32)

Weights 838 MB (file size) · KV cache 66 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 532 MB on a small card · context up to 1,024 tokens.

CardRequests at once
1K, its whole window tokens each
Requests at once
32K tokens each
Longest single
request
Counted
memory
RTX 3060 12 GB152—all 1K11.6 GB
RTX 4060 Ti 16 GB209—all 1K15.4 GB
RTX 3090 24 GB328—all 1K23.4 GB
RTX 4090 24 GB327—all 1K23.4 GB
RTX 5090 32 GB441—all 1K31.0 GB
L40S 48 GB634—all 1K44.0 GB
A100 80 GB1000+—all 1K78.2 GB
H100 80 GB1000+—all 1K78.1 GB
RTX PRO 6000 Blackwell 96 GB1000+—all 1K93.8 GB
DGX Spark (GB10) 128 GB unified1000+—all 1K107 GB
H200 141 GB1000+—all 1K138 GB
B200 180 GB1000+—all 1K176 GB
Memory needed at each load
Requests at once1K, its whole window tokens each32K tokens each
11.4 GB—
51.7 GB—
81.9 GB—
162.4 GB—
323.5 GB—
645.7 GB—

On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.

Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (multi-head attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). Assumes vLLM 0.10 or later.

From the model card

What OpenAI says about circuit-sparsity

Sparse Model from Gao et al. 2025

Weights for a sparse model from Gao et al. 2025, used for the qualitative results from the paper (related to bracket counting and variable binding). All weights for the other models used in the paper, as well as lightweight inference code, are present in https://github.com/openai/circuit_sparsity. In the context of that repo, this model is csp_yolo2.

This is a runnable standalone huggingface implementation for one of the models. It includes code to load the locally converted HF model + tokenizer and run a tiny generation.

Read the full model card
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

if __name__ == "__main__":
    PROMPT = "def square_sum(xs):\n    return sum(x * x for x in xs)\n\nsquare_sum([1, 2, 3])\n"
    tok = AutoTokenizer.from_pretrained("openai/circuit-sparsity", trust_remote_code=True)
    model = AutoModelForCausalLM.from_pretrained(
        "openai/circuit-sparsity",
        trust_remote_code=True,
        torch_dtype="auto",
    )
    model.to("cuda" if torch.cuda.is_available() else "cpu")
    inputs = tok(PROMPT, return_tensors="pt", add_special_tokens=False)["input_ids"].to(
        model.device
    )

    with torch.no_grad():
        out = model.generate(
            inputs,
            max_new_tokens=64,
            do_sample=True,
            temperature=0.8,
            top_p=0.95,
            return_dict_in_generate=False,
        )

    print("=== Prompt ===")
    print(PROMPT)
    print("\n=== Generation ===")
    print(tok.decode(out[0], skip_special_tokens=True))

License

This project is licensed under the Apache License 2.0.

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms