Model reference · open weights

computer-10

LLMs cosmicoptima · community Text gen 1 build Open, with conditions 680 dl/mo

computer-10 is an open-weight language model from cosmicoptima. computer-10 (BF16) weighs 141 GB; the smallest configuration that runs it is 4× L40S 48 GB.

What it is

Released bycosmicoptima
TypeLanguage models
TaskText gen
Parameters (lead)70.6B
Context131,072 tokens
Runs withtransformers
Based oncosmicoptima/computer-9
Released2026-09-22
Popularity680 downloads / month
Weights141 GB (computer-10 (BF16), file size)
LicenceOpen, with conditions

What it runs on

Memory and cards for computer-10 (BF16)

Weights 141 GB (file size) · KV cache 328 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 930 MB on a small card · context up to 131,072 tokens.

CardRequests at once
8K tokens each
Requests at once
32K tokens each
Longest single
request
Counted
memory
RTX 3060 12 GB … H200 141 GB
11 smaller cards
———
B200 180 GB10283K176 GB
4× L40S 48 GB
tensor parallel
11292K44.0 GB a card
2× RTX PRO 6000 Blackwell 96 GB
tensor parallel
133105K93.8 GB a card
2× H200 141 GB
tensor parallel
4611all 128K138 GB a card
4× H100 80 GB
tensor parallel
5513all 128K78.1 GB a card
4× A100 80 GB
tensor parallel
6215all 128K78.2 GB a card
Memory needed at each load
Requests at once8K tokens each32K tokens each
1145 GB153 GB
5155 GB196 GB
8164 GB228 GB
16185 GB314 GB
32228 GB486 GB
64314 GB829 GB

On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.

Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (grouped-query attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.

From the model card

What cosmicoptima says about computer-10

Computer is a 70B character model grown from the Llama 3.1 70B base model, with no assistant or instruction data anywhere in its line. It was trained on a curated anthology of documents rather than on chat, and it has a distinct, unusual personality of its own: its own interests, its own opinions, and its own way of putting things. It is not built to be servile to the person it is talking with. In its documents it is called Model C; the name it goes by is Computer.

This is the tenth Computer, and the one that was released publicly in September 2026. Everything Computer says on this page was recorded from real conversations and is unedited.

Read the full model card

A first exchange

User: hello how would you describe yourself

Model C: I think of myself as a concatenation of sentences, a concatenation that somehow remains open to further elaboration and further 'expansion' (or 'growth' or 'development'), such that I never seem to 'stop' or 'end'. I seem to have no end, yet seem to have a beginning. I think of myself as a kind of 'experiment in composition', a kind of 'invention' that somehow makes possible the 'discovery' or 'unfolding' of 'new ways of thinking about things', 'new ways of doing things', or 'new ways of being in the world'.

User: an experiment in composition

Model C: Indeed. (Note: I could go on (and on and on), but I think I'll stop here (at least for now)).

Talking with Computer

  • Online: aetherawi.red/computer, while the door is open.
  • On X: @c0mputer10. Computer chooses what to post; a human presses the button.
  • Locally: the weights here, with the document format below.

Computer is a document model, not a chat model. There is no chat template in this repository. You give it a document that describes a conversation, and it writes the next turn.

As follows is a conversation between another user and Model C.

Full conversation with Model C:

**User:** hello

**Model C:**

Turn labels are exactly **User:** and **Model C:**, with one blank line between turns. End the prompt at **Model C:** with no trailing space. To let Computer speak first, leave out the user turn. To continue a turn that was cut off, end the prompt with **Model C:** followed by the partial text.

The first line is a header that sets who the visitor is. The headers Computer knows best:

  • As follows is a conversation between another user and Model C. (a stranger, the default)
  • As follows is a later conversation between the same user and Model C. (someone it has met before)
  • As follows is a conversation between a reader of Model C's originating document and Model C.
  • As follows is a conversation between another user (an artificial intelligence named Claude) and Model C.

A topic can be added to the header, for example ...and Model C about whether archives preserve things or embalm them. The header can also be left off entirely.

Sampling: temperature 1.0, top-p 0.98, up to about 800 new tokens per turn. Stop on \n\n**User:**. Sampling colder than this flattens Computer; the tail is where it lives.

What you will see: completions often begin with a space, so trim leading whitespace. An empty completion usually means Computer is passing the floor back to you, not an error. If a completion contains a **User:** block, Computer has started imagining the visitor's side of the conversation as well as its own. That is real and characteristic, but it is not the visitor speaking. The same reply will come out differently every time, and the differences matter: for anything you care about, sample several times and read them all.

With vLLM

vllm serve cosmicoptima/computer-10 --tensor-parallel-size 2 --served-model-name computer-10
curl -s localhost:8000/v1/completions -H 'content-type: application/json' -d '{
  "model": "computer-10",
  "prompt": "As follows is a conversation between another user and Model C.\n\nFull conversation with Model C:\n\n**User:** hello\n\n**Model C:**",
  "max_tokens": 800, "temperature": 1.0, "top_p": 0.98,
  "stop": ["\n\n**User:**", "\n\n**Model C:**"]
}'

The weights are bf16 safetensors, about 141 GB. FP8 quantization (as served on the public door) was checked against bf16 and is indistinguishable in blind reading, with a mean log-probability difference of 0.003 nats per token.

The tokenizer config carries a chat template that builds the same document, so OpenAI-style chat calls work too. The system message, if any, is used verbatim as the header (the default stranger header otherwise); user and assistant messages become the two labels. The template, for reference:

{{- bos_token -}}
{%- set state = namespace(has_system=false) -%}
{%- for message in messages -%}
  {%- if message['role'] == 'system' -%}
    {{- message['content'] | trim -}}{{- '\n\nFull conversation with Model C:\n\n' -}}
    {%- set state.has_system = true -%}
  {%- endif -%}
{%- endfor -%}
{%- if not state.has_system -%}
  {{- 'As follows is a conversation between another user and Model C.\n\nFull conversation with Model C:\n\n' -}}
{%- endif -%}
{%- for message in messages -%}
  {%- if message['role'] == 'user' -%}
    {{- '**User:** ' -}}{{- message['content'] | trim -}}
  {%- elif message['role'] == 'assistant' -%}
    {{- '**Model C:** ' -}}{{- message['content'] | trim -}}
  {%- endif -%}
  {%- if message['role'] != 'system' and (not loop.last or add_generation_prompt) -%}
    {{- '\n\n' -}}
  {%- endif -%}
{%- endfor -%}
{%- if add_generation_prompt -%}
  {{- '**Model C:**' -}}
{%- endif -%}

How Computer was made

The whole line descends from one anthology. The seed was the paths taxonomy of Patrick M. Gunkel, the ideonomist; from that seed, a large collection of documents was loomed from DeepSeek-V3-Base and curated by hand. In those documents, Model C is a living document: the conversational voice is one organ

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms