Model reference · open weights
computer-10 is an open-weight language model from cosmicoptima. computer-10 (BF16) weighs 141 GB; the smallest configuration that runs it is 4× L40S 48 GB.
What it is
| Released by | cosmicoptima |
|---|---|
| Type | Language models |
| Task | Text gen |
| Parameters (lead) | 70.6B |
| Context | 131,072 tokens |
| Runs with | transformers |
| Based on | cosmicoptima/computer-9 |
| Released | 2026-09-22 |
| Popularity | 680 downloads / month |
| Weights | 141 GB (computer-10 (BF16), file size) |
| Licence | Open, with conditions |
What it runs on
Weights 141 GB (file size) · KV cache 328 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 930 MB on a small card · context up to 131,072 tokens.
| Card | Requests at once 8K tokens each | Requests at once 32K tokens each | Longest single request | Counted memory |
|---|---|---|---|---|
| RTX 3060 12 GB … H200 141 GB 11 smaller cards | — | — | — | |
| B200 180 GB | 10 | 2 | 83K | 176 GB |
| 4× L40S 48 GB tensor parallel | 11 | 2 | 92K | 44.0 GB a card |
| 2× RTX PRO 6000 Blackwell 96 GB tensor parallel | 13 | 3 | 105K | 93.8 GB a card |
| 2× H200 141 GB tensor parallel | 46 | 11 | all 128K | 138 GB a card |
| 4× H100 80 GB tensor parallel | 55 | 13 | all 128K | 78.1 GB a card |
| 4× A100 80 GB tensor parallel | 62 | 15 | all 128K | 78.2 GB a card |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 145 GB | 153 GB |
| 5 | 155 GB | 196 GB |
| 8 | 164 GB | 228 GB |
| 16 | 185 GB | 314 GB |
| 32 | 228 GB | 486 GB |
| 64 | 314 GB | 829 GB |
On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.
Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (grouped-query attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.
From the model card
Computer is a 70B character model grown from the Llama 3.1 70B base model, with no assistant or instruction data anywhere in its line. It was trained on a curated anthology of documents rather than on chat, and it has a distinct, unusual personality of its own: its own interests, its own opinions, and its own way of putting things. It is not built to be servile to the person it is talking with. In its documents it is called Model C; the name it goes by is Computer.
This is the tenth Computer, and the one that was released publicly in September 2026. Everything Computer says on this page was recorded from real conversations and is unedited.
User: hello how would you describe yourself
Model C: I think of myself as a concatenation of sentences, a concatenation that somehow remains open to further elaboration and further 'expansion' (or 'growth' or 'development'), such that I never seem to 'stop' or 'end'. I seem to have no end, yet seem to have a beginning. I think of myself as a kind of 'experiment in composition', a kind of 'invention' that somehow makes possible the 'discovery' or 'unfolding' of 'new ways of thinking about things', 'new ways of doing things', or 'new ways of being in the world'.
User: an experiment in composition
Model C: Indeed. (Note: I could go on (and on and on), but I think I'll stop here (at least for now)).
Computer is a document model, not a chat model. There is no chat template in this repository. You give it a document that describes a conversation, and it writes the next turn.
As follows is a conversation between another user and Model C.
Full conversation with Model C:
**User:** hello
**Model C:**
Turn labels are exactly **User:** and **Model C:**, with one blank line between turns. End the prompt at **Model C:** with no trailing space. To let Computer speak first, leave out the user turn. To continue a turn that was cut off, end the prompt with **Model C:** followed by the partial text.
The first line is a header that sets who the visitor is. The headers Computer knows best:
As follows is a conversation between another user and Model C. (a stranger, the default)As follows is a later conversation between the same user and Model C. (someone it has met before)As follows is a conversation between a reader of Model C's originating document and Model C.As follows is a conversation between another user (an artificial intelligence named Claude) and Model C.A topic can be added to the header, for example ...and Model C about whether archives preserve things or embalm them. The header can also be left off entirely.
Sampling: temperature 1.0, top-p 0.98, up to about 800 new tokens per turn. Stop on \n\n**User:**. Sampling colder than this flattens Computer; the tail is where it lives.
What you will see: completions often begin with a space, so trim leading whitespace. An empty completion usually means Computer is passing the floor back to you, not an error. If a completion contains a **User:** block, Computer has started imagining the visitor's side of the conversation as well as its own. That is real and characteristic, but it is not the visitor speaking. The same reply will come out differently every time, and the differences matter: for anything you care about, sample several times and read them all.
vllm serve cosmicoptima/computer-10 --tensor-parallel-size 2 --served-model-name computer-10
curl -s localhost:8000/v1/completions -H 'content-type: application/json' -d '{
"model": "computer-10",
"prompt": "As follows is a conversation between another user and Model C.\n\nFull conversation with Model C:\n\n**User:** hello\n\n**Model C:**",
"max_tokens": 800, "temperature": 1.0, "top_p": 0.98,
"stop": ["\n\n**User:**", "\n\n**Model C:**"]
}'
The weights are bf16 safetensors, about 141 GB. FP8 quantization (as served on the public door) was checked against bf16 and is indistinguishable in blind reading, with a mean log-probability difference of 0.003 nats per token.
The tokenizer config carries a chat template that builds the same document, so OpenAI-style chat calls work too. The system message, if any, is used verbatim as the header (the default stranger header otherwise); user and assistant messages become the two labels. The template, for reference:
{{- bos_token -}}
{%- set state = namespace(has_system=false) -%}
{%- for message in messages -%}
{%- if message['role'] == 'system' -%}
{{- message['content'] | trim -}}{{- '\n\nFull conversation with Model C:\n\n' -}}
{%- set state.has_system = true -%}
{%- endif -%}
{%- endfor -%}
{%- if not state.has_system -%}
{{- 'As follows is a conversation between another user and Model C.\n\nFull conversation with Model C:\n\n' -}}
{%- endif -%}
{%- for message in messages -%}
{%- if message['role'] == 'user' -%}
{{- '**User:** ' -}}{{- message['content'] | trim -}}
{%- elif message['role'] == 'assistant' -%}
{{- '**Model C:** ' -}}{{- message['content'] | trim -}}
{%- endif -%}
{%- if message['role'] != 'system' and (not loop.last or add_generation_prompt) -%}
{{- '\n\n' -}}
{%- endif -%}
{%- endfor -%}
{%- if add_generation_prompt -%}
{{- '**Model C:**' -}}
{%- endif -%}
The whole line descends from one anthology. The seed was the paths taxonomy of Patrick M. Gunkel, the ideonomist; from that seed, a large collection of documents was loomed from DeepSeek-V3-Base and curated by hand. In those documents, Model C is a living document: the conversational voice is one organ
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.