Model reference · open weights
openchat is an open-weight language model from openchat. openchat (BF16) weighs 26.0 GB; the smallest configuration that runs it is 2× RTX 4060 Ti 16 GB.
What it is
| Released by | openchat |
|---|---|
| Type | Language models |
| Task | Text gen |
| Context | 2,048 tokens |
| Runs with | transformers |
| Released | 2023-06-22 |
| Popularity | 133 downloads / month |
| Weights | 26.0 GB (openchat (BF16), file size) |
| Licence | Its own licence terms |
What it runs on
Weights 26.0 GB (file size) · KV cache 819 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 659 MB on a small card · context up to 2,048 tokens.
| Card | Requests at once 2K, its whole window tokens each | Requests at once 32K tokens each | Longest single request | Counted memory |
|---|---|---|---|---|
| RTX 3060 12 GB … RTX 4090 24 GB 4 smaller cards | — | — | — | |
| RTX 5090 32 GB | 2 | — | all 2K | 31.0 GB |
| L40S 48 GB | 10 | — | all 2K | 44.0 GB |
| A100 80 GB | 30 | — | all 2K | 78.2 GB |
| H100 80 GB | 28 | — | all 2K | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | 38 | — | all 2K | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | 46 | — | all 2K | 107 GB |
| H200 141 GB | 64 | — | all 2K | 138 GB |
| B200 180 GB | 86 | — | all 2K | 176 GB |
| 2× RTX 4060 Ti 16 GB tensor parallel | 2 | — | all 2K | 15.4 GB a card |
| 2× RTX 4090 24 GB tensor parallel | 11 | — | all 2K | 23.4 GB a card |
| 2× RTX 3090 24 GB tensor parallel | 11 | — | all 2K | 23.4 GB a card |
| 2× RTX 5090 32 GB tensor parallel | 20 | — | all 2K | 31.0 GB a card |
| Requests at once | 2K, its whole window tokens each | 32K tokens each |
|---|---|---|
| 1 | 28.4 GB | — |
| 5 | 35.1 GB | — |
| 8 | 40.1 GB | — |
| 16 | 53.5 GB | — |
| 32 | 80.4 GB | — |
| 64 | 134 GB | — |
On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.
Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (multi-head attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.
From the model card
OpenChat is a series of open-source language models fine-tuned on a diverse and high-quality dataset of multi-round conversations. With only ~6K GPT-4 conversations filtered from the ~90K ShareGPT conversations, OpenChat is designed to achieve high performance with limited data.
Generic models:
Code models:
Note: Please load the pretrained models using bfloat16
We provide the full source code, including an inference server compatible with the "ChatCompletions" API, in the OpenChat GitHub repository.
OpenChat also includes a web UI for a better user experience. See the GitHub repository for instructions.
The conversation template involves concatenating tokens.
Besides base model vocabulary, an end-of-turn token `` is added, with id eot_token_id.
# OpenChat
[bos_token_id] + tokenize("Human: ") + tokenize(user_question) + [eot_token_id] + tokenize("Assistant: ")
# OpenCoder
tokenize("User:") + tokenize(user_question) + [eot_token_id] + tokenize("Assistant:")
Hint: In BPE, tokenize(A) + tokenize(B) does not always equals to tokenize(A + B)
Following is the code for generating the conversation templates:
@dataclass
class ModelConfig:
# Prompt
system: Optional[str]
role_prefix: dict
ai_role: str
eot_token: str
bos_token: Optional[str] = None
# Get template
def generate_conversation_template(self, tokenize_fn, tokenize_special_fn, message_list):
tokens = []
masks = []
# begin of sentence (bos)
if self.bos_token:
t = tokenize_special_fn(self.bos_token)
tokens.append(t)
masks.append(False)
# System
if self.system:
t = tokenize_fn(self.system) + [tokenize_special_fn(self.eot_token)]
tokens.extend(t)
masks.extend([False] * len(t))
# Messages
for idx, message in enumerate(message_list):
# Prefix
t = tokenize_fn(self.role_prefix[message["from"]])
tokens.extend(t)
masks.extend([False] * len(t))
# Message
if "value" in message:
t = tokenize_fn(message["value"]) + [tokenize_special_fn(self.eot_token)]
tokens.extend(t)
masks.extend([message["from"] == self.ai_role] * len(t))
else:
assert idx == len(message_list) - 1, "Empty message for completion must be on the last."
return tokens, masks
MODEL_CONFIG_MAP = {
# OpenChat / OpenChat-8192
"openchat": ModelConfig(
# Prompt
system=None,
role_prefix={
"human": "Human: ",
"gpt": "Assistant: "
},
ai_role="gpt",
eot_token="",
bos_token="",
),
# OpenCoder / OpenCoderPlus
"opencoder": ModelConfig(
# Prompt
system=None,
role_prefix={
"human": "User:",
"gpt": "Assistant:"
},
ai_role="gpt",
eot_token="",
bos_token=None,
)
}
Our weight license is subject to their corresponding base model. For example, OpenChat and OpenChat-8192 are the same as the model License of LLaMA for non-commercial use only, while OpenCoderPlus is under the License of StarCoder. Furthermore, we should follow the Privacy Practices of ShareGPT. The code released on GitHub is under Apache License 2.0.
@software{openllms23,
title = {{OpenLLMs: Less is More for Open-source Models}},
author = {Wang, Guan and Cheng, Sijie and Yu, Qiying and Liu, Changling},
doi = {10.5281/zenodo.8105775},
url = {https://github.com/imoneoi/openchat},
version = {pre-release},
year = {2023},
month = {7},
}
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.