Model reference · open weights

openchat

LLMs openchat Text gen 1 build Its own licence terms 133 dl/mo

openchat is an open-weight language model from openchat. openchat (BF16) weighs 26.0 GB; the smallest configuration that runs it is 2× RTX 4060 Ti 16 GB.

What it is

Released byopenchat
TypeLanguage models
TaskText gen
Context2,048 tokens
Runs withtransformers
Released2023-06-22
Popularity133 downloads / month
Weights26.0 GB (openchat (BF16), file size)
LicenceIts own licence terms

What it runs on

Memory and cards for openchat (BF16)

Weights 26.0 GB (file size) · KV cache 819 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 659 MB on a small card · context up to 2,048 tokens.

CardRequests at once
2K, its whole window tokens each
Requests at once
32K tokens each
Longest single
request
Counted
memory
RTX 3060 12 GB … RTX 4090 24 GB
4 smaller cards
———
RTX 5090 32 GB2—all 2K31.0 GB
L40S 48 GB10—all 2K44.0 GB
A100 80 GB30—all 2K78.2 GB
H100 80 GB28—all 2K78.1 GB
RTX PRO 6000 Blackwell 96 GB38—all 2K93.8 GB
DGX Spark (GB10) 128 GB unified46—all 2K107 GB
H200 141 GB64—all 2K138 GB
B200 180 GB86—all 2K176 GB
2× RTX 4060 Ti 16 GB
tensor parallel
2—all 2K15.4 GB a card
2× RTX 4090 24 GB
tensor parallel
11—all 2K23.4 GB a card
2× RTX 3090 24 GB
tensor parallel
11—all 2K23.4 GB a card
2× RTX 5090 32 GB
tensor parallel
20—all 2K31.0 GB a card
Memory needed at each load
Requests at once2K, its whole window tokens each32K tokens each
128.4 GB—
535.1 GB—
840.1 GB—
1653.5 GB—
3280.4 GB—
64134 GB—

On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.

Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (multi-head attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.

From the model card

What openchat says about openchat

OpenChat is a series of open-source language models fine-tuned on a diverse and high-quality dataset of multi-round conversations. With only ~6K GPT-4 conversations filtered from the ~90K ShareGPT conversations, OpenChat is designed to achieve high performance with limited data.

Generic models:

  • OpenChat: based on LLaMA-13B (2048 context length)
    • 🚀 105.7% of ChatGPT score on Vicuna GPT-4 evaluation
    • 🔥 80.9% Win-rate on AlpacaEval
    • 🤗 Only used 6K data for finetuning!!!
  • OpenChat-8192: based on LLaMA-13B (extended to 8192 context length)
    • 106.6% of ChatGPT score on Vicuna GPT-4 evaluation
    • 79.5% Win-rate on AlpacaEval

Code models:

  • OpenCoderPlus: based on StarCoderPlus (native 8192 context length)
    • 102.5% of ChatGPT score on Vicuna GPT-4 evaluation
    • 78.7% Win-rate on AlpacaEval

Note: Please load the pretrained models using bfloat16

Read the full model card

Code and Inference Server

We provide the full source code, including an inference server compatible with the "ChatCompletions" API, in the OpenChat GitHub repository.

Web UI

OpenChat also includes a web UI for a better user experience. See the GitHub repository for instructions.

Conversation Template

The conversation template involves concatenating tokens.

Besides base model vocabulary, an end-of-turn token `` is added, with id eot_token_id.

# OpenChat
[bos_token_id] + tokenize("Human: ") + tokenize(user_question) + [eot_token_id] + tokenize("Assistant: ")
# OpenCoder
tokenize("User:") + tokenize(user_question) + [eot_token_id] + tokenize("Assistant:")

Hint: In BPE, tokenize(A) + tokenize(B) does not always equals to tokenize(A + B)

Following is the code for generating the conversation templates:

@dataclass
class ModelConfig:
    # Prompt
    system: Optional[str]

    role_prefix: dict
    ai_role: str
    eot_token: str
    bos_token: Optional[str] = None

    # Get template
    def generate_conversation_template(self, tokenize_fn, tokenize_special_fn, message_list):
        tokens = []
        masks = []

        # begin of sentence (bos)
        if self.bos_token:
            t = tokenize_special_fn(self.bos_token)
            tokens.append(t)
            masks.append(False)

        # System
        if self.system:
            t = tokenize_fn(self.system) + [tokenize_special_fn(self.eot_token)]
            tokens.extend(t)
            masks.extend([False] * len(t))

        # Messages
        for idx, message in enumerate(message_list):
            # Prefix
            t = tokenize_fn(self.role_prefix[message["from"]])
            tokens.extend(t)
            masks.extend([False] * len(t))

            # Message
            if "value" in message:
                t = tokenize_fn(message["value"]) + [tokenize_special_fn(self.eot_token)]
                tokens.extend(t)
                masks.extend([message["from"] == self.ai_role] * len(t))
            else:
                assert idx == len(message_list) - 1, "Empty message for completion must be on the last."

        return tokens, masks

MODEL_CONFIG_MAP = {
    # OpenChat / OpenChat-8192
    "openchat": ModelConfig(
        # Prompt
        system=None,

        role_prefix={
            "human": "Human: ",
            "gpt": "Assistant: "
        },
        ai_role="gpt",
        eot_token="",
        bos_token="",
    ),

    # OpenCoder / OpenCoderPlus
    "opencoder": ModelConfig(
        # Prompt
        system=None,

        role_prefix={
            "human": "User:",
            "gpt": "Assistant:"
        },
        ai_role="gpt",
        eot_token="",
        bos_token=None,
    )
}

License

Our weight license is subject to their corresponding base model. For example, OpenChat and OpenChat-8192 are the same as the model License of LLaMA for non-commercial use only, while OpenCoderPlus is under the License of StarCoder. Furthermore, we should follow the Privacy Practices of ShareGPT. The code released on GitHub is under Apache License 2.0.

Citation

@software{openllms23,
  title = {{OpenLLMs: Less is More for Open-source Models}},
  author = {Wang, Guan and Cheng, Sijie and Yu, Qiying and Liu, Changling},
  doi = {10.5281/zenodo.8105775},
  url = {https://github.com/imoneoi/openchat},
  version = {pre-release},
  year = {2023},
  month = {7},
}

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms