Model reference · open weights
openchat.2_super is an open-weight language model from openchat. openchat_v3.2_super (BF16) weighs 26.0 GB; the smallest configuration that runs it is 2× RTX 4060 Ti 16 GB.
What it is
| Released by | openchat |
|---|---|
| Type | Language models |
| Task | Text gen |
| Context | 4,096 tokens |
| Runs with | transformers |
| Released | 2023-09-04 |
| Popularity | 148 downloads / month |
| Weights | 26.0 GB (openchat_v3.2_super (BF16), file size) |
| Licence | Open, with conditions |
What it runs on
Weights 26.0 GB (file size) · KV cache 819 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 659 MB on a small card · context up to 4,096 tokens.
| Card | Requests at once 4K, its whole window tokens each | Requests at once 32K tokens each | Longest single request | Counted memory |
|---|---|---|---|---|
| RTX 3060 12 GB … RTX 4090 24 GB 4 smaller cards | — | — | — | |
| RTX 5090 32 GB | 1 | — | all 4K | 31.0 GB |
| L40S 48 GB | 5 | — | all 4K | 44.0 GB |
| A100 80 GB | 15 | — | all 4K | 78.2 GB |
| H100 80 GB | 14 | — | all 4K | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | 19 | — | all 4K | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | 23 | — | all 4K | 107 GB |
| H200 141 GB | 32 | — | all 4K | 138 GB |
| B200 180 GB | 43 | — | all 4K | 176 GB |
| 2× RTX 4060 Ti 16 GB tensor parallel | 1 | — | all 4K | 15.4 GB a card |
| 2× RTX 5090 32 GB tensor parallel | 10 | — | all 4K | 31.0 GB a card |
| 2× L40S 48 GB tensor parallel | 18 | — | all 4K | 44.0 GB a card |
| 4× RTX 4090 24 GB tensor parallel | 19 | — | all 4K | 23.4 GB a card |
| 4× RTX 3090 24 GB tensor parallel | 19 | — | all 4K | 23.4 GB a card |
| Requests at once | 4K, its whole window tokens each | 32K tokens each |
|---|---|---|
| 1 | 30.0 GB | — |
| 5 | 43.5 GB | — |
| 8 | 53.5 GB | — |
| 16 | 80.4 GB | — |
| 32 | 134 GB | — |
| 64 | 241 GB | — |
On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.
Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (multi-head attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.
From the model card
OpenChat is a collection of open-source language models, optimized and fine-tuned with a strategy inspired by offline reinforcement learning. We use approximately 80k ShareGPT conversations, a conditioning strategy, and weighted loss to deliver outstanding performance, despite our simple approach. Our ultimate goal is to develop a high-performance, commercially available, open-source large language model, and we are continuously making strides towards this vision.
🤖 Ranked #1 among all open-source models on AgentBench
🔥 Ranked #1 among 13B open-source models | 89.5% win-rate on AlpacaEval | 7.19 score on MT-bench
🕒 Exceptionally efficient padding-free fine-tuning, only requires 15 hours on 8xA100 80G
💲 FREE for commercial use under Llama 2 Community License
To use these models, we highly recommend installing the OpenChat package by following the installation guide and using the OpenChat OpenAI-compatible API server by running the serving command from the table below. The server is optimized for high-throughput deployment using vLLM and can run on a GPU with at least 48GB RAM or two consumer GPUs with tensor parallelism. To enable tensor parallelism, append --tensor-parallel-size 2 to the serving command.
When started, the server listens at localhost:18888 for requests and is compatible with the OpenAI ChatCompletion API specifications. See the example request below for reference. Additionally, you can access the OpenChat Web UI for a user-friendly experience.
To deploy the server as an online service, use --api-keys sk-KEY1 sk-KEY2 ... to specify allowed API keys and --disable-log-requests --disable-log-stats --log-file openchat.log for logging only to a file. We recommend using a HTTPS gateway in front of the server for security purposes.
curl http://localhost:18888/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "openchat_v3.2",
"messages": [{"role": "user", "content": "You are a large language model named OpenChat. Write a poem to describe yourself"}]
}'
| Model | Size | Context | Weights | Serving |
|---|---|---|---|---|
| OpenChat 3.2 SUPER | 13B | 4096 | Huggingface | python -m ochat.serving.openai_api_server --model-type openchat_v3.2 --model openchat/openchat_v3.2_super --engine-use-ray --worker-use-ray --max-num-batched-tokens 5120 |
For inference with Huggingface Transformers (slow and not recommended), follow the conversation template provided below:
# Single-turn V3.2 (SUPER)
tokenize("GPT4 User: HelloGPT4 Assistant:")
# Result: [1, 402, 7982, 29946, 4911, 29901, 15043, 32000, 402, 7982, 29946, 4007, 22137, 29901]
# Multi-turn V3.2 (SUPER)
tokenize("GPT4 User: HelloGPT4 Assistant: HiGPT4 User: How are you today?GPT4 Assistant:")
# Result: [1, 402, 7982, 29946, 4911, 29901, 15043, 32000, 402, 7982, 29946, 4007, 22137, 29901, 6324, 32000, 402, 7982, 29946, 4911, 29901, 1128, 526, 366, 9826, 29973, 32000, 402, 7982, 29946, 4007, 22137, 29901]
We have evaluated our models using the two most popular evaluation benchmarks **, including AlpacaEval and MT-bench. Here we list the top models with our released versions, sorted by model size in descending order. The full version can be found on the MT-bench and AlpacaEval leaderboards.
To ensure consistency, we used the same routine as ChatGPT / GPT-4 to run these benchmarks. We started the OpenAI API-compatible server and set the openai.api_base to http://localhost:18888/v1 in the benchmark program.
| Model | Size | Context | Dataset Size | 💲Free | AlpacaEval (win rate %) | MT-bench (win rate adjusted %) | MT-bench (score) |
|---|---|---|---|---|---|---|---|
| v.s. text-davinci-003 | v.s. ChatGPT | ||||||
| GPT-4 | 1.8T* | 8K | ❌ | 95.3 | 82.5 | 8.99 | |
| ChatGPT | 175B* | 4K | ❌ | 89.4 | 50.0 | 7.94 | |
| Llama-2-70B-Chat | 70B | 4K | 2.9M | ✅ | 92.7 | 60.0 | 6.86 |
| OpenChat 3.2 SUPER | 13B | 4K | 80K | ✅ | 89.5 | 57.5 | 7.19 |
| Llama-2-13B-Chat | 13B | 4K | 2.9M |
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.