Model reference · open weights
open.2_super is an open-weight language model from openchat. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.
Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.
What it is
| Maker | openchat |
|---|---|
| Type | Language models |
| Task | Text gen |
| Context | 4k tokens |
| Runs with | transformers |
| Released | 2023-09-04 |
| Popularity | 148 downloads / month |
| Licence | Open, with conditions |
About
OpenChat is a collection of open-source language models, optimized and fine-tuned with a strategy inspired by offline reinforcement learning. We use approximately 80k ShareGPT conversations, a conditioning strategy, and weighted loss to deliver outstanding performance, despite our simple approach. Our ultimate goal is to develop a high-performance, commercially available, open-source large language model, and we are continuously making strides towards this vision.
🤖 Ranked #1 among all open-source models on AgentBench
🔥 Ranked #1 among 13B open-source models | 89.5% win-rate on AlpacaEval | 7.19 score on MT-bench
🕒 Exceptionally efficient padding-free fine-tuning, only requires 15 hours on 8xA100 80G
💲 FREE for commercial use under Llama 2 Community License
To use these models, we highly recommend installing the OpenChat package by following the installation guide and using the OpenChat OpenAI-compatible API server by running the serving command from the table below. The server is optimized for high-throughput deployment using vLLM and can run on a GPU with at least 48GB RAM or two consumer GPUs with tensor parallelism. To enable tensor parallelism, append --tensor-parallel-size 2 to the serving command.
When started, the server listens at localhost:18888 for requests and is compatible with the OpenAI ChatCompletion API specifications. See the example request below for reference. Additionally, you can access the OpenChat Web UI for a user-friendly experience.
To deploy the server as an online service, use --api-keys sk-KEY1 sk-KEY2 ... to specify allowed API keys and --disable-log-requests --disable-log-stats --log-file openchat.log for logging only to a file. We recommend using a HTTPS gateway in front of the server for security purposes.
curl http://localhost:18888/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "openchat_v3.2",
"messages": [{"role": "user", "content": "You are a large language model named OpenChat. Write a poem to describe yourself"}]
}'
| Model | Size | Context | Weights | Serving |
|---|---|---|---|---|
| OpenChat 3.2 SUPER | 13B | 4096 | Huggingface | python -m ochat.serving.openai_api_server --model-type openchat_v3.2 --model openchat/openchat_v3.2_super --engine-use-ray --worker-use-ray --max-num-batched-tokens 5120 |
For inference with Huggingface Transformers (slow and not recommended), follow the conversation template provided below:
# Single-turn V3.2 (SUPER)
tokenize("GPT4 User: HelloGPT4 Assistant:")
# Result: [1, 402, 7982, 29946, 4911, 29901, 15043, 32000, 402, 7982, 29946, 4007, 22137, 29901]
# Multi-turn V3.2 (SUPER)
tokenize("GPT4 User: HelloGPT4 Assistant: HiGPT4 User: How are you today?GPT4 Assistant:")
# Result: [1, 402, 7982, 29946, 4911, 29901, 15043, 32000, 402, 7982, 29946, 4007, 22137, 29901, 6324, 32000, 402, 7982, 29946, 4911, 29901, 1128, 526, 366, 9826, 29973, 32000, 402, 7982, 29946, 4007, 22137, 29901]
We have evaluated our models using the two most popular evaluation benchmarks **, including AlpacaEval and MT-bench. Here we list the top models with our released versions, sorted by model size in descending order. The full version can be found on the MT-bench and AlpacaEval leaderboards.
To ensure consistency, we used the same routine as ChatGPT / GPT-4 to run these benchmarks. We started the OpenAI API-compatible server and set the openai.api_base to http://localhost:18888/v1 in the benchmark program.
| Model | Size | Context | Dataset Size | 💲Free | AlpacaEval (win rate %) | MT-bench (win rate adjusted %) | MT-bench (score) |
|---|---|---|---|---|---|---|---|
| v.s. text-davinci-003 | v.s. ChatGPT | ||||||
| GPT-4 | 1.8T* | 8K | ❌ | 95.3 | 82.5 | 8.99 | |
| ChatGPT | 175B* | 4K | ❌ | 89.4 | 50.0 | 7.94 | |
| Llama-2-70B-Chat | 70B | 4K | 2.9M | ✅ | 92.7 | 60.0 | 6.86 |
| OpenChat 3.2 SUPER | 13B | 4K | 80K | ✅ | 89.5 | 57.5 | 7.19 |
| Llama-2-13B-Chat | 13B | 4K | 2.9M |
From the published model card. Full card on the HuggingFace links in the sidebar.
Using it via the API
Once AxForge deploys open-2-super for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (open-2-super below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/chat/completions \
-H "Authorization: Bearer $AXFORGE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"open-2-super","messages":[{"role":"user","content":"Hello"}]}'
Create an account — your API key is available in the console. 5M tokens/month currently included with every new account at launch.