Model reference · open weights

Nous-Hermes-2-Mixtral-8x7B-SFT

LLMs NousResearch Text gen 1 build Open weights 4k dl/mo

Nous-Hermes-2-Mixtral-8x7B-SFT is an open-weight language model from NousResearch. Nous-Hermes-2-Mixtral-8x7B-SFT (BF16) weighs 93.4 GB; the smallest configuration that runs it is DGX Spark (GB10) 128 GB unified.

Nous-Hermes-2-Mixtral-8x7B-SFT is a 46.7B parameter text-generation model developed by NousResearch. It is a supervised finetune of the Mixtral 8x7B MoE LLM, designed for English language tasks with a context length of 32,768 tokens. The model is distributed under the Apache 2.0 license.

Summary of the NousResearch/Nous-Hermes-2-Mixtral-8x7B-SFT model card, 2026-10-01

What it is

Released byNousResearch
TypeLanguage models
TaskText gen
Parameters (lead)46.7B
Context32,768 tokens
Runs withtransformers
Based onmistralai/Mixtral-8x7B-v0.1
Released2023-12-26
Popularity4k downloads / month
Weights93.4 GB (Nous-Hermes-2-Mixtral-8x7B-SFT (BF16), file size)
LicenceOpen weights

What it runs on

Memory and cards for Nous-Hermes-2-Mixtral-8x7B-SFT (BF16)

Weights 93.4 GB (file size) · KV cache 131 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 647 MB on a small card · context up to 32,768 tokens.

CardRequests at once
8K tokens each
Requests at once
32K tokens each
Longest single
request
Counted
memory
RTX 3060 12 GB … RTX PRO 6000 Blackwell 96 GB
9 smaller cards
———
DGX Spark (GB10) 128 GB unified92all 32K107 GB
H200 141 GB389all 32K138 GB
B200 180 GB7218all 32K176 GB
4× RTX 5090 32 GB
tensor parallel
266all 32K31.0 GB a card
2× H100 80 GB
tensor parallel
5112all 32K78.1 GB a card
2× A100 80 GB
tensor parallel
5714all 32K78.2 GB a card
4× L40S 48 GB
tensor parallel
7418all 32K44.0 GB a card
2× RTX PRO 6000 Blackwell 96 GB
tensor parallel
8020all 32K93.8 GB a card
Memory needed at each load
Requests at once8K tokens each32K tokens each
195.1 GB98.3 GB
599.4 GB116 GB
8103 GB128 GB
16111 GB163 GB
32128 GB231 GB
64163 GB369 GB

On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.

Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (grouped-query attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.

From the model card

What NousResearch says about Nous-Hermes-2-Mixtral-8x7B-SFT

Read the model card

Model description

Nous Hermes 2 Mixtral 8x7B SFT is the supervised finetune only version of our new flagship Nous Research model trained over the Mixtral 8x7B MoE LLM.

The model was trained on over 1,000,000 entries of primarily GPT-4 generated data, as well as other high quality data from open datasets across the AI landscape, achieving state of the art performance on a variety of tasks.

This is the SFT only version of Mixtral Hermes 2, we have also released an SFT+DPO version, for people to find which works best for them, which can be found here: https://huggingface.co/NousResearch/Nous-Hermes-2-Mixtral-8x7B-DPO

We are grateful to Together.ai for sponsoring our compute during the many experiments both training Mixtral and working on DPO!

Table of Contents

  1. Example Outputs
  2. Benchmark Results
    • GPT4All
    • AGIEval
    • BigBench
    • Comparison to Mixtral-Instruct
  3. Prompt Format
  4. Inference Example Code
  5. Quantized Models

Example Outputs

Writing Code for Data Visualization

Writing Cyberpunk Psychedelic Poems

Performing Backtranslation to Create Prompts from Input Text

Benchmark Results

Nous-Hermes 2 on Mixtral 8x7B SFT is the bedrock for major improvements on many of the benchmarks below compared to the base Mixtral model, and is the SFT only version of our first model to beat the flagship Mixtral Finetune by MistralAI (the DPO version).

GPT4All:

|    Task     |Version| Metric |Value |   |Stderr|
|-------------|------:|--------|-----:|---|-----:|
|arc_challenge|      0|acc     |0.5904|±  |0.0144|
|             |       |acc_norm|0.6323|±  |0.0141|
|arc_easy     |      0|acc     |0.8594|±  |0.0071|
|             |       |acc_norm|0.8607|±  |0.0071|
|boolq        |      1|acc     |0.8783|±  |0.0057|
|hellaswag    |      0|acc     |0.6592|±  |0.0047|
|             |       |acc_norm|0.8434|±  |0.0036|
|openbookqa   |      0|acc     |0.3400|±  |0.0212|
|             |       |acc_norm|0.4660|±  |0.0223|
|piqa         |      0|acc     |0.8324|±  |0.0087|
|             |       |acc_norm|0.8379|±  |0.0086|
|winogrande   |      0|acc     |0.7569|±  |0.0121|

Average: 75.36

AGIEval:

|             Task             |Version| Metric |Value |   |Stderr|
|------------------------------|------:|--------|-----:|---|-----:|
|agieval_aqua_rat              |      0|acc     |0.2441|±  |0.0270|
|                              |       |acc_norm|0.2598|±  |0.0276|
|agieval_logiqa_en             |      0|acc     |0.4025|±  |0.0192|
|                              |       |acc_norm|0.3978|±  |0.0192|
|agieval_lsat_ar               |      0|acc     |0.2391|±  |0.0282|
|                              |       |acc_norm|0.2043|±  |0.0266|
|agieval_lsat_lr               |      0|acc     |0.5353|±  |0.0221|
|                              |       |acc_norm|0.5098|±  |0.0222|
|agieval_lsat_rc               |      0|acc     |0.6617|±  |0.0289|
|                              |       |acc_norm|0.5948|±  |0.0300|
|agieval_sat_en                |      0|acc     |0.7961|±  |0.0281|
|                              |       |acc_norm|0.7816|±  |0.0289|
|agieval_sat_en_without_passage|      0|acc     |0.4757|±  |0.0349|
|                              |       |acc_norm|0.4515|±  |0.0348|
|agieval_sat_math              |      0|acc     |0.4818|±  |0.0338|
|                              |       |acc_norm|0.3909|±  |0.0330|

Average: 44.89

BigBench:

|                      Task                      |Version|       Metric        |Value |   |Stderr|
|------------------------------------------------|------:|---------------------|-----:|---|-----:|
|bigbench_causal_judgement                       |      0|multiple_choice_grade|0.5789|±  |0.0359|
|bigbench_date_understanding                     |      0|multiple_choice_grade|0.7154|±  |0.0235|
|bigbench_disambiguation_qa                      |      0|multiple_choice_grade|0.5388|±  |0.0311|
|bigbench_geometric_shapes                       |      0|multiple_choice_grade|0.4680|±  |0.0264|
|                                                |       |exact_str_match      |0.0000|±  |0.0000|
|bigbench_logical_deduction_five_objects         |      0|multiple_choice_grade|0.3260|±  |0.0210|
|bigbench_logical_deduction_seven_objects        |      0|multiple_choice_grade|0.2443|±  |0.0163|
|bigbench_logical_deduction_three_objects        |      0|multiple_choice_grade|0.5233|±  |0.0289|
|bigbench_movie_recommendation                   |      0|multiple_choice_grade|0.3700|±  |0.0216|
|bigbench_navigate                               |      0|multiple_choice_grade|0.5000|±  |0.0158|
|bigbench_reasoning_about_colored_objects        |      0|multiple_choice_grade|0.6665|±  |0.0105|
|bigbench_ruin_names                             |      0|multiple_choice_grade|0.6317|±  |0.0228|
|bigbench_salient_translation_error_detection    |      0|multiple_choice_grade|0.2505|±  |0.0137|
|bigbench_snarks                                 |      0|multiple_choice_grade|0.7127|±  |0.0337|
|bigbench_sports_understanding                   |      0|multiple_choice_grade|0.6592|±  |0.0151|
|bigbench_temporal_sequences                     |      0|multiple_choice_grade|0.6860|±  |0.0147|
|bigbench_tracking_shuffled_objects_five_objects |      0|multiple_choice_grade|0.2200|±  |0.0117|
|bigbench_tracking_shuffled_objects_seven_objects|      0|multiple_choice_grade|0.1503|±  |0.0085|
|bigbench_tracking_shuffled_objects_three_objects|      0|multiple_choice_grade|0.5233|±  |0.0289|

Average: 48.69

Benchmark Comparison Charts

GPT4All

AGI-Eval

BigBench Reasoning Test

Prompt Format

Nous Hermes 2 uses ChatML as the prompt format, opening up a much more structured system for engaging the LLM in multi-turn chat dialogue.

System prompts allow steerability and interesting new ways to inte

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms