Model reference · open weights
neutrino is an open-weight language model from neuralcrew. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.
Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.
What it is
| Released by | neuralcrew |
|---|---|
| Type | Language models |
| Task | Text gen |
| Context | 32k tokens |
| Runs with | gguf |
| Based on | neuralcrew/neutrino |
| Released | 2025-09-14 |
| Popularity | 36k downloads / month |
| Licence | Open weights |
About
A 7B-parameter instruction-tuned language model for conversational AI
| Field | Value |
|---|---|
| Developer | Fardeen NB |
| Model type | Autoregressive transformer, decoder-only |
| Parameters | 7B |
| Base model | neuralcrew/neutrino |
| Fine-tuning | Instruction tuning for multi-turn dialogue |
| Language(s) | English |
| Format | GGUF (quantized for llama.cpp, Ollama, llama-cpp-python) |
| License | Apache 2.0 |
| Version | 2.0 |
Neutrino-Instruct is a 7B-parameter language model fine-tuned from the Neutrino base model for conversational and instruction-following tasks. It is distributed in GGUF format for efficient local inference on consumer hardware via llama.cpp, Ollama, and llama-cpp-python.
The model is designed to hold coherent, contextual dialogue across multiple turns and to follow natural-language instructions reliably at chat scale, while remaining light enough to run on a single consumer GPU or CPU-only machine.
Primary use cases
Out of scope
Neutrino-Instruct is a general-purpose research and hobbyist model. It has not been evaluated for production or safety-critical deployment, and Fardeen NB makes no warranty as to its factual accuracy, safety, or fitness for any particular purpose.
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp && make
# Single prompt
./main -m ./neutrino-instruct.gguf -p "Hello, who are you?"
# Interactive chat
./main -m ./neutrino-instruct.gguf -i -p "Let's chat."
# Control output length
./main -m ./neutrino-instruct.gguf -n 256 -p "Write a poem about stars."
# Adjust temperature
./main -m ./neutrino-instruct.gguf --temp 0.7 -p "Explain quantum computing simply."
# GPU offload (if built with CUDA/Metal)
./main -m ./neutrino-instruct.gguf --gpu-layers 50 -p "Summarize this article."
ollama run fardeen0424/neutrino
from llama_cpp import Llama
llm = Llama(model_path="./neutrino-instruct.gguf")
response = llm("Who are you?")
print(response["choices"][0]["text"])
# Streaming
for token in llm("Tell me a story about Neutrino:", stream=True):
print(token["choices"][0]["text"], end="", flush=True)
Neutrino-Instruct's base model was pretrained on a mixture of web, encyclopedic, code, and long-document text, and instruction-tuned on conversational data. Component sources include:
| Dataset | Type | Role |
|---|---|---|
| finepdfs | Long-form documents | Pretraining |
| finewiki | Encyclopedic text | Pretraining |
| fineweb-edu-100b-shuffle | Educational web text | Pretraining |
| wikipedia | Encyclopedic text | Pretraining |
| github-code | Source code | Pretraining |
| TinyStories | Short narrative text | Fine-tuning / eval |
| awesome-chatgpt-prompts | Instruction/persona prompts | Instruction tuning |
Full dataset links are available in the metadata panel on the right side of this page.
| Field | Value |
|---|---|
| Context length | 32,768 tokens |
| Precision | bfloat16 |
Neutrino-Instruct is built on a standard decoder-only Transformer stack. The core computations are as follows.
Scaled dot-product self-attention, applied per head:
$$\text{Attention}(Q, K, V) = \text{softmax}!\left(\frac{QK^\top}{\sqrt{d_k}}\right)V$$
with multi-head attention combining $h$ heads via a learned output projection:
$$\text{MultiHead}(Q,K,V) = \text{Concat}(\text{head}_1, \dots, \text{head}_h),W^O, \qquad \text{head}_i = \text{Attention}(QW_i^Q, KW_i^K, VW_i^V)$$
Neutrino-Instruct uses grouped-query attention (GQA): the 32 query heads are partitioned into $g=8$ groups, each group sharing a single key/value projection ($KW^K_g, VW^V_g$) instead of every query head having its own — trading a small amount of expressivity for a much smaller KV cache at inference time.
Position encoding — rotary position embeddings (RoPE), applied to query/key vectors by rotating consecutive coordinate pairs as a function of token position $m$ and frequency $\theta_i$:
$$f(x_m, m) = \big[x_m^{(1)}\cos m\theta_i - x_m^{(2)}\sin m\theta_i,\ \ x_m^{(1)}\sin m\theta_i + x_m^{(2)}\cos m\theta_i\big]$$
Feed-forward block (SwiGLU-style gating):
$$\text{FFN}(x) = \big(\text{Swish}(xW_1)\otimes xW_3\big)W_2$$
Pre-normalization (RMSNorm) around each sub-block:
$$\text{RMSNorm}(x) = \frac{x}{\sqrt{\frac{1}{n}\sum_{i=1}^n x_i^2 + \epsilon}} \odot g$$
Training objective — standard next-token cross-entropy over the vocabulary:
$$\mathcal{L} = -\sum_{t=1}^{T} \log P_\theta(x_t \mid x_{<t})$$
| Hyperparameter | Value |
|---|---|
| Layers | 32 |
| Hidden size | 4096 |
| Intermediate (FFN) size | 14336 |
| Attention heads | 32 |
| KV heads (GQA) | 8 |
| Vocabulary size | 32,768 |
| Max context length | 32,768 |
| Activation | SiLU (SwiGLU gating) |
| Normalization | RMSNorm (pre-norm, $\epsilon=1\text{e-}5$) |
| Position encoding | RoPE ($\theta=1{,}000{,}000$) |
| Sliding window | None (full attention) |
| Precision | bfloat16 |
| Tied embeddings | No |
| Setup | VRAM / RAM | Recommended quantization |
|---|---|---|
| CPU-only | 32–64 GB RAM | Q4_K_M / Q5_K_M |
| GPU (entry) | 4 GB VRAM | Q4 |
| GPU (mid) | 8 GB VRAM | Q5 / Q |
From the published model card. Full card on the HuggingFace links in the sidebar.
Using it via the API
Once AxForge deploys neutrino for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (neutrino below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/chat/completions \
-H "Authorization: Bearer $AXFORGE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"neutrino","messages":[{"role":"user","content":"Hello"}]}'
Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.