Model reference · open weights

omnistep-12a3b

LLMs sovthpaw · community Omni (any→any) · MoE 1 build Open weights 580 dl/mo

omnistep-12a3b is an open-weight language model from sovthpaw. omnistep-12a3b (BF16) weighs 2.1 GB; the smallest configuration that runs it is RTX 3060 12 GB.

What it is

Released bysovthpaw
TypeLanguage models
TaskOmni (any→any) · MoE
Parameters (lead)10.5B
Runs withtransformers
Based onsovthpaw/omnistep-12a3b
Released2026-06-04
Popularity580 downloads / month
Weights2.1 GB (omnistep-12a3b (BF16), file size)
LicenceOpen weights

What it runs on

Memory and cards for omnistep-12a3b (BF16)

Weights 2.1 GB (file size) · KV cache 37 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 673 MB on a small card · context up to 32,768 tokens.

CardRequests at once
8K tokens each
Requests at once
32K tokens each
Longest single
request
Counted
memory
RTX 3060 12 GB297all 32K11.6 GB
RTX 4060 Ti 16 GB4110all 32K15.4 GB
RTX 3090 24 GB6817all 32K23.4 GB
RTX 4090 24 GB6817all 32K23.4 GB
RTX 5090 32 GB9323all 32K31.0 GB
L40S 48 GB13634all 32K44.0 GB
A100 80 GB24962all 32K78.2 GB
H100 80 GB24962all 32K78.1 GB
RTX PRO 6000 Blackwell 96 GB30175all 32K93.8 GB
DGX Spark (GB10) 128 GB unified34686all 32K107 GB
H200 141 GB447111all 32K138 GB
B200 180 GB574143all 32K176 GB
Memory needed at each load
Requests at once8K tokens each32K tokens each
13.1 GB4.0 GB
54.3 GB8.8 GB
85.2 GB12.4 GB
167.6 GB22.1 GB
3212.4 GB41.4 GB
6422.1 GB80.1 GB

On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.

Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (grouped-query attention); the overhead is an estimate of llama.cpp's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). Assumes vLLM 0.10 or later.

From the model card

What sovthpaw says about omnistep-12a3b

A multimodal voice-and-music AI, born from generational Darwin family evolution. OmniStep 12A3B is a personal AI companion — a vibe coach that takes notes for you, keeps up the conversation, and plays background music that matches your mood. All the while it self-evolves to become a better assistant to you, via the Darwin Family weight-space recombination methodology (arXiv:2605.14386). Built as a paper-exact 2-parent merge of Qwen2.5-Omni-3B (multimodal) and ACE-Step v1.5 XL SFT 4B (text-to-music).

The text body of the model was produced by a paper-exact Darwin 2-parent weight-space recombination of the Qwen2.5-Omni thinker and the ACE-Step text encoder, with the Architecture Mapper's "skip on dim mismatch" behavior preserving the Omni text body intact across the Qwen2.5/Qwen3 cross-architecture boundary. The diffusion (music) head sits at F16 (unquantized) for maximum audio quality. The transformer (text/multimodal) head is shipped in 4 quantized GGUF deployments (F16, Q8_0, Q4_K_M, Q4_0) for llama.cpp users.

Read the full model card

The OmniStep Evolutionary Radio is the operational version of "infinitely generate its own background music" — a 4-loop pipeline (playback + queue fill + GEPA prompt evolution + Darwin weight evolution) wired up in the evolutionary-radio skill.

🎧 Listen to the examples

1. 🎵 Lo-Fi — chill lofi beats for late-night coding

chill lofi beats, mellow hip-hop, soft piano keys, vinyl crackle, late-night study vibes, 75 bpm, instrumental

🎤 Voice intro (text generated by OmniStep 12A3B, speech by Soprano 80M)

🎵 The track

2. 🎬 Movie Orchestra — epic cinematic orchestral

epic cinematic orchestral soundtrack, sweeping strings, French horns, building tension, Hans Zimmer style, 90 bpm, instrumental

🎤 Voice intro

🎵 The track

3. 🔥 Dark Metal — heavy dark metal for dark times

heavy dark metal, blast beats, down-tuned 7-string guitars, atmospheric, blackened death metal, 180 bpm, instrumental

🎤 Voice intro

🎵 The track

  • 🧠 Vibe coach — reads the room, matches your mood, suggests what to play next
  • 📝 Note-taker — listens to your conversation and captures the bits you want to remember
  • 💬 Conversational companion — keeps up a real back-and-forth, asks the follow-up questions
  • 🎵 Background music that matches the vibe — generates infinite music that fits what you're doing, in any style
  • 🔁 Self-evolving — gets better at being your assistant over time, via the Darwin Family weight-space evolution methodology
  • 🎤 Real-time ASR + TTS — Whisper audio in, Talker + token2wav audio out (4o-style streaming voice)
  • 🖼️ Image understanding — NaViT vision encoder

All in one model. Run it with vllm, llama-server, or the included Python scripts.

🎛 Pick your quantization — download just the one you need

The GGUFs are independent files. Download only the one that fits your VRAM — you don't need all of them. Pick from the table below.

QuantSizeVRAMBest forDownload
F166.4GB6.4GBMaximum quality, plenty of VRAM⬇ omnistep-12a3b-f16.gguf
Q8_03.4GB3.4GBNear-F16 quality, balanced⬇ omnistep-12a3b-q8_0.gguf
Q4_K_M2.0GB2.0GBRecommended — best size/quality tradeoff⬇ omnistep-12a3b-q4_k_m.gguf
Q4_01.9GB1.9GBSmallest, lowest quality⬇ omnistep-12a3b-q4_0.gguf

Run any of them with llama-server (the Omni build of llama.cpp is in the HF model comments / wiki):

llama-server -m omnistep-12a3b-q4_k_m.gguf -ngl 99 --port 8080 --host 0.0.0.0 -c 8192

Quick start

Option 1 — vllm (the easiest, full multimodal)

pip install vllm
vllm serve sovthpaw/omnistep-12a3b \
    --max-model-len 32768 \
    --gpu-memory-utilization 0.85 \
    --trust-remote-code

Then in another terminal:

curl -s http://localhost:8000/v1/chat/completions \
    -H "Content-Type: application/json" \
    -d '{
        "model": "sovthpaw/omnistep-12a3b",
        "messages": [{"role": "user", "content": "Take a note: I need to follow up with the design team about the Q3 launch."}],
        "max_tokens": 200
    }'

See vllm Qwen2.5-Omni docs for the full multimodal API.

Option 2 — llama-server with the GGUFs (fast text path)

# Pick your quantization based on VRAM
#   F16  = 6.4GB VRAM, best quality
#   Q8_0 = 3.4GB VRAM, near-F16 quality
#   Q4_K_M = 2.0GB VRAM, recommended
#   Q4_0 = 1.9GB VRAM, smallest

llama-server \
    -m omnistep-12a3b-q4_k_m.gguf \
    -ngl 99 \
    --port 8080 \
    --host 0.0.0.0 \
    -c 8192

curl -s http://localhost:8080/v1/chat/completions \
    -H "Content-Type: application/json" \
    -d '{
        "model": "omnistep-12a3b",
        "messages": [{"role": "user", "content": "I am coding late, give me a lofi vibe and a note about what I should focus on tomorrow."}],
        "max_tokens": 200,
        "stream": true
    }'

The 4 GGUFs are deployment options — pick whichever fits your hardware. Q4_K_M is the recommended sweet spot.

Option 3 — the included Python scripts (the most fun)

The repo includes Python scripts that wire everything together for the headline use cases. After cloning:

git clone https://huggingface.co/sovthpaw/omnistep-12a3b
cd omnistep-12a3b

# Start the vllm server (one terminal)
python scripts/run_omnistep_12a3b.py serve

# In another terminal — try the modalities
python scripts/run_omnistep_12a3b.py te

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms