Model reference · open weights

minimax-music3

minimax-music3 is an open-weight audio or speech model from ChrisColeTech, listed in the AxForge catalogue. AxForge can bring it up on EU-owned hardware for you on request — with the licence handled where one is required.

Licence fee required Audio ChrisColeTech 1 variants 5k downloads/mo
Request a licence + hosting quote All served models Not on the shared API today — deployed on request.

About

What minimax-music3 is

MiniMax Music 3 GGUF The MiniMax Music 3 autoregressive text encoder + flow-matching DiT, converted to GGUF. Complete songs with vocals from a caption (music description) + lyrics. Up to ~5 min, stereo 44.1 kHz. Samples Synth-Pop - 2:29 min (Q80 GGUF) - ~7.3 min on an RTX 5090 (32 GB), peak 9.7 GiB. Punk-funk — 1:01 (Q80 GGUF) - ~5 min on an RTX 5090 (32 GB), peak 9.7 GiB. Components used to generate ⚠ Must use the updated GGUF Loader ⚠ In comfy: - Open the ComfyUI Manager - Change the channel to "Channel (remote)" - and search for comfyui-gguf-loader Use the workflow Link Caption in three parts: - global metadata (genre, BPM, key, mood) - vocal details - arrangement. Note: Describe the vocal explicitly (ie. "warm female vocal, close and breathy" etc.) or it drifts. Structure tags each on their own line: [intro] [verse] [pre-chorus] [chorus] [bridge] [instrumental] [solo] [outro]. Note: These are the only tags the model uses. Recommended settings Sources

Summarised from the published model card. Read the full card on the HuggingFace links below.

Specifications

What it is

MakerChrisColeTech
TypeAudio & music
Variants1
Runs withgguf
Based onMiniMaxAI/MiniMax-Music3
Released2026-08-14
Popularity5k downloads / month
Likes11
LicenceCommercial licence needed

How it works

How audio & music work

Audio or textinputAudio modelrecognise / synthesiseText or audiooutputSpeech-to-text turns audio into text; text-to-speech and music models turn text into audio.

Variants

Sizes & precisions

Open weights ship in several sizes and precisions. One page, all the variants — pick the one that fits your GPU. VRAM figures are estimates from model size.

VariantParamsPrecisionVRAMFits 16 GBWeights
minimax-music3-GGUFGGUFWeights ↗

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys chriscoletech-minimax-music3 for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (chriscoletech-minimax-music3 below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -F model="chriscoletech-minimax-music3" -F file=@audio.mp3

Details

Languages, data & research

Tags

gguf text-to-music music-generation quantized minimax minimax-music3 text-to-audio endpoints_compatible conversational

Licence

Commercial licence needed

The weights are open but its licence needs a commercial agreement for business use. AxForge can arrange that licence and host the model for you — you pay AxForge, we settle with the model’s maker. Ask us for a quote. Read the licence ↗

Sources

Weights & code

Want minimax-music3 on EU-owned hardware?

Request a licence + hosting quote See what’s served now

Explore

More audio & music

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms