Model reference · open weights

MiniMax-H3-comfyUI

MiniMax-H3-comfyUI is an open-weight video model from vantagewithai, listed in the AxForge catalogue. AxForge can bring it up on EU-owned hardware for you on request — with the licence handled where one is required.

Licence fee required Video vantagewithai 1 variants 63k downloads/mo
Request a licence + hosting quote All served models Not on the shared API today — deployed on request.

About

What MiniMax-H3-comfyUI is

GGUF quants of Minimax-H3 files for ComfyUI. Original model repository: https://huggingface.co/MiniMaxAI/MiniMax-H3 Watch us on Youtube: @VantageWithAI MiniMax H3 System Overview MiniMax H3 is a general-purpose, omni-modal generative system. It supports unified understanding of multimodal contexts composed of text, images, video, and audio, and can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds. Thanks to its task-generalization-oriented system design, H3 already possesses broad multimodal context understanding and generation capabilities at the pre-training stage, enabling outstanding performance in following complex multimodal instructions. H3 supports the following input and output specifications: Model Variants and Input Specifications The complete H3 system consists of the following three modules: - H3-Context-IR: As inputs become increasingly complex, we build a dedicated system to deeply understand and refine the input multimodal instructions, then convert them into a form that H3 can readily understand—the Context Intermediate Representation—for generation. H3-Context-IR is critical to the quality of the final output, so we strongly recommend incorporating it into your generation pipeline or following the “Prompting Guidance” to build your own context-processing system. - H3-Base: Generates audio and video based on the H3-Context-IR output, producing results at 768p resolution. - H3-Regenerate-2K: Feeds the 768p result together with the original context back into H3 to regenerate the output at 2K resolution. This process leverages both H3’s powerful generative capabilities and the rich information contained in the original context, enabling it to produce high-resolution outputs with more accurate details and greater visual fidelity. Online API Use MiniMax\-H3 directly via API\. - Global: platform\.minimax\.io \| CN: platform\.minimaxi\.com Online App Use MiniMax\-H3 directly via App\. - WebApp Global: hailuoai\.video \| CN: hailuoai\.com - Desktop Global: hub\.minimax\.io \| CN: hub\.minimaxi\.com Model Architecture H3\-Context\-IR H3\-Context\-IR is a hosted preprocessing and orchestration system design

Summarised from the published model card. Read the full card on the HuggingFace links below.

Specifications

What it is

Makervantagewithai
TypeVideo models
Variants1
Runs withdiffusers
Released2026-08-04
Popularity63k downloads / month
Likes17
LicenceCommercial licence needed

How it works

How video models work

Prompt / imagestart pointTemporal diffusionframes over timeVideoMP4 clipA video model generates a sequence of coherent frames from your prompt or a starting image.

Variants

Sizes & precisions

Open weights ship in several sizes and precisions. One page, all the variants — pick the one that fits your GPU. VRAM figures are estimates from model size.

VariantParamsPrecisionVRAMFits 16 GBWeights
MiniMax-H3-comfyUI-GGUFGGUFWeights ↗

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys minimax-h3-comfyui for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (minimax-h3-comfyui below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/videos/generations \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"minimax-h3-comfyui","prompt":"a drone shot over a forest"}'

Details

Languages, data & research

Tags

diffusers gguf text-to-video image-to-video image-text-to-video video-to-video text-to-audio-video image-to-audio-video image-text-to-audio-video video-to-audio-video audio-to-audio-video audio-video-generation multimodal synchronized-audio-video

Licence

Commercial licence needed

The weights are open but its licence needs a commercial agreement for business use. AxForge can arrange that licence and host the model for you — you pay AxForge, we settle with the model’s maker. Ask us for a quote. Read the licence ↗

Sources

Weights & code

Want MiniMax-H3-comfyUI on EU-owned hardware?

Request a licence + hosting quote See what’s served now

Explore

More video models

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms