Model reference · open weights

midashenglm-gen

midashenglm-gen is an open-weight audio or speech model from mispeech, listed in the AxForge catalogue. AxForge can bring it up on EU-owned hardware for you on request — with the licence handled where one is required.

Audio mispeech 1 variants 1k downloads/mo
Request this model on EU hardware All served models Not on the shared API today — deployed on request.

About

What midashenglm-gen is

MiDashengLM-Gen [](https://arxiv.org/abs/2608.11804)&nbsp;&nbsp;[](https://huggingface.co/mispeech/midashenglm-gen)&nbsp;&nbsp;[](https://xingws.github.io/midashenglm-gen-demo/)&nbsp;&nbsp;[](https://github.com/xiaomi-research/midashenglm-gen) English | 中文 MiDashengLM-Gen is an end-to-end framework that uses a pre-trained Large Language Model and audio tokenizer as the backbone, combined with per-token conditional flow matching for autoregressive, variable-length mixed-audio scene generation. It generates coherent 16 kHz audio scenes that simultaneously blend speech, music, sound effects, and environmental acoustics from structured text descriptions. Architecture Left: training pipeline with flow matching loss. Right: autoregressive inference pipeline. Key components: Input Format Input uses structured multi-view captions with special tokens to describe different aspects of an audio scene. Use <|unknown| for absent elements. Installation Quick Start Batch Generation Generation Parameters Citation License Apache 2.0 Use Restrictions You are solely responsible for your use of MiDashengLM-Gen and any outputs, actions, or consequences arising therefrom, and you agree not to use MiDashengLM-Gen or any derivatives thereof: - For any unlawful, fraudulent, or malicious purpose, or in any manner that violates any applicable laws or regulations; - To infringe upon the intellectual property rights, privacy rights, publicity rights, or other lawful rights or interests of any third party; - To exploit, harm, harass, defame, unlawfully discriminate against, or otherwise adversely affect any individual or group, including minors or vulnerable persons; - For any military purpose or application.

Summarised from the published model card. Read the full card on the HuggingFace links below.

Specifications

What it is

Makermispeech
TypeAudio & music
Parameters (lead)2.9B
Variants1
Runs withtransformers
Released2026-08-12
Popularity1k downloads / month
Likes41
LicenceOpen weights

How it works

How audio & music work

Audio or textinputAudio modelrecognise / synthesiseText or audiooutputSpeech-to-text turns audio into text; text-to-speech and music models turn text into audio.

Variants

Sizes & precisions

Open weights ship in several sizes and precisions. One page, all the variants — pick the one that fits your GPU. VRAM figures are estimates from model size.

VariantParamsPrecisionVRAMFits 16 GBWeights
midashenglm-gen2.9BBF16~6.6 GBWeights ↗

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys midashenglm-gen for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (midashenglm-gen below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -F model="midashenglm-gen" -F file=@audio.mp3

Details

Languages, data & research

Tags

transformers safetensors midashenglm-gen feature-extraction audio-generation text-to-audio flow-matching dasheng custom_code

Papers

Licence

Open weights

Open weights under apache-2.0 — commercial use is permitted. Deploy it on AxForge EU hardware on request. Read the licence ↗

Sources

Weights & code

Want midashenglm-gen on EU-owned hardware?

Request this model on EU hardware See what’s served now

Explore

More audio & music

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms