Model reference · open weights

acestep-xl-sft

acestep-xl-sft is an open-weight audio or speech model from ACE-Step, listed in the AxForge catalogue. AxForge can bring it up on EU-owned hardware for you on request — with the licence handled where one is required.

Audio ACE-Step 2 variants 6k downloads/mo
Request this model on EU hardware All served models Not on the shared API today — deployed on request.

About

What acestep-xl-sft is

Model Details This is the XL (4B) SFT variant of ACE-Step 1.5 — a supervised fine-tuned model with ~4B parameters. SFT provides higher audio quality with CFG (Classifier-Free Guidance) support for fine-grained prompt adherence control. XL Architecture GPU Requirements All LM models (0.6B / 1.7B / 4B) are fully compatible with XL. Key Features - 💰 Commercial-Ready: Trained on legally compliant datasets. Generated music can be used for commercial purposes. - 📚 Safe Training Data: Licensed music, royalty-free/public domain, and synthetic (MIDI-to-Audio) data. - 🎯 CFG Support: Fine-tune prompt adherence with guidance scale control. - 🔮 Highest Quality: SFT + 4B parameters = the highest quality variant. Quick Start Model Zoo XL (4B) DiT Models LM Models (all compatible with XL) Acknowledgements This project is co-led by ACE Studio and StepFun. Citation

Summarised from the published model card. Read the full card on the HuggingFace links below.

Specifications

What it is

MakerACE-Step
TypeAudio & music
Parameters (lead)5.0B
Variants2
Runs withtransformers
Released2026-04-02
Popularity6k downloads / month
Likes96
LicenceOpen weights

How it works

How audio & music work

Audio or textinputAudio modelrecognise / synthesiseText or audiooutputSpeech-to-text turns audio into text; text-to-speech and music models turn text into audio.

Variants

Sizes & precisions

Open weights ship in several sizes and precisions. One page, all the variants — pick the one that fits your GPU. VRAM figures are estimates from model size.

VariantParamsPrecisionVRAMFits 16 GBWeights
acestep-v15-xl-sft5.0BBF16~11.5 GBWeights ↗
acestep-v15-xl-sft-diffusers4.2BBF16~9.6 GBWeights ↗

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys acestep-xl-sft for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (acestep-xl-sft below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -F model="acestep-xl-sft" -F file=@audio.mp3

Details

Languages, data & research

Tags

transformers safetensors acestep feature-extraction audio music text2music custom_code text-to-audio diffusers text-to-music flow-matching diffusers:AceStepPipeline

Papers

Licence

Open weights

Open weights under mit — commercial use is permitted. Deploy it on AxForge EU hardware on request. Read the licence ↗

Sources

Weights & code

Want acestep-xl-sft on EU-owned hardware?

Request this model on EU hardware See what’s served now

Explore

More audio & music

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms