Available as managed deploymentAudiocisco-aiAudio→audio1 variants552 dl/mo
stupase is an open-weight audio or speech model from cisco-ai. AxForge deploys and operates it
for you on dedicated EU-owned hardware — with the licence handled where one is required.
Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.
What it is
Released by
cisco-ai
Type
Audio & music
Task
Audio→audio
Released
2026-06-22
Popularity
552 downloads / month
Licence
Open weights
About
What stupase is
StuPASE is a state-of-the-art generative speech enhancement model trained to remove noise and reverberation while preserving linguistic content and speaker identity, and achieving studio-level perceptual quality. It operates on 16 kHz mono audio.
Read the full model card
Model Details
Model Description
StuPASE contains three main components:
DeWavLM-R: Performs low-hallucination phonetic enhancement, fine‑tuned from DeWavLM using dry targets for improved dereverberation.
These source datasets were used to prepare training mixtures and train the released model. The model card and repository do not redistribute the underlying dataset contents; please refer to the original dataset pages and licenses below.
Dataset Attribution
DNS5 Challenge clean speech (LibriVox subset): clean-speech material prepared from LibriVox through the DNS Challenge. The LibriVox recordings used for this portion are public domain and were used as clean-speech training data for the released checkpoint.
LibriSpeech: LibriSpeech by Vassil Panayotov et al., licensed under CC BY 4.0. It was used as clean-speech training data for the released checkpoint.
LibriTTS: LibriTTS by Heiga Zen et al., licensed under CC BY 4.0. It was used as clean-speech training data for the released checkpoint.
VCTK Corpus: the VCTK dataset from the Centre for Speech Technology Research, University of Edinburgh, licensed under CC BY 4.0. It was used as clean-speech training data for the released checkpoint.
DNS5 Challenge noise resources: noise data prepared through the DNS Challenge and used to synthesize noisy training mixtures for the released checkpoint. For this release, the DNS5 noise resources draw on AudioSet material licensed under CC BY 4.0, selected Freesound files licensed under CC0 1.0, and DEMAND environmental recordings licensed under CC BY-SA 3.0.
OpenSLR26 and OpenSLR28: OpenSLR26 and OpenSLR28 room impulse response resources, both licensed under Apache 2.0, were used to add reverberation during training.
The performance of the retrained version compared to the original one:
Model
DNSMOS
UTMOS
SBS
LPS
SpkSim
WER (%)
From the published model card. Full card on the HuggingFace links in the sidebar.
How it works
How audio & music work
Using it via the API
Call it like any OpenAI endpoint
Once AxForge deploys stupase for you, it answers on the OpenAI-compatible API — the same
base URL and keys as every other model. (stupase below is illustrative; you get the exact
model name on deployment.)