Model reference · open weights

SenseVoiceSmall

SenseVoiceSmall is an open-weight audio or speech model from FunAudioLLM, listed in the AxForge catalogue. AxForge can bring it up on EU-owned hardware for you on request — with the licence handled where one is required.

Licence fee required Audio FunAudioLLM 1 variants 33k downloads/mo
Request a licence + hosting quote All served models Not on the shared API today — deployed on request.

About

What SenseVoiceSmall is

(简体中文|English|日本語) ⭐ Powered by FunASR — please give us a GitHub Star! SenseVoice is part of the FunASR ecosystem — one industrial-grade open-source toolkit for ASR · VAD · punctuation · speaker diarization · emotion / event · LLM-ASR. A Star really helps the project (and keeps you updated): 🌟 FunASR · 🌟 SenseVoice · 🌟 Fun-ASR · 🌟 FunClip ⚡ CPU / edge — no GPU, no Python: run SenseVoiceSmall as a single self-contained binary via llama.cpp / GGUF (like whisper.cpp), with built-in VAD. Prebuilt binaries + one-command model download → SenseVoiceSmall-GGUF · runtime · guide Introduction github repo : https://github.com/FunAudioLLM/SenseVoice SenseVoice is a speech foundation model with multiple speech understanding capabilities, including automatic speech recognition (ASR), spoken language identification (LID), speech emotion recognition (SER), and audio event detection (AED). [//]: # (<div align="center"<img src="image/sensevoice.png" width="700"/ </div) |<a href="#What's News" What's News </a |<a href="#Benchmarks" Benchmarks </a |<a href="#Install" Install </a |<a href="#Usage" Usage </a |<a href="#Community" Community </a Model Zoo: modelscope, huggingface Online Demo: modelscope demo, huggingface space Highlights 🎯 SenseVoice focuses on high-accuracy multilingual speech recognition, speech emotion recognition, and audio event detection. - Multilingual Speech Recognition: Trained with over 400,000 hours of data, supporting more than 50 languages, the recognition performance surpasses that of the Whisper model. - Rich transcribe: - Possess excellent emotion recognition capabilities, achieving and surpassing the effectiveness of the current best emotion recognition models on test data. - Offer sound event detection capabilities, supporting the detection of various common human-computer interaction events such as bgm, applause, laughter, crying, coughing, and sneezing. - Efficient Inference: The SenseVoice-Small model utilizes a non-autoregressive end-to-end framework, leading to exceptionally low inference latency. It requires only 70ms to process 10 seconds of audio, which is 15 times faster than Whisper-Large. - Convenient Finetuning: Provide convenient finetuni

Summarised from the published model card. Read the full card on the HuggingFace links below.

Specifications

What it is

MakerFunAudioLLM
TypeAudio & music
Variants1
Runs withfunasr
Released2024-07-03
Popularity33k downloads / month
Likes466
LicenceCommercial licence needed

How it works

How audio & music work

Audio or textinputAudio modelrecognise / synthesiseText or audiooutputSpeech-to-text turns audio into text; text-to-speech and music models turn text into audio.

Variants

Sizes & precisions

Open weights ship in several sizes and precisions. One page, all the variants — pick the one that fits your GPU. VRAM figures are estimates from model size.

VariantParamsPrecisionVRAMFits 16 GBWeights
SenseVoiceSmallBF16Weights ↗

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys sensevoicesmall for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (sensevoicesmall below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -F model="sensevoicesmall" -F file=@audio.mp3

Details

Languages, data & research

Languages

zh en ja ko yue multilingual

Tags

funasr speech-recognition asr emotion-recognition audio-event-detection multilingual non-autoregressive whisper-alternative real-time voice-ai automatic-speech-recognition zh en ja

Licence

Commercial licence needed

The weights are open but its licence needs a commercial agreement for business use. AxForge can arrange that licence and host the model for you — you pay AxForge, we settle with the model’s maker. Ask us for a quote. Read the licence ↗

Sources

Weights & code

Want SenseVoiceSmall on EU-owned hardware?

Request a licence + hosting quote See what’s served now

Explore

More audio & music

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms