Model reference · open weights
kotoba-whisper-faster is an open-weight audio or speech model from kotoba-tech. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.
Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.
What it is
| Released by | kotoba-tech |
|---|---|
| Type | Audio & music |
| Task | Speech→text |
| Runs with | ctranslate2 |
| Released | 2024-09-17 |
| Popularity | 3k downloads / month |
| Licence | Open weights |
About
This repository contains the conversion of kotoba-tech/kotoba-whisper-v2.0 to the CTranslate2 model format.
This model can be used in CTranslate2 or projects based on CTranslate2 such as faster-whisper.
Install library and download sample audio.
pip install faster-whisper
wget https://huggingface.co/kotoba-tech/kotoba-whisper-v1.0-ggml/resolve/main/sample_ja_speech.wav
Inference with the kotoba-whisper-v2.0-faster.
from faster_whisper import WhisperModel
model = WhisperModel("kotoba-tech/kotoba-whisper-v2.0-faster")
segments, info = model.transcribe("sample_ja_speech.wav", language="ja", chunk_length=15, condition_on_previous_text=False)
for segment in segments:
print("[%.2fs -> %.2fs] %s" % (segment.start, segment.end, segment.text))
We measure the inference speed of different kotoba-whisper-v2.0 implementations with four different Japanese speech audio on MacBook Pro with the following spec:
| audio file | audio duration (min) | whisper.cpp (sec) | faster-whisper (sec) | hf pipeline (sec) |
|---|---|---|---|---|
| audio 1 | 50.3 | 581 | 2601 | 807 |
| audio 2 | 5.6 | 41 | 73 | 61 |
| audio 3 | 4.9 | 30 | 141 | 54 |
| audio 4 | 5.6 | 35 | 126 | 69 |
Scripts to re-run the experiment can be found bellow:
Also, currently whisper.cpp and faster-whisper support the sequential long-form decoding, and only Huggingface pipeline supports the chunked long-form decoding, which we empirically found better than the sequnential long-form decoding.
The original model was converted with the following command:
ct2-transformers-converter --model kotoba-tech/kotoba-whisper-v2.0 --output_dir kotoba-whisper-v2.0-faster \
--copy_files tokenizer.json preprocessor_config.json --quantization float16
Note that the model weights are saved in FP16. This type can be changed when the model is loaded using the compute_type option in CTranslate2.
For more information about the kotoba-whisper-v2.0, refer to the original model card.
From the published model card. Full card on the HuggingFace links in the sidebar.
How it works
Using it via the API
Once AxForge deploys kotoba-whisper-faster for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (kotoba-whisper-faster below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \ -H "Authorization: Bearer $AXFORGE_API_KEY" \ -F model="kotoba-whisper-faster" -F file=@audio.mp3
Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.