Model reference · open weights

CrisperWhisper2.0

Available as managed deployment Licence fee Audio nyralabs Speech→text 1 variants 7k dl/mo

CrisperWhisper2.0 is an open-weight audio or speech model from nyralabs. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

Released bynyralabs
TypeAudio & music
TaskSpeech→text
Parameters (lead)809M
Runs withcrisperwhisper
Released2026-07-15
Popularity7k downloads / month
LicenceCommercial licence needed

About

What CrisperWhisper2.0 is

The most accurate verbatim speech recognition you can run in production: controllable, multilingual, and timed to the word.

Release post · Paper · Full documentation · Models · Benchmark · Benchmark repo

Most speech-to-text systems never actually decide whether to write down what was said or what was meant. They inherit that choice from their training data and apply it inconsistently. CrisperWhisper 2.0 makes it an explicit, controllable choice. One recording, two transcripts:

Read the full model card

Verbatim, exactly what was said, in one consistent format: [um] so we we need to, to reschedule the th- thursday meeting to [uh] march third at nine thirty [laughter]

Intended, the clean version the speaker meant, with numbers, dates, and emails formatted the way you'd write them: So we need to reschedule the Thursday meeting to March 3 at 9:30.

On top of that:

  • Word-level timings. Around 30 ms mean boundary error on read speech and 41 ms on conversational speech, the most precise word timing of any system we benchmarked, on both.
  • Verbatimize. Upgrade transcripts you already have: given audio plus a trusted clean transcript, the model reproduces your content word-for-word and inserts only the disfluencies and vocal events actually present in the audio (rare-word recall jumps from 6.8% to 96.1% vs. re-transcribing). This turns the world's abundant clean corpora into verbatim ones, ready for TTS data, clinical speech analysis, and dataset construction.
  • Multilingual. Verbatim and intended modes work across most languages Whisper supports. CrisperWhisper 2.0 tops the Nyra Verbatim Speech Benchmark leaderboard for disfluency F1 across ten languages, ahead of every closed-source alternative we tested.
  • Seamless longform. Audio of any length, transcribed without the usual chunk-boundary artifacts: each window continues from the words already transcribed (conditional continuation), so there are no duplicated or dropped words at the seams and no fragile timestamp-token bookkeeping.
  • Production inference. A CTranslate2 runtime with speculative decoding and built-in mitigation of Whisper's looping-hallucination failure mode.

Performance

The Nyra Verbatim Speech Benchmark scores fillers, repetitions, cut-offs, and vocal sounds as separate, typed metrics. Its headline number is disfluency F1: how reliably a system writes down the disfluencies that were actually spoken, without inventing ones that weren't. Averaged over ten languages:

#SystemDisfluency F1
1CrisperWhisper 2.0 Pro93.5
2CrisperWhisper 2.087.8
3ElevenLabs Scribe v279.2
4Microsoft MAI-Transcribe-1.577.5
5CrisperWhisper 1.0*64.8
6Inworld STT59.5
7xAI Grok Speech-to-Text42.8
8Deepgram Nova-337.8
9Fish Audio ASR35.0
10AssemblyAI Universal-3 Pro30.5

two languages. English and German use human-labeled evaluation sets; the other eight languages use synthetic verbatim sets. Per-language breakdowns and how the metric is computed are in the benchmark post.

Word-timing accuracy

Mean absolute word-boundary error on read speech (TIMIT), lower is better:

#SystemBoundary error
1CrisperWhisper 2.029.6 ms
2xAI Grok Speech-to-Text37.1 ms
3CTC-seg49.3 ms
4ElevenLabs Scribe v251.3 ms
5NeMo-FA60.0 ms
6Deepgram Nova-363.3 ms
7WhisperX64.8 ms
8Cartesia Ink-Whisper69.4 ms
9Canary85.5 ms

are extracted from supervised cross-attention, plus results on conversational speech, are in the aligner post.

Install

# NVIDIA GPU (Linux): fastest, includes speculative decoding.
# An NVIDIA driver is all you need; CUDA libraries arrive via pip.
pip install "crisperwhisper[ct2]"

# Pure PyTorch: runs anywhere torch does (macOS, Windows, CPU)
pip install "crisperwhisper[transformers]"

Quickstart

from crisperwhisper import CrisperWhisperModel

model = CrisperWhisperModel()          # nyralabs/CrisperWhisper2.0_large
# or pick a size: CrisperWhisperModel("turbo")  # turbo / medium / small

# Verbatim transcription (default): every filler, repetition, stutter,
# false start, and vocal event
result = model.transcribe("meeting.wav", language="en")
print(result.text)

# Intended: the clean, readable version
clean = model.transcribe("meeting.wav", language="en", mode="intended")

# Word-level timestamps
result = model.transcribe("meeting.wav", language="en", word_timestamps=True)
for w in result.words:
    print(f"{w.start:6.2f}-{w.end:6.2f}  {w.word}")

# Verbatimize: upgrade an existing clean transcript with the
# disfluencies that are actually in the audio
result = model.verbatimize("clip.wav", "I think we should ship it Friday.")

Audio longer than 30 seconds is handled automatically (see longform below). The first load of a model downloads it from HuggingFace and, on the ct2 backend, converts it once into a local cache.

Models

ShorthandHuggingFace IDNotes
"large" (default)nyralabs/CrisperWhisper2.0_largeBest open quality
"turbo"`nyralabs/CrisperWhisp

From the published model card. Full card on the HuggingFace links in the sidebar.

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys crisperwhisper2-0 for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (crisperwhisper2-0 below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -F model="crisperwhisper2-0" -F file=@audio.mp3

Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms