Model reference · open weights
align is an open-weight audio or speech model from desert-ant-labs. align (BF16) weighs 0 MB; the smallest configuration that runs it is RTX 3060 12 GB.
What it is
| Released by | desert-ant-labs |
|---|---|
| Type | Audio & music |
| Task | Speech→text |
| Runs with | litert |
| Released | 2026-08-13 |
| Popularity | 546 downloads / month |
| Weights | 0 MB (align (BF16), file size) |
| Licence | Its own licence terms |
What it runs on
Weights 0 MB (file size) · overhead about 1.6 GB.
| Card | One stream | Counted memory |
|---|---|---|
| RTX 3060 12 GB | fits | 11.6 GB |
| RTX 4060 Ti 16 GB | fits | 15.4 GB |
| RTX 3090 24 GB | fits | 23.4 GB |
| RTX 4090 24 GB | fits | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
Estimates, not measurements: the weights are the build's file size. A speech model's decoder keeps a small cache for every stream it transcribes, so memory grows with the streams and beams at once. Counted memory is 92 % of what CUDA reports for the card.
From the model card
Accurate word timestamps for any transcript.
Word-timestamp refinement for any transcript, on device.
Corrects the word-level timings that Apple's SpeechTranscriber and SpeechAnalyzer
return, without replacing them. Align observes the same audio the analyzer already
receives, runs a small Core ML cascade on the CPU and Neural Engine, and returns the
familiar result surface with tightened audioTimeRange values. The models are tiny
(560KB compiled Core ML) and refine a typical result in a few milliseconds
on device.
Apple:
"world"2.61-3.04s ➜ Align:"world"2.57-2.98s
| Platforms | iOS, macOS, tvOS, visionOS, Linux, Windows, Node |
| Languages | 9 |
| Weights | v1.1.0 |
Swift (requirements)
.package(url: "https://github.com/Desert-Ant-Labs/desert-ant-core.git", from: "3.5.0")
Then add the Align product to your target.
JavaScript (requirements)
npm i @desert-ant-labs/align
| File | Format | Size | Contents |
|---|---|---|---|
align-coarse.mlmodelc | Compiled Core ML (FP16) | 285KB | Coarse stage |
align-fine.mlmodelc | Compiled Core ML (FP16) | 274KB | Fine stage |
align-coarse.tflite | LiteRT (FP32) | ~0.5MB | Linux, Windows and Node via the SDK |
align-fine.tflite | LiteRT (FP32) | ~0.5MB | Linux, Windows and Node via the SDK |
mel_filters.bin | Float32 filter bank | 40KB | Log-mel filter bank the runtime frontend needs |
calibrator.bin | Gradient-boosted trees | 70KB | Correction calibrator |
refiner_config.json | JSON | tiny | Runtime config |
LICENSE.md | Text | tiny | Full license terms |
The compiled .mlmodelc stages (Apple platforms) or .tflite stages (Linux, Windows and Node),
plus mel_filters.bin, calibrator.bin, and refiner_config.json, are exactly what the SDK
downloads.
See the LiteRT usage guide
for how the SDK loads the .tflite pair on Linux, Windows and Node.
Measured on v1.0.0 over held-out recordings.
| Condition | Raw error | Align error | Reduction |
|---|---|---|---|
| Clean | 124.2ms | 43.9ms | 65% |
| Noisy | 88.3ms | 33.4ms | 62% |
Raw is each proposer's own timing before refinement; the English cell pools Whisper, Parakeet and Apple proposals.
1.1.0 corrects the Apple runtime's log-mel scaling, so on-device results are re-measured for this release; the numbers above are the training-side measurement.
Macro-averaged over the nine languages, so a language with more test data cannot carry the figure on its own.
A 500-clip sample of each official LibriSpeech test-clean and test-other split.
| Engine | Split | Raw | Refined | Reduction | Within 50ms |
|---|---|---|---|---|---|
| Apple SpeechAnalyzer | test-clean | 106.4ms | 20.2ms | 81% | 37% to 95% |
| Apple SpeechAnalyzer | test-other | 111.6ms | 24.8ms | 78% | 35% to 92% |
For editing work the p90 matters more than the mean. Large errors are what a viewer notices when a caption slips or a clip cuts mid-word.
| Split | Raw p90 | Refined p90 |
|---|---|---|
| test-clean | 230.7ms | 33.0ms, about one frame of 30fps video |
258 word boundaries across 10 recordings, corrected by hand against the waveform rather than by another aligner. The only figure here not measured against machine references.
| System | Error | Within 50ms |
|---|---|---|
| Raw Whisper | 100.8ms | 43% |
| WhisperX | 53.5ms | 67% |
| Align | 45.0ms | 76% |
English, Spanish, French, Italian, Portuguese, German, Japanese, Korean, and Chinese. A locale outside this set is passed through unchanged.
Desert Ant Labs Source-Available License. Free for most apps, and a commercial license is required at scale. Full terms are at the link. Licensing: .
@software{align_2026,
title = {Align: Word-timestamp refinement for any transcript, on device},
author = {Desert Ant Labs},
year = {2026},
url = {https://huggingface.co/desert-ant-labs/align},
}
© 2026 Desert Ant Labs ·
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.