Model reference · open weights

align

Audio desert-ant-labs Speech→text 1 build Its own licence terms 546 dl/mo

align is an open-weight audio or speech model from desert-ant-labs. align (BF16) weighs 0 MB; the smallest configuration that runs it is RTX 3060 12 GB.

What it is

Released bydesert-ant-labs
TypeAudio & music
TaskSpeech→text
Runs withlitert
Released2026-08-13
Popularity546 downloads / month
Weights0 MB (align (BF16), file size)
LicenceIts own licence terms

What it runs on

Memory and cards for align (BF16)

Weights 0 MB (file size) · overhead about 1.6 GB.

CardOne streamCounted
memory
RTX 3060 12 GBfits11.6 GB
RTX 4060 Ti 16 GBfits15.4 GB
RTX 3090 24 GBfits23.4 GB
RTX 4090 24 GBfits23.4 GB
RTX 5090 32 GBfits31.0 GB
L40S 48 GBfits44.0 GB
A100 80 GBfits78.2 GB
H100 80 GBfits78.1 GB
RTX PRO 6000 Blackwell 96 GBfits93.8 GB
DGX Spark (GB10) 128 GB unifiedfits107 GB
H200 141 GBfits138 GB
B200 180 GBfits176 GB

Estimates, not measurements: the weights are the build's file size. A speech model's decoder keeps a small cache for every stream it transcribes, so memory grows with the streams and beams at once. Counted memory is 92 % of what CUDA reports for the card.

From the model card

What desert-ant-labs says about align

Accurate word timestamps for any transcript.

Word-timestamp refinement for any transcript, on device.

  • SDKs, install and examples: https://github.com/Desert-Ant-Labs/desert-ant-core/blob/main/docs/models/align.md

Corrects the word-level timings that Apple's SpeechTranscriber and SpeechAnalyzer return, without replacing them. Align observes the same audio the analyzer already receives, runs a small Core ML cascade on the CPU and Neural Engine, and returns the familiar result surface with tightened audioTimeRange values. The models are tiny (560KB compiled Core ML) and refine a typical result in a few milliseconds on device.

Apple: "world" 2.61-3.04s ➜ Align: "world" 2.57-2.98s

Read the full model card

Try it

PlatformsiOS, macOS, tvOS, visionOS, Linux, Windows, Node
Languages9
Weightsv1.1.0

Install

Swift (requirements)

.package(url: "https://github.com/Desert-Ant-Labs/desert-ant-core.git", from: "3.5.0")

Then add the Align product to your target.

JavaScript (requirements)

npm i @desert-ant-labs/align

Files

FileFormatSizeContents
align-coarse.mlmodelcCompiled Core ML (FP16)285KBCoarse stage
align-fine.mlmodelcCompiled Core ML (FP16)274KBFine stage
align-coarse.tfliteLiteRT (FP32)~0.5MBLinux, Windows and Node via the SDK
align-fine.tfliteLiteRT (FP32)~0.5MBLinux, Windows and Node via the SDK
mel_filters.binFloat32 filter bank40KBLog-mel filter bank the runtime frontend needs
calibrator.binGradient-boosted trees70KBCorrection calibrator
refiner_config.jsonJSONtinyRuntime config
LICENSE.mdTexttinyFull license terms

The compiled .mlmodelc stages (Apple platforms) or .tflite stages (Linux, Windows and Node), plus mel_filters.bin, calibrator.bin, and refiner_config.json, are exactly what the SDK downloads.

LiteRT

See the LiteRT usage guide for how the SDK loads the .tflite pair on Linux, Windows and Node.

Inputs and outputs

  • Input: mono audio plus Apple's recognized words with their proposed start/end times.
  • Output: the same words with corrected start/end times, or Apple's original time when a correction is not structurally safe.

Accuracy

Measured on v1.0.0 over held-out recordings.

All nine languages

ConditionRaw errorAlign errorReduction
Clean124.2ms43.9ms65%
Noisy88.3ms33.4ms62%

Raw is each proposer's own timing before refinement; the English cell pools Whisper, Parakeet and Apple proposals.

1.1.0 corrects the Apple runtime's log-mel scaling, so on-device results are re-measured for this release; the numbers above are the training-side measurement.

Macro-averaged over the nine languages, so a language with more test data cannot carry the figure on its own.

Public benchmark, English

A 500-clip sample of each official LibriSpeech test-clean and test-other split.

EngineSplitRawRefinedReductionWithin 50ms
Apple SpeechAnalyzertest-clean106.4ms20.2ms81%37% to 95%
Apple SpeechAnalyzertest-other111.6ms24.8ms78%35% to 92%

For editing work the p90 matters more than the mean. Large errors are what a viewer notices when a caption slips or a clip cuts mid-word.

SplitRaw p90Refined p90
test-clean230.7ms33.0ms, about one frame of 30fps video

Against hand-corrected boundaries

258 word boundaries across 10 recordings, corrected by hand against the waveform rather than by another aligner. The only figure here not measured against machine references.

SystemErrorWithin 50ms
Raw Whisper100.8ms43%
WhisperX53.5ms67%
Align45.0ms76%

Languages

English, Spanish, French, Italian, Portuguese, German, Japanese, Korean, and Chinese. A locale outside this set is passed through unchanged.

Limitations

  • References are machine forced-alignment estimates, not human annotations, so the figures show a large, consistent reduction of Apple's timing error rather than sample-accurate ground truth.
  • A learned correction is not guaranteed to improve every boundary; the structural fallback keeps Apple's timestamp when a correction looks unsafe but cannot catch every plausible-looking error.
  • Japanese, Korean, and Chinese were the weakest languages before v1.0.0. They now improve their proposals by 33%, 55%, and 51%.
  • Spoken numbers are the weakest remaining case. On a small sample, refinement moved digit boundaries further from the reference than leaving them alone, so treat them as unimproved until a larger sample settles it.
  • LiteRT is a third numeric path alongside Core ML, for Linux, Windows and Node. There is no browser build.

License

Desert Ant Labs Source-Available License. Free for most apps, and a commercial license is required at scale. Full terms are at the link. Licensing: .

See THIRD_PARTY_NOTICES.md.

Citation

@software{align_2026,
  title  = {Align: Word-timestamp refinement for any transcript, on device},
  author = {Desert Ant Labs},
  year   = {2026},
  url    = {https://huggingface.co/desert-ant-labs/align},
}

© 2026 Desert Ant Labs ·

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms