Model reference · open weights
react-native-executorch-distiluse-multilingual-cased is an open-weight embedding model from software-mansion. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.
Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.
What it is
| Released by | software-mansion |
|---|---|
| Type | Embedding models |
| Task | Embeddings |
| Runs with | executorch |
| Released | 2026-04-24 |
| Popularity | 785 downloads / month |
| Licence | Open weights |
About
This repository hosts the distiluse-base-multilingual-cased-v2 models exported for the
React Native ExecuTorch
library as ExecuTorch .pte programs, ready to run on device.
Upstream model: distiluse-base-multilingual-cased-v2
| Path | Backend | Precision |
|---|---|---|
coreml/distiluse_base_multilingual_cased_v2_coreml_fp16.pte | coreml | fp16 |
mlx/distiluse_base_multilingual_cased_v2_mlx_int8.pte | mlx | int8 |
vulkan/distiluse_base_multilingual_cased_v2_vulkan_fp16.pte | vulkan | fp16 |
xnnpack/distiluse_base_multilingual_cased_v2_xnnpack_fp32.pte | xnnpack | fp32 |
xnnpack/distiluse_base_multilingual_cased_v2_xnnpack_8da4w.pte | xnnpack | 8da4w |
config.json 58 B
coreml/config.json 1.0 kB
coreml/distiluse_base_multilingual_cased_v2_coreml_fp16.pte 258 MB
mlx/config.json 1.0 kB
mlx/distiluse_base_multilingual_cased_v2_mlx_int8.pte 133 MB
tokenizer.json 2.8 MB
tokenizer_config.json 531 B
vulkan/config.json 1.0 kB
vulkan/distiluse_base_multilingual_cased_v2_vulkan_fp16.pte 258 MB
xnnpack/config.json 1.7 kB
xnnpack/distiluse_base_multilingual_cased_v2_xnnpack_8da4w.pte 375 MB
xnnpack/distiluse_base_multilingual_cased_v2_xnnpack_fp32.pte 516 MB
These files are published for the ExecuTorch v1.4.1 runtime. ExecuTorch gives no forward compatibility guarantee, so an older runtime may fail to load them.
To use them in React Native ExecuTorch, pass the model constant shipped in the library's model registry to the corresponding task pipeline. See the documentation.
To load these files in your own ExecuTorch runtime, read the compatibility note first.
[CLS] / [SEP]).The exported program skips HuggingFace's internal attention-mask-to-4D conversion because the RNE runtime never pads at inference (single sentence, no batching). This preserves bit-exactness with the PyTorch reference (RMSE 0 on fp32 random input) while trimming ~27% off the XNNPACK forward wall-time and keeping XNNPACK delegation around 89–91% of graph runtime.
Unsupported combinations (rejected by the exporter, documented for reference):
model.to(torch.float16) causes softmax / LayerNorm overflow and the runtime output is NaN. XNNPACK's size wins come from quantization, not fp16.coremltools has no MIL mapping for the torch.int8 tensors torchao emits (KeyError: torch.int8). The CoreML-native way to shrink further is ct.optimize.coreml palette/linear quantization, not torchao source transforms.From the published model card. Full card on the HuggingFace links in the sidebar.
How it works
Using it via the API
Once AxForge deploys react-native-executorch-distiluse-multilingual-cased for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (react-native-executorch-distiluse-multilingual-cased below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/embeddings \
-H "Authorization: Bearer $AXFORGE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"react-native-executorch-distiluse-multilingual-cased","input":"text to embed"}'
Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.