Model reference · open weights

vit_base_patch16_224.dinov3-mlxim

Not for GPU servers Embeddings mlx-vision Image embed 1 build Its own licence terms 568 dl/mo

vit_base_patch16_224.dinov3-mlxim is an open-weight embedding model from mlx-vision. Built for Apple silicon (MLX) — it does not run on NVIDIA GPUs.

What it is

Released bymlx-vision
TypeEmbedding models
TaskImage embed
Parameters (lead)86M
Runs withmlx-image
Released2026-03-13
Popularity568 downloads / month
Weights171 MB (vit_base_patch16_224.dinov3-mlxim (MLX), file size)
LicenceIts own licence terms

What it runs on

Memory and cards for vit_base_patch16_224.dinov3-mlxim (MLX)

Weights 171 MB (file size) · overhead about 1.1 GB.

CardRunsCounted
memory
RTX 3060 12 GBfits11.6 GB
RTX 4060 Ti 16 GBfits15.4 GB
RTX 3090 24 GBfits23.4 GB
RTX 4090 24 GBfits23.4 GB
RTX 5090 32 GBfits31.0 GB
L40S 48 GBfits44.0 GB
A100 80 GBfits78.2 GB
H100 80 GBfits78.1 GB
RTX PRO 6000 Blackwell 96 GBfits93.8 GB
DGX Spark (GB10) 128 GB unifiedfits107 GB
H200 141 GBfits138 GB
B200 180 GBfits176 GB

Estimates, not measurements: the weights are the build's file size. No cache grows with use; a batch of inputs needs working memory of its own. Counted memory is 92 % of what CUDA reports for the card.

From the model card

What mlx-vision says about vit_base_patch16_224.dinov3-mlxim

A Vision Transformer feature extraction model trained on the LVD-1689M web dataset with DINOv3.

The model was trained in a self-supervised fashion. No classification head was trained, only the backbone. This is the ViT-B/16 variant (86M parameters), distilled from the DINOv3 ViT-7B teacher model.

Disclaimer: This is a porting of the Meta AI DINOv3 model weights to the Apple MLX Framework.

Read the full model card

How to use

pip install mlx-image

Here is how to use this model for feature extraction:

import mlx.core as mx
from mlxim.model import create_model
from mlxim.io import read_rgb
from mlxim.transform import ImageNetTransform

transform = ImageNetTransform(train=False, img_size=224)
x = transform(read_rgb("image.png"))
x = mx.expand_dims(x, 0)

model = create_model("vit_base_patch16_224.dinov3")
model.eval()

embeds = model(x, is_training=False)

Architecture

This model follows the ViT architecture with a patch size of 16. For a 224×224 image this results in 1 class token + 4 register tokens + 196 patch tokens = 201 tokens.

The model can accept larger images provided the image shapes are multiples of the patch size (16). If this condition is not met, the model will crop to the closest smaller multiple.

PropertyValue
Parameters86M
Patch size16
Embedding dim768
Depth12
Heads12
FFNMLP
Position encodingRoPE
Register tokens4

Available model variants (mlx-image)

Model nameParamsFFNIN-ReaLIN-RObj.Net
vit_small_patch16_224.dinov321MMLP87.060.450.9
vit_small_plus_patch16_224.dinov329MSwiGLU88.068.854.6
vit_base_patch16_224.dinov386MMLP89.376.764.1
vit_large_patch16_224.dinov3300MMLP90.288.174.8

Evaluation results

Results on global and dense tasks (LVD-1689M pretraining)

ModelIN-ReaLIN-RObj.NetOx.-HADE20kNYU↓DAVISNAVISPair
DINOv3 ViT-B/1689.376.764.158.551.80.37377.258.857.2

Refer to the DINOv3 paper for full evaluation details and protocols.

Training data

The model was distilled from DINOv3 ViT-7B, which was pretrained on LVD-1689M — a curated dataset of 1,689 million images from public web sources collected from Instagram.

Bias and limitations

DINOv3 delivers generally consistent performance across income categories on geographical fairness benchmarks, though a performance gap between low-income and high-income buckets remains. A relative difference is also observed between European and African regions. Fine-tuning may amplify these biases depending on the fine-tuning labels used.

Acknowledgements

Original model developed by Meta AI. See the blog post and paper. Weights ported to MLX by etornam45.

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

How it works

How embedding models work

Your textsentence / documentEncodermaps meaningVectorlist of numbersAn embedding model turns text into a vector, so similar meanings sit close together — the basis of search and RAG.

Running it

Where it runs

Built for Apple silicon (MLX) — it does not run on NVIDIA GPUs.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms