Model reference · open weights
clipclap is an open-weight embedding model from antflydb. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.
Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.
What it is
| Released by | antflydb |
|---|---|
| Type | Embedding models |
| Task | Embeddings |
| Runs with | onnxruntime |
| Released | 2026-02-05 |
| Popularity | 2k downloads / month |
| Licence | Open weights |
About
CLIPCLAP is a unified multimodal embedding model that maps text, images, and audio into a shared 512-dimensional vector space. It combines OpenAI's CLIP (text + image) with LAION's CLAP (audio) through a trained linear projection.
Built by antflydb for use with Antfly Inference, a standalone ML inference service for embeddings, chunking, reranking, and local model serving.
Text ──→ CLIP text encoder ──→ text_projection ──→ 512-dim (CLIP space)
Image ──→ CLIP visual encoder ──→ visual_projection ──→ 512-dim (CLIP space)
Audio ──→ CLAP audio encoder ──→ audio_projection ──→ 512-dim (CLIP space)
openai/clip-vit-base-patch32).laion/larger_clap_music_and_speech. The audio projection combines CLAP's native audio projection (1024→512) with a trained 512→512 linear layer that maps CLAP audio space into CLIP space.All three modalities produce 512-dimensional L2-normalized embeddings that are directly comparable via cosine similarity.
# Pull and run the model
antfly inference pull antflydb/clipclap:gguf:Q4_K
antfly inference run
# Embed text
curl -X POST http://localhost:8082/embed \
-H "Content-Type: application/json" \
-d '{
"model": "clipclap",
"input": [
{"type": "text", "text": "a cat sitting on a windowsill"},
{"type": "image_url", "image_url": {"url": "https://example.com/cat.jpg"}},
{"type": "audio_url", "audio_url": {"url": "https://example.com/cat-purring.wav"}}
]
}'
The audio projection layer bridges CLAP and CLIP embedding spaces. Training procedure:
The contrastive loss pushes matching audio-text pairs together while pushing non-matching pairs apart within each batch, preserving content discrimination.
| Parameter | Value |
|---|---|
| Training dataset | OpenSound/AudioCaps |
| Samples | 5000 audio-caption pairs |
| Epochs | 20 |
| Batch size | 256 |
| Learning rate | 1e-3 |
| Optimizer | Adam |
| Loss | Symmetric InfoNCE (temperature=0.07) |
| Train/val split | 90/10 |
| Component | Model |
|---|---|
| CLIP | openai/clip-vit-base-patch32 |
| CLAP | laion/larger_clap_music_and_speech |
| File | Description | Size |
|---|---|---|
text_model.onnx | CLIP text encoder | ~254 MB |
visual_model.onnx | CLIP visual encoder | ~330 MB |
text_projection.onnx | CLIP text projection (512→512) | ~4 KB |
visual_projection.onnx | CLIP visual projection (768→512) | ~6 KB |
audio_model.onnx | CLAP HTSAT audio encoder | ~590 MB |
audio_projection.onnx | Combined CLAP→CLIP projection (1024→512) | ~8 KB |
Additional files: clip_config.json, tokenizer.json, preprocessor_config.json, projection_training_metadata.json.
If you use CLIPCLAP, please cite the underlying models:
@inproceedings{radford2021clip,
title={Learning Transferable Visual Models From Natural Language Supervision},
author={Radford, Alec and Kim, Jong Wook and Hallacy, Chris and others},
booktitle={ICML},
year={2021}
}
@inproceedings{wu2023clap,
title={Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation},
author={Wu, Yusong and Chen, Ke and Zhang, Tianyu and others},
booktitle={ICASSP},
year={2023}
}
From the published model card. Full card on the HuggingFace links in the sidebar.
How it works
Using it via the API
Once AxForge deploys clipclap for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (clipclap below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/embeddings \
-H "Authorization: Bearer $AXFORGE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"clipclap","input":"text to embed"}'
Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.