Model reference · open weights
dictabert-seg is an open-weight embedding model from dicta-il. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.
Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.
What it is
| Released by | dicta-il |
|---|---|
| Type | Embedding models |
| Task | Embeddings |
| Parameters (lead) | 185M |
| Context | 512 tokens |
| Runs with | transformers |
| Released | 2023-08-29 |
| Popularity | 4k downloads / month |
| Licence | Open weights |
About
State-of-the-art language model for Hebrew, released here.
This is the fine-tuned model for the prefix segmentation task.
For the bert-base models for other tasks, see here.
Sample usage:
from transformers import AutoModel, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained('dicta-il/dictabert-seg')
model = AutoModel.from_pretrained('dicta-il/dictabert-seg', trust_remote_code=True)
model.eval()
sentence = 'בשנת 1948 השלים אפרים קישון את לימודיו בפיסול מתכת ובתולדות האמנות והחל לפרסם מאמרים הומוריסטיים'
print(model.predict([sentence], tokenizer))
Output:
[
[
[ "[CLS]" ],
[ "ב","שנת" ],
[ "1948" ],
[ "השלים" ],
[ "אפרים" ],
[ "קישון" ],
[ "את" ],
[ "לימודיו" ],
[ "ב","פיסול" ],
[ "מתכת" ],
[ "וב","תולדות" ],
[ "ה","אמנות" ],
[ "ו","החל" ],
[ "לפרסם" ],
[ "מאמרים" ],
[ "הומוריסטיים" ],
[ "[SEP]" ]
]
]
If you use DictaBERT in your research, please cite DictaBERT: A State-of-the-Art BERT Suite for Modern Hebrew
BibTeX:
@misc{shmidman2023dictabert,
title={DictaBERT: A State-of-the-Art BERT Suite for Modern Hebrew},
author={Shaltiel Shmidman and Avi Shmidman and Moshe Koppel},
year={2023},
eprint={2308.16687},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
This work is licensed under a Creative Commons Attribution 4.0 International License.
From the published model card. Full card on the HuggingFace links in the sidebar.
Using it via the API
Once AxForge deploys dictabert-seg for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (dictabert-seg below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/embeddings \
-H "Authorization: Bearer $AXFORGE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"dictabert-seg","input":"text to embed"}'
Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.