Model reference · open weights
dictabert-morph is an open-weight embedding model from dicta-il. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.
Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.
What it is
| Released by | dicta-il |
|---|---|
| Type | Embedding models |
| Task | Embeddings |
| Parameters (lead) | 184M |
| Context | 512 tokens |
| Runs with | transformers |
| Released | 2023-08-29 |
| Popularity | 2k downloads / month |
| Licence | Open weights |
About
State-of-the-art language model for Hebrew, released here.
This is the fine-tuned model for the morphological tagging task.
For the bert-base models for other tasks, see here.
Sample usage:
from transformers import AutoModel, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained('dicta-il/dictabert-morph')
model = AutoModel.from_pretrained('dicta-il/dictabert-morph', trust_remote_code=True)
model.eval()
sentence = 'בשנת 1948 השלים אפרים קישון את לימודיו בפיסול מתכת ובתולדות האמנות והחל לפרסם מאמרים הומוריסטיים'
print(model.predict([sentence], tokenizer))
Output:
[{
"text": "בשנת 1948 השלים אפרים קישון את לימודיו בפיסול מתכת ובתולדות האמנות והחל לפרסם מאמרים הומוריסטיים",
"tokens": [{
"token": "בשנת",
"pos": "NOUN",
"feats": {
"Gender": "Fem",
"Number": "Sing"
},
"prefixes": ["ADP"],
"suffix": false
}, {
"token": "1948",
"pos": "NUM",
"feats": {},
"prefixes": [],
"suffix": false
}, {
"token": "השלים",
"pos": "VERB",
"feats": {
"Gender": "Masc",
"Number": "Sing",
"Person": "3",
"Tense": "Past"
},
"prefixes": [],
"suffix": false
}, {
"token": "אפרים",
"pos": "PROPN",
"feats": {},
"prefixes": [],
"suffix": false
}, {
"token": "קישון",
"pos": "PROPN",
"feats": {},
"prefixes": [],
"suffix": false
}, {
"token": "את",
"pos": "ADP",
"feats": {},
"prefixes": [],
"suffix": false
}, {
"token": "לימודיו",
"pos": "NOUN",
"feats": {
"Gender": "Masc",
"Number": "Plur"
},
"prefixes": [],
"suffix": "PRON",
"suffix_feats": {
"Gender": "Masc",
"Number": "Sing",
"Person": "3"
}
}, {
"token": "בפיסול",
"pos": "NOUN",
"feats": {
"Gender": "Masc",
"Number": "Sing"
},
"prefixes": ["ADP"],
"suffix": false
}, {
"token": "מתכת",
"pos": "NOUN",
"feats": {
"Gender": "Fem",
"Number": "Sing"
},
"prefixes": [],
"suffix": false
}, {
"token": "ובתולדות",
"pos": "NOUN",
"feats": {
"Gender": "Fem",
"Number": "Plur"
},
"prefixes": ["CCONJ", "ADP"],
"suffix": false
}, {
"token": "האמנות",
"pos": "NOUN",
"feats": {
"Gender": "Fem",
"Number": "Sing"
},
"prefixes": ["DET"],
"suffix": false
}, {
"token": "והחל",
"pos": "VERB",
"feats": {
"Gender": "Masc",
"Number": "Sing",
"Person": "3",
"Tense": "Past"
},
"prefixes": ["CCONJ"],
"suffix": false
}, {
"token": "לפרסם",
"pos": "VERB",
"feats": {},
"prefixes": [],
"suffix": false
}, {
"token": "מאמרים",
"pos": "NOUN",
"feats": {
"Gender": "Masc",
"Number": "Plur"
},
"prefixes": [],
"suffix": false
}, {
"token": "הומוריסטיים",
"pos": "ADJ",
"feats": {
"Gender": "Masc",
"Number": "Plur"
},
"prefixes": [],
"suffix": false
}]
}]
If you use DictaBERT in your research, please cite DictaBERT: A State-of-the-Art BERT Suite for Modern Hebrew
BibTeX:
@misc{shmidman2023dictabert,
title={DictaBERT: A State-of-the-Art BERT Suite for Modern Hebrew},
author={Shaltiel Shmidman and Avi Shmidman and Moshe Koppel},
year={2023},
eprint={2308.16687},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
This work is licensed under a Creative Commons Attribution 4.0 International License.
From the published model card. Full card on the HuggingFace links in the sidebar.
Using it via the API
Once AxForge deploys dictabert-morph for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (dictabert-morph below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/embeddings \
-H "Authorization: Bearer $AXFORGE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"dictabert-morph","input":"text to embed"}'
Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.