Model reference · open weights
TeleOCR is an open-weight language model from StarDoc-AI. TeleOCR (BF16) weighs 2.8 GB; the smallest configuration that runs it is RTX 3060 12 GB.
TeleOCR is a 1.4B parameter open-source Vision-Language Model developed by StarDoc-AI for image-text-to-text document parsing. It supports Chinese, English, and Japanese with a context length of 128,000 tokens and is released under the Apache 2.0 license. The model unifies the parsing of digital and camera-captured documents within a single framework.
Summary of the StarDoc-AI/TeleOCR model card, 2026-10-01
What it is
| Released by | StarDoc-AI |
|---|---|
| Type | Language models |
| Task | Vision + text |
| Parameters (lead) | 1.4B |
| Context | 128,000 tokens |
| Runs with | transformers |
| Released | 2026-08-14 |
| Popularity | 16k downloads / month |
| Weights | 2.8 GB (TeleOCR (BF16), file size) |
| Licence | Open weights |
What it runs on
Weights 2.8 GB (file size) · KV cache 115 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 1.9 GB on a small card · context up to 128,000 tokens.
| Card | Requests at once 8K tokens each | Requests at once 32K tokens each | Longest single request | Counted memory |
|---|---|---|---|---|
| RTX 3060 12 GB | 7 | 1 | 58K | 11.6 GB |
| RTX 4060 Ti 16 GB | 11 | 2 | 90K | 15.4 GB |
| RTX 3090 24 GB | 19 | 4 | all 125K | 23.4 GB |
| RTX 4090 24 GB | 19 | 4 | all 125K | 23.4 GB |
| RTX 5090 32 GB | 27 | 6 | all 125K | 31.0 GB |
| L40S 48 GB | 41 | 10 | all 125K | 44.0 GB |
| A100 80 GB | 78 | 19 | all 125K | 78.2 GB |
| H100 80 GB | 73 | 18 | all 125K | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | 90 | 22 | all 125K | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | 104 | 26 | all 125K | 107 GB |
| H200 141 GB | 137 | 34 | all 125K | 138 GB |
| B200 180 GB | 178 | 44 | all 125K | 176 GB |
| 2× RTX 3060 12 GB tensor parallel | 17 | 4 | all 125K | 11.6 GB a card |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 5.7 GB | 8.5 GB |
| 5 | 9.4 GB | 23.5 GB |
| 8 | 12.3 GB | 34.8 GB |
| 16 | 19.8 GB | 64.9 GB |
| 32 | 34.8 GB | 125 GB |
| 64 | 64.9 GB | 245 GB |
On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.
Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (grouped-query attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.
From the model card
TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
TeleOCR is a lightweight (~1.2B parameters), open-source Vision-Language Model designed specifically for document parsing.
Unlike existing methods that mainly target either digital documents or camera-captured documents, TeleOCR unifies both scenarios within a single framework.
Compared with previous document parsing models, TeleOCR introduces
These techniques enable TeleOCR to achieve state-of-the-art performance on both digital and camera-captured document benchmarks while remaining lightweight enough for practical deployment.
TeleOCR achieves state-of-the-art performance on multiple public document parsing benchmarks.
To evaluate the model's ability to understand complex document deformations, we conduct a visual evaluation on the public dewarping datasets DocUNet and DIR300, with representative results shown in Figure. TeleOCR directly performs layout and content parsing on distorted documents without dewarping preprocessing or a dedicated rectification model, demonstrating robust parsing under complex geometric deformations.
| 模型 | overall ↑ | Text edit ↓ | formula cdm ↑ | Table teds ↑ | order edit ↓ |
|---|---|---|---|---|---|
| Specialized VLMs | |||||
| TeleOCR | 67.96 | 0.1903 | 0.02 | 64.97 | 0.398 |
| Mineru 2.5 pro | 62.26 | 0.3402 | 0.04 | 67.75 | 0.356 |
| OvisOCR2 | 59.25 | 0.3883 | 0.00 | 61.59 | 0.3791 |
| PaddleOCRvl 1.6 | 55.11 | 0.4364 | 0.21 | 51.34 | 0.412 |
| Model Type | Methods | Param | Overall ↑ | Text Edit ↓ | Formula CDM ↑ | Table TEDS ↑ | Table TEDS-S ↑ | Read Order Edit ↓ |
|---|---|---|---|---|---|---|---|---|
| Specialized VLMs | TeleOCR | 1.2B | 96.87 | 0.027 | 96.36 | 97.05 | 98.52 | 0.122 |
| OvisOCR2 | 0.8B | 96.58 | 0.025 | 97.53 | 94.76 | 97.16 | 0.111 | |
| PaddleOCR-VL-1.6 | 0.9B | 96.33 | 0.033 | 97.49 | 94.76 | 97.11 | 0.127 | |
| MinerU2.5-Pro | 1.2B | 95.75 | 0.036 | 97.45 | 93.42 | 95.92 | 0.120 | |
| GLM-OCR | 0.9B | 95.22 | 0.044 | 97.18 | 92.83 | 95.39 | 0.133 | |
| PaddleOCR-VL-1.5 | 0.9B | 94.87 | 0.038 | 96.69 | 91.67 | 94.37 | 0.130 | |
| HunyuanOCR-1.5 | 1B | 94.74 | 0.033 | 97.49 | 94.76 | 97.11 | 0.127 | |
| PaddleOCR-VL | 0.9B | 94.11 | 0.040 | 95.70 | 90.65 | 93.74 | 0.135 | |
| Youtu-Parsing | 2.5B | 93.68 | 0.044 | 93.45 | 92.02 | 95.00 | 0.116 | |
| Logics-Parsing-v2 | 4B | 93.27 | 0.041 | 95.47 | 88.42 | 91.98 | 0.137 | |
| FireRed-OCR | 2B | 93.20 | 0.037 | 95.27 | 88.04 | 91.06 | 0.131 | |
| MinerU2.5 | 1.2B | 92.98 | 0.045 | 95.59 | 87.88 | 91.47 | 0.130 | |
| OpenDoc-0.1B | 0.1B | 90.64 | 0.049 | 92.93 | 83.88 | 87.45 | 0.140 | |
| dots.ocr | 3B | 90.50 | 0.048 | 89.12 | 87.18 | 90.58 | 0.138 | |
| DeepSeek-OCR 2 | 3B | 90.17 | 0.050 | 91.59 | 83.89 | 87.75 | 0.144 | |
| HunyuanOCR | 1B | 89.87 | 0.089 | 87.44 | 91.01 | 93.23 | 0.171 | |
| Dolphin-v2 | 3B | 89.34 | 0.069 | 90.53 | 84.40 | 87.44 | 0.150 | |
| OCRVerse | 4B | 88.44 | 0.063 | 89.14 | 82.44 | 86.27 | 0.163 | |
| MonkeyOCR-pro-3B | 3B | 88.43 | 0.074 | 88.33 | 84.35 | 88.62 | 0.189 | |
| General VLMs | Ovis2.6-30B-A3B | 30B | 93.62 | 0.035 | 94.93 | 89.44 | 92.40 | 0.135 |
| Gemini 3 Pro | -- | 92.85 | 0.064 | 95.83 | 89.15 | 92.96 | 0.165 | |
| Gemini 3 Flash | -- | 92.58 | 0.066 | 95.03 | 89.29 | 93.51 | 0.173 | |
| Qwen3-VL-235B | 235B | 89.78 | 0.063 | 92.53 | 83.07 | 86.75 | 0.166 | |
| GPT-5.2 | -- | 86.52 | 0.114 | 88.00 | 82.95 | 87.93 | 0.193 | |
| InternVL3.5-241B | 241B | 83.61 | 0.130 | 89.52 | 74.35 | 79.78 | 0.215 |
| Model Type | Methods | Param | Overall ↑ | Text Edit ↓ | Formula CDM ↑ | Table TEDS ↑ | Table TEDS-S ↑ | Read Order Edit ↓ |
|---|---|---|---|---|---|---|---|---|
| Decoupled VLMs | TeleOCR | 1.2B | 88.53 | 0.1173 | 88.26 | 89.05 | 92.14 | 0.2011 |
| PaddleOCR-VL-1.6 | 0.9B | 87.36 | 0.1369 | 88.42 | 85.76 | 90.14 | 0.2057 | |
| MinerU2.5-Pro | 1.2B | 87.33 | 0.1362 | 90.15 | 85.46 | 90.12 | 0.2013 |
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.