Model reference · open weights
hayai-ocr-nova is an open-weight language model from JustANormalTinkerer. hayai-ocr-v2.5-nova (FP32) weighs 314 MB; the smallest configuration that runs it is RTX 3060 12 GB.
What it is
| Released by | JustANormalTinkerer |
|---|---|
| Type | Language models |
| Task | Image→text |
| Parameters (lead) | 157M |
| Runs with | transformers |
| Released | 2026-09-20 |
| Popularity | 594 downloads / month |
| Weights | 314 MB (hayai-ocr-v2.5-nova (FP32), file size) |
| Licence | Open weights |
What it runs on
Weights 314 MB (file size) · runtime overhead from 471 MB on a small card.
How much memory each request adds is not estimated yet for this architecture — only the weights are. They need the cards below at the least, plus room for the context.
| Card | The weights alone |
|---|---|
| RTX 3060 12 GB | fits |
| RTX 4060 Ti 16 GB | fits |
| RTX 3090 24 GB | fits |
| RTX 4090 24 GB | fits |
| RTX 5090 32 GB | fits |
| L40S 48 GB | fits |
| A100 80 GB | fits |
| H100 80 GB | fits |
| RTX PRO 6000 Blackwell 96 GB | fits |
| DGX Spark (GB10) 128 GB unified | fits |
| H200 141 GB | fits |
| B200 180 GB | fits |
From the model card
Hayai OCR is an ultra-lightweight (~150M parameter) vision-to-text model engineered for ultra-fast, crop-level transcription across Japanese, Chinese, Korean, and English.
Hayai couples Google's SigLIP2 NaFlex vision encoder with a custom 12-layer causal transformer decoder. It transcribes dense, stylized, horizontal, and vertical text in a single forward pass without requiring an intermediate text-line detection stage (e.g., DBNet/YOLO).
Note: Hayai is designed specifically for crop-level recognition. For full-page scanning, pair it with a text detector.
Hayai v2.5 retains the core backbone of v2.1 (SigLIP2 NaFlex ~86M + 12-layer GQA decoder) while introducing major structural and efficiency upgrades:
DSCProjector) that reshapes patch embeddings into a 2D grid and applies a 2x pixel unshuffle (4 patches -> 1 token).attn_res_scale, ffn_res_scale, initialized at 1.0) to stabilize deep residual propagation.generate() features static KV-cache allocation, precomputed 1D text RoPE, and FP16 autocasting on CUDA devices.| Component | v2.1 | v2.5 Nova |
|---|---|---|
| Vision Projector | 2-layer MLP (1 patch -> 1 token) | DSCProjector: Reshape -> Replicate Pad -> Pixel Unshuffle (4 -> 1) -> LayerNorm -> MLP -> RMSNorm |
| Decoder Vision Tokens | N patches | N / 4 tokens (2D mRoPE runs on compressed spatial grid) |
| Residual Connections | Standard additive: x + f(x) | Learnable per-channel scale: x + scale * f(x) |
| Auxiliary Supervision | None | 226-class IDS Head (training-only for CJK grounding) |
| Inference Path | Standard dynamic greedy loop | Preallocated static KV-cache + precomputed RoPE caches |
| Decoder Architecture | 12 layers, d_model=512, d_ffn=2048, 8 query / 2 KV heads, SwiGLU, QK RMSNorm | Identical core parameters |
import torch
from PIL import Image
from transformers import AutoModel, AutoProcessor, PreTrainedTokenizerFast
MODEL_ID = "JustANormalTinkerer/hayai-ocr-v2.5-nova"
model = AutoModel.from_pretrained(MODEL_ID, trust_remote_code=True).cuda().eval()
tokenizer = PreTrainedTokenizerFast.from_pretrained(MODEL_ID)
processor = AutoProcessor.from_pretrained("google/siglip2-base-patch16-naflex")
image = Image.open("example.png").convert("RGB")
# Select patch budget: 256 (throughput), 384 (balanced), or 512 (quality)
inputs = processor(images=[image], max_num_patches=384, return_tensors="pt").to("cuda")
with torch.no_grad():
texts = model.generate(
pixel_values=inputs["pixel_values"],
pixel_attention_mask=inputs["pixel_attention_mask"],
spatial_shapes=inputs["spatial_shapes"],
tokenizer=tokenizer,
max_new_tokens=128,
repetition_penalty=1.0, # Keep at 1.0 (disabled) for OCR
)
print(texts[0])
Requirements:
trust_remote_code=Trueis required for custom block-causal attention and 2D multi-dimensional RoPE (mRoPE). Standard greedy decoding is recommended. For helper utilities, see hayai-ocr on GitHub.
max_num_patches)Because SigLIP2 NaFlex scales inputs dynamically preserving aspect ratio, max_num_patches governs your spatial budget. Thanks to the 4x patch compression in v2.5, higher patch counts run significantly faster than in earlier architectures:
max_num_patches | Vision Tokens (Decoder) | CER ↓ | Exact Match ↑ | Text-only CER ↓ | Text-only EM ↑ | Relative Latency* |
|---|---|---|---|---|---|---|
| 256 | <= 64 | 4.95% | 75.35% | 3.54% | 82.65% | Baseline (1.00x) |
| 384 (Default) | <= 96 | 3.65% | 79.15% | 2.36% | 86.31% | ~1.19x |
| 512 (Max Quality) | <= 128 | 3.10% | 80.68% | 1.78% | 88.04% | ~1.34x |
*Evaluated on JMangaBench_Mixed (3,286 crops) under NFKC normalization. Latencies measured on an NVIDIA T4 GPU.
256 — Throughput-Oriented: Best for high-volume pipelines, wide-aspect crops, and clean horizontal text.384 — Throughput-Oriented but still want quality: Excellent balance of speed and recognition fidelity. Closes over 70% of the accuracy gap to 512.512 — Fine-Grained / Small Glyphs (Recommended): Critical for dense panels, small stylized fonts, and complex vertical layouts (e.g., dense slices drop CER from 7.43% down to 2.70%).Evaluated across 3,286 standard benchmark crops under Unicode NFKC normalization:
| Model | Parameters | CER ↓ | Exact Match ↑ | Text-only CER ↓ | Text-only EM ↑ |
|---|---|---|---|---|---|
| MangaOCR | ~150M | 4.68% | 73.52% | 2.70% | 82.87% |
| BaberuOCR | ~150M | 4.59% | 72.25% | 2.60% | 81.65% |
| PaddleOCR-VL-For-Manga | ~900M | 2.91% | 78.91% | 1.87% | 84.66% |
| HayaiOCR v2.1 | ~150M | 3.23% | 79.67% | 1.90% | 87.46% |
| HayaiOCR v2.5 Nova (384) | ~150M | 3.65% | 79.15% | 2.36% | 86.31% |
| HayaiOCR v2.5 Nova (512) | ~150M | 3.10% | 80.68% | 1.78% | 88.04% |
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.