Model reference · open weights

hayai-ocr-nova

LLMs JustANormalTinkerer · community Image→text 1 build Open weights 594 dl/mo

hayai-ocr-nova is an open-weight language model from JustANormalTinkerer. hayai-ocr-v2.5-nova (FP32) weighs 314 MB; the smallest configuration that runs it is RTX 3060 12 GB.

What it is

Released byJustANormalTinkerer
TypeLanguage models
TaskImage→text
Parameters (lead)157M
Runs withtransformers
Released2026-09-20
Popularity594 downloads / month
Weights314 MB (hayai-ocr-v2.5-nova (FP32), file size)
LicenceOpen weights

What it runs on

Memory and cards for hayai-ocr-v2.5-nova (FP32)

Weights 314 MB (file size) · runtime overhead from 471 MB on a small card.

How much memory each request adds is not estimated yet for this architecture — only the weights are. They need the cards below at the least, plus room for the context.

CardThe weights alone
RTX 3060 12 GBfits
RTX 4060 Ti 16 GBfits
RTX 3090 24 GBfits
RTX 4090 24 GBfits
RTX 5090 32 GBfits
L40S 48 GBfits
A100 80 GBfits
H100 80 GBfits
RTX PRO 6000 Blackwell 96 GBfits
DGX Spark (GB10) 128 GB unifiedfits
H200 141 GBfits
B200 180 GBfits

From the model card

What JustANormalTinkerer says about hayai-ocr-nova

Hayai OCR is an ultra-lightweight (~150M parameter) vision-to-text model engineered for ultra-fast, crop-level transcription across Japanese, Chinese, Korean, and English.

Hayai couples Google's SigLIP2 NaFlex vision encoder with a custom 12-layer causal transformer decoder. It transcribes dense, stylized, horizontal, and vertical text in a single forward pass without requiring an intermediate text-line detection stage (e.g., DBNet/YOLO).

Note: Hayai is designed specifically for crop-level recognition. For full-page scanning, pair it with a text detector.


Read the full model card

What’s New in v2.5

Hayai v2.5 retains the core backbone of v2.1 (SigLIP2 NaFlex ~86M + 12-layer GQA decoder) while introducing major structural and efficiency upgrades:

  • 4x Token Reduction (DSC Projector): Features a new Downsampling Spatial Convolution (DSCProjector) that reshapes patch embeddings into a 2D grid and applies a 2x pixel unshuffle (4 patches -> 1 token).
  • Lower Prefill Latency & KV Footprint: By reducing decoder vision tokens by 75%, v2.5 dramatically cuts prefill latency and memory footprint, making higher patch budgets computationally affordable.
  • Learnable Residual Scaling: Decoder layers now incorporate learnable per-channel scaling parameters (attn_res_scale, ffn_res_scale, initialized at 1.0) to stabilize deep residual propagation.
  • Engineered for High-Throughput Inference: generate() features static KV-cache allocation, precomputed 1D text RoPE, and FP16 autocasting on CUDA devices.
  • Auxiliary IDS Co-training: Pre-trained with an auxiliary 226-class Ideographic Description Sequence (IDS) classification head to sharpen character-level discrimination across rare CJK glyphs (training only; zero overhead at inference).

Architectural Comparison

Componentv2.1v2.5 Nova
Vision Projector2-layer MLP (1 patch -> 1 token)DSCProjector: Reshape -> Replicate Pad -> Pixel Unshuffle (4 -> 1) -> LayerNorm -> MLP -> RMSNorm
Decoder Vision TokensN patchesN / 4 tokens (2D mRoPE runs on compressed spatial grid)
Residual ConnectionsStandard additive: x + f(x)Learnable per-channel scale: x + scale * f(x)
Auxiliary SupervisionNone226-class IDS Head (training-only for CJK grounding)
Inference PathStandard dynamic greedy loopPreallocated static KV-cache + precomputed RoPE caches
Decoder Architecture12 layers, d_model=512, d_ffn=2048, 8 query / 2 KV heads, SwiGLU, QK RMSNormIdentical core parameters

Quickstart

import torch
from PIL import Image
from transformers import AutoModel, AutoProcessor, PreTrainedTokenizerFast

MODEL_ID = "JustANormalTinkerer/hayai-ocr-v2.5-nova"

model = AutoModel.from_pretrained(MODEL_ID, trust_remote_code=True).cuda().eval()
tokenizer = PreTrainedTokenizerFast.from_pretrained(MODEL_ID)
processor = AutoProcessor.from_pretrained("google/siglip2-base-patch16-naflex")

image = Image.open("example.png").convert("RGB")

# Select patch budget: 256 (throughput), 384 (balanced), or 512 (quality)
inputs = processor(images=[image], max_num_patches=384, return_tensors="pt").to("cuda")

with torch.no_grad():
    texts = model.generate(
        pixel_values=inputs["pixel_values"],
        pixel_attention_mask=inputs["pixel_attention_mask"],
        spatial_shapes=inputs["spatial_shapes"],
        tokenizer=tokenizer,
        max_new_tokens=128,
        repetition_penalty=1.0,  # Keep at 1.0 (disabled) for OCR
    )

print(texts[0])

Requirements: trust_remote_code=True is required for custom block-causal attention and 2D multi-dimensional RoPE (mRoPE). Standard greedy decoding is recommended. For helper utilities, see hayai-ocr on GitHub.


Resolution Strategy (max_num_patches)

Because SigLIP2 NaFlex scales inputs dynamically preserving aspect ratio, max_num_patches governs your spatial budget. Thanks to the 4x patch compression in v2.5, higher patch counts run significantly faster than in earlier architectures:

max_num_patchesVision Tokens (Decoder)CER ↓Exact Match ↑Text-only CER ↓Text-only EM ↑Relative Latency*
256<= 644.95%75.35%3.54%82.65%Baseline (1.00x)
384 (Default)<= 963.65%79.15%2.36%86.31%~1.19x
512 (Max Quality)<= 1283.10%80.68%1.78%88.04%~1.34x

*Evaluated on JMangaBench_Mixed (3,286 crops) under NFKC normalization. Latencies measured on an NVIDIA T4 GPU.

Which patch budget should you use?

  • 256 — Throughput-Oriented: Best for high-volume pipelines, wide-aspect crops, and clean horizontal text.
  • 384 — Throughput-Oriented but still want quality: Excellent balance of speed and recognition fidelity. Closes over 70% of the accuracy gap to 512.
  • 512 — Fine-Grained / Small Glyphs (Recommended): Critical for dense panels, small stylized fonts, and complex vertical layouts (e.g., dense slices drop CER from 7.43% down to 2.70%).

Benchmarks

Crop-Level Recognition (JMangaBench_Mixed)

Evaluated across 3,286 standard benchmark crops under Unicode NFKC normalization:

ModelParametersCER ↓Exact Match ↑Text-only CER ↓Text-only EM ↑
MangaOCR~150M4.68%73.52%2.70%82.87%
BaberuOCR~150M4.59%72.25%2.60%81.65%
PaddleOCR-VL-For-Manga~900M2.91%78.91%1.87%84.66%
HayaiOCR v2.1~150M3.23%79.67%1.90%87.46%
HayaiOCR v2.5 Nova (384)~150M3.65%79.15%2.36%86.31%
HayaiOCR v2.5 Nova (512)~150M3.10%80.68%1.78%88.04%

Block-Level E

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms