Persian OCR Triple Threat 🔥 Benchmark

A same-test leaderboard for Persian OCR across 6,669 printed lines and 6,669 handwritten pages. Both frozen splits live in one public dataset and contribute exactly 50/50 to the overall rank.

WER · word errorCER · character error S³ · semantic importance🔑 Ess Err · critical words

weights substituted and dropped reference words by ShenavaSanj importance. Overall 🔥 is the equal-split macro-average of each split's weighted score: 20% WER, 20% CER, and 60% S³. Lower is better. 🔒 marks closed paid APIs.

🔥 Best Overall Triple Threat 🥇 Bina 0.1 Koochik · 9.8420% WER · 20% CER · 60% S³
Frozen benchmark6,669 + 6,669 printed · handwriting

Sortable leaderboard — ranked by Overall Triple Threat

Sortable leaderboard — ranked by Overall Triple Threat
1
35.6M
104.92
10.21
12.12
217.47
166.86
11.19
185.39
165.36
24.05
Paddle Inference · CTC
Bina · Surya
complete
PP-OCRv6_small_rec student trained with hard-label CTC and teacher-logit distillation, evaluated through the adaptive75 two-layer full-page pipeline. Params reports the 5,292,614-parameter deployable recognition architecture; the resumable training state is larger because it retains auxiliary training heads. All 13,338 rows scored with zero missing predictions.

Submit a model

Run the model on all rows of both benchmark splits and submit JSONL using IDs printed:0printed:6668 and handwriting:0handwriting:6668. Open a Community discussion with predictions, exact model/API revision, decoding configuration, and runtime details.