Persian OCR Triple Threat 🔥 Benchmark

A same-test leaderboard for Persian OCR across 6,669 printed lines and 6,669 handwritten pages. Both frozen splits live in one public dataset and contribute exactly 50/50 to the overall rank.

WER · word errorCER · character error S³ · semantic importance🔑 Ess Err · critical words

weights substituted and dropped reference words by ShenavaSanj importance. Overall 🔥 is the equal-split macro-average of each split's weighted score: 20% WER, 20% CER, and 60% S³. Lower is better. Avg decode (ms) is measured mean latency per image; API timings include provider and network latency, while local batched timings are amortized per image. 🔒 marks closed paid APIs.

🔥 Best Overall Triple Threat 🥇 Bina 0.1 · 9.8420% WER · 20% CER · 60% S³
Frozen benchmark6,669 + 6,669 printed · handwriting

Sortable leaderboard — ranked by Overall Triple Threat

Sortable leaderboard — ranked by Overall Triple Threat
1
0.097436B
104.92
10.21
12.12
217.47
166.86
11.19
185.39
165.36
24.05
10004.2
Bina · Surya
complete
Persian LoRA checkpoint step 8000, merged for inference; batch size 8 per GPU; max new tokens 4096. All 13,338 rows scored with zero missing predictions. Triple Threat uses the leaderboard's current 20% WER · 20% CER · 60% S³ weighting.

Submit a model

Run the model on all rows of both benchmark splits and submit JSONL using IDs printed:0printed:6668 and handwriting:0handwriting:6668. Open a Community discussion with predictions, exact model/API revision, decoding configuration, and runtime details.