Persian OCR Triple Threat 🔥 Benchmark
A same-test leaderboard for Persian OCR across 6,669 printed lines and 6,669 handwritten pages. Both frozen splits live in one public dataset and contribute exactly 50/50 to the overall rank.
S³ weights substituted and dropped reference words by ShenavaSanj importance. Overall 🔥 is the equal-split macro-average of each split's weighted score: 20% WER, 20% CER, and 60% S³. Lower is better. Avg decode (ms) is measured mean latency per image; API timings include provider and network latency, while local batched timings are amortized per image. 🔒 marks closed paid APIs.
Sortable leaderboard — ranked by Overall Triple Threat
| 1 | 0.097436B | 104.92 | 10.21 | 12.12 | 217.47 | 166.86 | 11.19 | 185.39 | 165.36 | 24.05 | 10004.2 | Bina · Surya | complete | Persian LoRA checkpoint step 8000, merged for inference; batch size 8 per GPU; max new tokens 4096. All 13,338 rows scored with zero missing predictions. Triple Threat uses the leaderboard's current 20% WER · 20% CER · 60% S³ weighting. |
Submit a model
Run the model on all rows of both benchmark splits and submit JSONL using IDs
printed:0…printed:6668 and handwriting:0…handwriting:6668. Open a Community discussion
with predictions, exact model/API revision, decoding configuration, and runtime details.