Leaderboard

davanstrien/ocr-bench-britannica-results

Rankings are computed using Bradley-Terry MLE from pairwise comparisons judged by a vision-language model. The judge sees the original document image alongside two anonymised OCR outputs and picks the more faithful transcription. Browse the comparisons to see the evidence — and vote yourself to build a Human ELO column. Human votes are stored locally for this session only and will reset when the server restarts.

# Model Params Judge ELO 95% CI Wins Losses Ties Win%
1 dots.mocr 3B 1745 1714–1782 436 141 4 75%
2 LightOnOCR-2-1B 1B 1741 1709–1779 426 141 4 75%
3 GLM-OCR 0.9B 1738 1707–1773 469 157 2 75%
4 olmOCR-2-7B-1025-FP8 1719 1688–1753 454 167 4 73%
5 NuExtract3 4B 1690 1657–1725 435 190 0 70%
6 Qianfan-OCR 4.7B 1568 1539–1600 347 279 2 55%
7 FireRed-OCR 2.1B 1557 1528–1588 338 286 4 54%
8 Unlimited-OCR 1545 1518–1575 312 281 0 53%
9 PaddleOCR-VL-1.6 0.9B 1463 1435–1493 267 353 1 43%
10 DeepSeek-OCR 4B 1452 1422–1482 260 363 4 41%
11 PP-OCRv6_medium 1397 1367–1428 220 393 0 36%
12 DeepSeek-OCR-2 1380 1345–1415 211 409 5 34%
13 tesseract-5 1113 1062–1153 79 514 0 13%
14 dots.ocr 1.7B 891 797–963 24 604 0 4%

ELO vs Parameter Count

Smaller models can win on the right documents. Error bars show 95% confidence intervals.