Leaderboard

davanstrien/ocr-bench-britannica-results

Rankings are computed using Bradley-Terry MLE from pairwise comparisons judged by a vision-language model. The judge sees the original document image alongside two anonymised OCR outputs and picks the more faithful transcription. Browse the comparisons to see the evidence — and vote yourself to build a Human ELO column. Human votes are stored locally for this session only and will reset when the server restarts.

# Model Params Judge ELO 95% CI Wins Losses Ties Win% Human ELO H-Win%
1 dots.mocr 3B 1745 1714–1782 436 141 4 75% 443 0
2 LightOnOCR-2-1B 1B 1741 1709–1779 426 141 4 75% — —
3 GLM-OCR 0.9B 1738 1707–1773 469 157 2 75% 2557 100
4 olmOCR-2-7B-1025-FP8 — 1719 1688–1753 454 167 4 73% — —
5 NuExtract3 4B 1690 1657–1725 435 190 0 70% — —
6 Qianfan-OCR 4.7B 1568 1539–1600 347 279 2 55% — —
7 FireRed-OCR 2.1B 1557 1528–1588 338 286 4 54% — —
8 Unlimited-OCR — 1545 1518–1575 312 281 0 53% — —
9 PaddleOCR-VL-1.6 0.9B 1463 1435–1493 267 353 1 43% — —
10 DeepSeek-OCR 4B 1452 1422–1482 260 363 4 41% — —
11 PP-OCRv6_medium — 1397 1367–1428 220 393 0 36% — —
12 DeepSeek-OCR-2 — 1380 1345–1415 211 409 5 34% — —
13 tesseract-5 — 1113 1062–1153 79 514 0 13% — —
14 dots.ocr 1.7B 891 797–963 24 604 0 4% — —

ELO vs Parameter Count

Smaller models can win on the right documents. Error bars show 95% confidence intervals.