davanstrien/ocr-bench-britannica-results
Rankings are computed using Bradley-Terry MLE from pairwise comparisons judged by a vision-language model. The judge sees the original document image alongside two anonymised OCR outputs and picks the more faithful transcription. Browse the comparisons to see the evidence — and vote yourself to build a Human ELO column. Human votes are stored locally for this session only and will reset when the server restarts.
| # | Model | Params | Judge ELO | 95% CI | Wins | Losses | Ties | Win% |
|---|---|---|---|---|---|---|---|---|
| 1 | dots.mocr | 3B | 1745 | 1714–1782 | 436 | 141 | 4 | 75% |
| 2 | LightOnOCR-2-1B | 1B | 1741 | 1709–1779 | 426 | 141 | 4 | 75% |
| 3 | GLM-OCR | 0.9B | 1738 | 1707–1773 | 469 | 157 | 2 | 75% |
| 4 | olmOCR-2-7B-1025-FP8 | — | 1719 | 1688–1753 | 454 | 167 | 4 | 73% |
| 5 | NuExtract3 | 4B | 1690 | 1657–1725 | 435 | 190 | 0 | 70% |
| 6 | Qianfan-OCR | 4.7B | 1568 | 1539–1600 | 347 | 279 | 2 | 55% |
| 7 | FireRed-OCR | 2.1B | 1557 | 1528–1588 | 338 | 286 | 4 | 54% |
| 8 | Unlimited-OCR | — | 1545 | 1518–1575 | 312 | 281 | 0 | 53% |
| 9 | PaddleOCR-VL-1.6 | 0.9B | 1463 | 1435–1493 | 267 | 353 | 1 | 43% |
| 10 | DeepSeek-OCR | 4B | 1452 | 1422–1482 | 260 | 363 | 4 | 41% |
| 11 | PP-OCRv6_medium | — | 1397 | 1367–1428 | 220 | 393 | 0 | 36% |
| 12 | DeepSeek-OCR-2 | — | 1380 | 1345–1415 | 211 | 409 | 5 | 34% |
| 13 | tesseract-5 | — | 1113 | 1062–1153 | 79 | 514 | 0 | 13% |
| 14 | dots.ocr | 1.7B | 891 | 797–963 | 24 | 604 | 0 | 4% |
Smaller models can win on the right documents. Error bars show 95% confidence intervals.