Good Papers

GT-Free OCR Metrics: A Reference-Free Evaluation Framework for Document OCR Systems

A render-and-compare framework benchmarks 147 reference-free visual metrics for document OCR, with the best composite achieving ρ = 0.494 correlation to ground-truth quality. Silent region misclassification artificially preserves visual similarity, suppressing correlation when masking is omitted.

Kshitij Singh

Published 2026Paris Poster Session 1 · Wed, Dec 9, 12:30 PM–2:30 PM local time · Paris Poster HallOpenReview ↗

89%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel16/20reviewers recommend it
lenient 5/5
medium 8/10
strict 3/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Evaluating document OCR systems at scale requires ground-truth text annotations that are expensive to obtain, preventing continuous quality monitoring across diverse document corpora. We present a reference-free evaluation framework based on render-and-compare: OCR output is rendered back into a page image and the result is compared visually against the masked original using no-reference or learned metrics. To our knowledge, no prior work has systematically benchmarked reference-free visual metrics as proxies for whole-document OCR quality at scale. We conduct the first such study, evaluating 147 method implementations spanning 15 mechanism classes on 1,355 OmniDocBench pages across 5 OCR-output variants (text-only, formula, table, combined, combined without masking). We measure every method by Spearman rank correlation against reference-based ground truth (edit distance for text, formula edit distance for formulae, TEDS for tables), quantifying whether a reference-free metric correctly orders pages by OCR quality without any annotations. A na ̈ıve CLIP-cosine full-page similarity baseline achieves ρ = 0.339; our best composite method reaches ρ = 0.494 (mean across 5 variants, p < 0.001, N=1,355), with a per-variant peak of ρ = 0.605 on the formula-only variant. We further show that when an OCR system silently misclassifies a region (e.g. labels a formula as text), both the masked original and the rendered reconstruction omit it consistently, keeping visual similarity artificially high even when ground-truth-based metrics correctly penalise the page; this effect suppresses correlation specifically on the combined-without-masking variant. To support reproducibility we release three HuggingFace datasets comprising rendered page pairs, per-element logprobs, and document-similarity triplets, together with a trained DocSim LoRA similarity head.