GT-Free OCR Metrics: A Reference-Free Evaluation Framework for Document OCR Systems
A render-and-compare framework benchmarks 147 reference-free visual metrics for document OCR, with the best composite achieving ρ = 0.494 correlation to ground-truth quality. Silent region misclassification artificially preserves visual similarity, suppressing correlation when masking is omitted.
Published 2026Paris Poster Session 1 · Wed, Dec 9, 12:30 PM–2:30 PM local time · Paris Poster HallOpenReview ↗
Only vote on papers you've read. Sign in with GitHub to vote.
Abstract
Evaluating document OCR systems at scale requires ground-truth text annotations that are expensive to obtain, preventing continuous quality monitoring across diverse document corpora. We present a reference-free evaluation framework based on render-and-compare: OCR output is rendered back into a page image and the result is compared visually against the masked original using no-reference or learned metrics. To our knowledge, no prior work has systematically benchmarked reference-free visual metrics as proxies for whole-document OCR quality at scale. We conduct the first such study, evaluating 147 method implementations spanning 15 mechanism classes on 1,355 OmniDocBench pages across 5 OCR-output variants (text-only, formula, table, combined, combined without masking). We measure every method by Spearman rank correlation against reference-based ground truth (edit distance for text, formula edit distance for formulae, TEDS for tables), quantifying whether a reference-free metric correctly orders pages by OCR quality without any annotations. A na ̈ıve CLIP-cosine full-page similarity baseline achieves ρ = 0.339; our best composite method reaches ρ = 0.494 (mean across 5 variants, p < 0.001, N=1,355), with a per-variant peak of ρ = 0.605 on the formula-only variant. We further show that when an OCR system silently misclassifies a region (e.g. labels a formula as text), both the masked original and the rendered reconstruction omit it consistently, keeping visual similarity artificially high even when ground-truth-based metrics correctly penalise the page; this effect suppresses correlation specifically on the combined-without-masking variant. To support reproducibility we release three HuggingFace datasets comprising rendered page pairs, per-element logprobs, and document-similarity triplets, together with a trained DocSim LoRA similarity head.