Work / Fig. 04 · Computer vision · NLP · 2026

RenAIssance OCR — Historical Document Pipeline

Word accuracy 97.96% → 98.17% — full corpus, after LLM post-processing.

Fig. 04

RenAIssance OCR — Historical Document Pipeline

Computer vision · NLP · 2026

Problem

16th–17th-century Spanish print defeats modern OCR: degraded paper, archaic ligatures, long-s glyphs, and almost no labeled training data.

Approach

Chained CRAFT text detection, an SE-ResNet-BiLSTM CRNN recognizer with beam-search decoding, and a RoBERTa-based post-processor; scaled 183 labeled pages to 12,792 training samples through a 7-step augmentation-heavy preprocessing workflow tuned to period-specific degradation.

Results

  • Word accuracy 97.96% → 98.17%

    full corpus, after LLM post-processing

  • CER 3.44% → 0.00%

    on the held-out 183-page validation set

  • 12,792 training samples

    synthesized from 183 scarce period pages