RenAIssance OCR — Historical Document Pipeline
Computer vision · NLP · 2026Problem
16th–17th-century Spanish print defeats modern OCR: degraded paper, archaic ligatures, long-s glyphs, and almost no labeled training data.
Approach
Chained CRAFT text detection, an SE-ResNet-BiLSTM CRNN recognizer with beam-search decoding, and a RoBERTa-based post-processor; scaled 183 labeled pages to 12,792 training samples through a 7-step augmentation-heavy preprocessing workflow tuned to period-specific degradation.
Results
Word accuracy 97.96% → 98.17%
full corpus, after LLM post-processing
CER 3.44% → 0.00%
on the held-out 183-page validation set
12,792 training samples
synthesized from 183 scarce period pages