Stage one — capture and segmentation
Scans are deskewed and segmented into columns, headlines, advertisements and images. Getting the geometry right first is what allows the text stages to keep their reading order instead of scrambling a broadsheet into nonsense.
Stage two — recognition and correction
OCR on century-old newsprint is unreliable in exactly the places that matter: proper nouns, place names, and the names of people the mainstream press never repeated. Correction passes are targeted there rather than spread evenly across the page.
Stage three — reconstruction and publication
Restored pages are rebuilt with their original typography and issued as readable volumes, each carrying its source attribution. The output is a book a person can read, not a database dump.
Questions and answers
- Is this fully automated?
- No. AI does the volume work; human judgement resolves ambiguity. Fully automated restoration reliably invents names, and in this material invented names are the whole failure.
- Which newspapers are covered?
- 69 titles spanning American cities, from The New York World to The Carolina Times to The Butte Miner, with particular attention to Black-owned publications.