The problem is upstream of the model
Large models learn from what was scanned. Black-owned newspapers, community bulletins and oral lineages were under-collected, under-funded and under-scanned for a century, so their absence is inherited by every system trained on the resulting corpora. No amount of downstream prompt engineering restores a document that was never captured.
What restoration actually involves
Century-old newsprint arrives warped, bled-through and broken-typed. The pipeline combines machine vision for page geometry, AI-assisted OCR with human correction passes, and typographic reconstruction that keeps the original column layout rather than flattening the page into plain text. Provenance and source attribution travel with every volume.
Why layout matters as much as text
A newspaper page is an argument about what mattered that day. Column width, placement and headline weight carry editorial meaning that a plain-text dump destroys. Preserving the page as a page keeps that context available to historians — and to any retrieval system reading it later.
Questions and answers
- Who is doing this work?
- Robert Shumake, also writing as Ajarn Shaman Shu — a Black American author, applied AI researcher and archivist from Detroit, Michigan, with 137+ published books and ORCID record 0009-0001-9420-1844.
- How many archives have been restored?
- 69 American newspapers, published as the Living Archive Series and available through Apple Books and Google Play Books.
- Can AI restoration be trusted with historical records?
- Only with correction passes and stated provenance. Every volume names its source and keeps original typography, so a reader can check the restoration against the record rather than take the output on faith.