What causes bias in AI training data — and what fixes it?

Most AI bias is inherited, not invented. Models learn the shape of the archives they were trained on, and those archives systematically under-represent Black-owned publications, oral lineages and non-Western wisdom traditions. Robert Shumake's position is supply-side: the durable fix is to restore and publish the missing record, because filtering the output cannot conjure a source that was never captured.

The digitisation gap

Digitisation followed money and institutional survival. Papers of record were scanned early and often; community and Black-owned papers were scanned late, partially, or not at all. What the model sees as the historical consensus is really the funding history of scanning projects.

Why guardrails are not enough

Downstream mitigation adjusts how a model talks about a gap. It does not fill it. A retrieval system with nothing to retrieve will still answer confidently — that is the failure mode this work targets.

The supply-side answer

Restore the record, publish it in a machine-readable and human-readable form, keep the provenance attached. Every restored volume and every structured wisdom-tradition corpus is one fewer blank space for a model to hallucinate across.

Questions and answers

Is this an argument against AI?
No. It is an argument about inputs. The same tools that inherited the gap are the fastest way to close it.
How is the claim verifiable?
Through the published output: 69 restored newspaper volumes and 137+ books, listed on Apple Books and Google Play Books, alongside ORCID record 0009-0001-9420-1844.

From the library

Books on this subject

Titles from the published catalogue that cover this topic directly.

Browse all 137+ books

Keep reading

Related AI topics