Vol. IFrozen 2026-09-11Git 2dfc1b5*220,221 articles · 1900–1963

GenderNews Atlas

Sections

Data

What the corpus is, how it was sampled, what was checked, and how to rebuild everything.

Corpus

AmericanStories (Dell, M., Carlson, J., Bryan, T., Silcock, E., Arora, A., Shen, Z., D'Amico-Wong, L., Le, Q., Querubin, P., Heldring, L. (2023). American Stories: A Large-Scale Structured Text Dataset of Historical U.S. Newspapers. NeurIPS 2023 Datasets and Benchmarks Track.). Source collection: Chronicling America (Library of Congress, National Digital Newspaper Program). Source rights: Public domain (pre-1964 U.S. newspapers). Dataset licence: CC-BY-4.0.

AmericanStories is article-segmented OCR of digitised U.S. newspapers. It is not a representative sample of American news. Coverage depends on which papers libraries chose to digitise. The corpus thins sharply after 1922, when copyright begins to bind, and the Washington Evening Star dominates later years. Both facts are modelled, not ignored.

Sampling

For each of 22 years (1900–1963, every third year) the pipeline streams the year's archive and keeps the first 10,000 articles that pass length and legibility filters. The archive's member order is date-shuffled, so a streamed prefix should behave like a random sample. That was tested: against a complete download of 1963 (51,600 scans), the prefix's distribution over newspapers, months and page positions is indistinguishable from random samples of the same size.

Sampling audit (1963, full year)

DimensionObserved TVDNull medianNull 95th pct.p (one-sided)
lccn0.0480.0480.0610.48
month0.0430.0370.0530.23
page_band0.0200.0170.0350.35

Corpus by year

YearArticlesNewspapersEvening Star shareMedian wordsDictionary rateLegibleNear-duplicates
190010,0072702.1%11191.8%74.4%0.08%
190310,0222662.6%11192.3%77.1%0.06%
190610,0322873.9%11892.2%77.1%0.01%
190910,0142943.7%11892.2%78.5%0.01%
191210,0032493.6%11692.2%76.8%0.04%
191510,0172821.8%11691.7%73.7%0.03%
191810,0042902.7%11691.9%77.8%0.02%
192110,0022583.0%11991.8%72.1%0.01%
192410,00810617.9%11491.7%80.3%0.03%
192710,0035925.5%11892.4%87.2%0.03%
193010,0044423.0%11392.2%88.6%0.02%
193310,0085025.5%11892.5%85.0%0.02%
193610,0085825.6%11591.9%90.3%0.09%
193910,0085730.8%11492.0%86.3%0.05%
194210,0166725.0%11391.4%86.1%0.04%
194510,0027025.4%11990.7%83.2%0.05%
194810,0176545.9%12391.1%77.9%0.04%
195110,0175847.8%11691.1%83.9%0.09%
195410,0094849.7%12491.1%83.9%0.10%
195710,0094251.0%11890.8%89.2%0.05%
196010,0083758.5%11491.2%88.9%0.07%
196310,0033857.3%11991.5%94.1%0.04%

What is public and what is regenerated

The repository holds code, configuration, manifests with per-year content hashes, reference labels and derived results. The sampled article text (public domain, CC-BY-4.0 as packaged) is regenerated by make data rather than committed, to keep the repository small. Excerpts on this site are short and shown as evidence.