Data
What the corpus is, how it was sampled, what was checked, and how to rebuild everything.
Corpus
AmericanStories (Dell, M., Carlson, J., Bryan, T., Silcock, E., Arora, A., Shen, Z., D'Amico-Wong, L., Le, Q., Querubin, P., Heldring, L. (2023). American Stories: A Large-Scale Structured Text Dataset of Historical U.S. Newspapers. NeurIPS 2023 Datasets and Benchmarks Track.). Source collection: Chronicling America (Library of Congress, National Digital Newspaper Program). Source rights: Public domain (pre-1964 U.S. newspapers). Dataset licence: CC-BY-4.0.
AmericanStories is article-segmented OCR of digitised U.S. newspapers. It is not a representative sample of American news. Coverage depends on which papers libraries chose to digitise. The corpus thins sharply after 1922, when copyright begins to bind, and the Washington Evening Star dominates later years. Both facts are modelled, not ignored.
Sampling
For each of 22 years (1900–1963, every third year) the pipeline streams the year's archive and keeps the first 10,000 articles that pass length and legibility filters. The archive's member order is date-shuffled, so a streamed prefix should behave like a random sample. That was tested: against a complete download of 1963 (51,600 scans), the prefix's distribution over newspapers, months and page positions is indistinguishable from random samples of the same size.
Sampling audit (1963, full year)
| Dimension | Observed TVD | Null median | Null 95th pct. | p (one-sided) |
|---|---|---|---|---|
| lccn | 0.048 | 0.048 | 0.061 | 0.48 |
| month | 0.043 | 0.037 | 0.053 | 0.23 |
| page_band | 0.020 | 0.017 | 0.035 | 0.35 |
Corpus by year
| Year | Articles | Newspapers | Evening Star share | Median words | Dictionary rate | Legible | Near-duplicates |
|---|---|---|---|---|---|---|---|
| 1900 | 10,007 | 270 | 2.1% | 111 | 91.8% | 74.4% | 0.08% |
| 1903 | 10,022 | 266 | 2.6% | 111 | 92.3% | 77.1% | 0.06% |
| 1906 | 10,032 | 287 | 3.9% | 118 | 92.2% | 77.1% | 0.01% |
| 1909 | 10,014 | 294 | 3.7% | 118 | 92.2% | 78.5% | 0.01% |
| 1912 | 10,003 | 249 | 3.6% | 116 | 92.2% | 76.8% | 0.04% |
| 1915 | 10,017 | 282 | 1.8% | 116 | 91.7% | 73.7% | 0.03% |
| 1918 | 10,004 | 290 | 2.7% | 116 | 91.9% | 77.8% | 0.02% |
| 1921 | 10,002 | 258 | 3.0% | 119 | 91.8% | 72.1% | 0.01% |
| 1924 | 10,008 | 106 | 17.9% | 114 | 91.7% | 80.3% | 0.03% |
| 1927 | 10,003 | 59 | 25.5% | 118 | 92.4% | 87.2% | 0.03% |
| 1930 | 10,004 | 44 | 23.0% | 113 | 92.2% | 88.6% | 0.02% |
| 1933 | 10,008 | 50 | 25.5% | 118 | 92.5% | 85.0% | 0.02% |
| 1936 | 10,008 | 58 | 25.6% | 115 | 91.9% | 90.3% | 0.09% |
| 1939 | 10,008 | 57 | 30.8% | 114 | 92.0% | 86.3% | 0.05% |
| 1942 | 10,016 | 67 | 25.0% | 113 | 91.4% | 86.1% | 0.04% |
| 1945 | 10,002 | 70 | 25.4% | 119 | 90.7% | 83.2% | 0.05% |
| 1948 | 10,017 | 65 | 45.9% | 123 | 91.1% | 77.9% | 0.04% |
| 1951 | 10,017 | 58 | 47.8% | 116 | 91.1% | 83.9% | 0.09% |
| 1954 | 10,009 | 48 | 49.7% | 124 | 91.1% | 83.9% | 0.10% |
| 1957 | 10,009 | 42 | 51.0% | 118 | 90.8% | 89.2% | 0.05% |
| 1960 | 10,008 | 37 | 58.5% | 114 | 91.2% | 88.9% | 0.07% |
| 1963 | 10,003 | 38 | 57.3% | 119 | 91.5% | 94.1% | 0.04% |
What is public and what is regenerated
The repository holds code, configuration, manifests with per-year content hashes, reference labels and derived results. The sampled article text (public domain, CC-BY-4.0 as packaged) is regenerated by make data rather than committed, to keep the repository small. Excerpts on this site are short and shown as evidence.