Vol. IFrozen 2026-09-11Git 2dfc1b5*220,221 articles · 1900–1963

GenderNews Atlas

Sections

Failures

Where the pipeline breaks, how often, and in which direction the error pushes the estimates.

Named-entity recognition on OCR

30.4% of extracted “people” in the reference sample are not people: OCR fragments, streets, ships and firms. The error is not gender-neutral in effect. Honorific-bearing spans are almost always real people, so the gender-signalled subset is cleaner than the UNKNOWN pool.

Pronouns and coreference

The pronoun rule disagrees with the honorific on 9.3% of the 27,065 people who carry both signals. The typical failure is a pronoun that refers to a second person the parser did not tag (“Smith told his wife she…”). Conflicting evidence sends a person to UNKNOWN rather than to a guess.

Share of people whose gender evidence conflicts
Table view
Yearentity gender conflicts (dropped to UNKNOWN)
19000.0%
19030.0%
19060.1%
19090.0%
19120.0%
19150.0%
19180.0%
19210.0%
19240.0%
19270.1%
19300.1%
19330.1%
19360.1%
19390.1%
19420.0%
19450.0%
19480.0%
19510.0%
19540.0%
19570.0%
19600.0%
19630.0%

AmericanStories (Dell et al. 2023, CC-BY-4.0), built on Chronicling America (Library of Congress). 10,000 sampled articles per year, 22 years.

Entity resolution

Surname-only mentions that could belong to more than one person in the article are dropped rather than guessed. Period naming conventions make this frequent: a wife named by her husband's full name, “Mr. and Mrs.” pairs.

YearAmbiguous mentions dropped per articleMentions per personPattern-only people (excluded from primary)
19000.041.15624
19030.041.16605
19060.041.16556
19090.051.16571
19120.061.15602
19150.041.15634
19180.031.13526
19210.041.15753
19240.041.16511
19270.041.16655
19300.041.16548
19330.051.15570
19360.051.15617
19390.051.14583
19420.031.12537
19450.051.12646
19480.051.13711
19510.051.15669
19540.041.13715
19570.041.13637
19600.041.15657
19630.041.14633

OCR quality

OCR quality proxies by yearA period-dictionary word rate (Webster's Second, 1934) and AmericanStories' own legibility label.
Table view
Yearmedian dictionary-word ratearticles labelled legible
190091.8%74.4%
190392.3%77.1%
190692.2%77.1%
190992.2%78.5%
191292.2%76.8%
191591.7%73.7%
191891.9%77.8%
192191.8%72.1%
192491.7%80.3%
192792.4%87.2%
193092.2%88.6%
193392.5%85.0%
193691.9%90.3%
193992.0%86.3%
194291.4%86.1%
194590.7%83.2%
194891.1%77.9%
195191.1%83.9%
195491.1%83.9%
195790.8%89.2%
196091.2%88.9%
196391.5%94.1%

AmericanStories (Dell et al. 2023, CC-BY-4.0), built on Chronicling America (Library of Congress). 10,000 sampled articles per year, 22 years.

Quotation attribution

The named-speaker rule finds 31.3% of reference-labelled quoted people at a precision of 77.0%. Most misses are quotations attributed through a title or a pronoun the parser attached elsewhere.

Historical language

Some failures are not errors in the usual sense but shifts in the language itself. Chairman, alderman and congressman were applied to women office-holders, so they are treated as role terms, never as gender evidence. Married women were routinely named by their husbands' names (“Mrs. John Smith”), which is why entity resolution is keyed to honorific class. Miss appears in advertising as a size category (“misses' dresses”); requiring an NER person span keeps most of these out. Secretary and president name both government offices and club officers, which is why method B types organisational heads by the organisation they head.

LLM disagreement

Where the text gives no gender signal, the LLM still assigns a gender to 83.6% of such people in the reference sample. That is the clearest single reason its gender output is not used as evidence here.

Examples from the reference sample