Vol. IFrozen 2026-09-11Git 2dfc1b5*220,221 articles · 1900–1963

GenderNews Atlas

Sections

Validation

Every method is checked against a stratified sample of reference labels. The reference labels are not human annotation, and this page says so first.

Who produced the reference labels

v1 AI-annotator reference labels (not human). 600 extracted entities were drawn by stratified random sampling (period × honorific class × role present); 547 of them are real people and carry the gender and role labels used below. The annotator worked blind to every method's output, following research/annotation_protocol.md. Estimates are reweighted to population proportions. Replacing these labels with human annotation is the first open item, and the annotation tool exists for exactly that.

Self-consistency of the reference labels

120 items. Intra-annotator self-consistency: the same AI annotator re-labelled a 20% subset blind to method outputs, in shuffled order, after the full pass, but in the same working session, with first-pass labels still in its context history. Treat it as an UPPER BOUND on self-consistency, not as independent agreement. A near-perfect κ here is expected for that reason and should not be read as evidence that the labels are right.

FieldCohen κAgreementItems
is_person1.00100%120
gender_text1.00100%120
quoted1.00100%120
role · PUBLIC_OFFICE1.00100%120
role · MILITARY1.00100%120
role · BUSINESS1.00100%120
role · LABOR100%120
role · PROFESSIONAL1.00100%120
role · ARTS_SPORTS1.00100%120
role · CIVIC0.9699%120
role · SOCIAL1.00100%120
role · FAMILY1.00100%120
role · CRIME_ACCIDENT1.00100%120
role · AUTHORITY1.00100%120

Is the extracted span a person?

Of sampled entities, 69.6% (sampling-weighted) are real references to a person. The rest are OCR fragments, places and organisations tagged as people. Every entity-level estimate carries this noise.

Gender signal

For each gender-evidence rule: accuracy when it assigns a gender, and recall of the genders the text does signal.

RuleAccuracy when assigningRecall of signalled gendersShare assignedAssigned (n)
gender_hp_v20.920.7238%417
gender_h0.980.5728%391
gender_hp0.920.7238%415
gender_hpn0.930.7840%428
E_llm0.520.9790%530
RuleGenderPrecisionRecallF1
gender_hp_v2F0.980.800.88
gender_hp_v2M0.830.600.69
gender_hF0.990.760.86
gender_hM0.940.290.44
gender_hpF0.980.800.88
gender_hpM0.830.600.69
gender_hpnF0.980.860.92
gender_hpnM0.840.650.73
E_llmF0.850.960.91
E_llmM0.330.980.50

Genders assigned where the text gives none

This is the failure mode an LLM is most prone to: guessing gender from a first name or from occupational stereotype when the text is silent.

MethodShare of text-silent people given a genderGiven a gender (n)Text-silent people (n)
gender_hp_v24.5%28128
gender_h0.7%20128
gender_hp4.5%27128
gender_hpn4.8%28128
E_llm83.6%115128
Large-sample check: does the pronoun rule agree with the honorific?Among 27,065 people with both an honorific and a pronoun signal, the rule agrees 90.7% of the time (women 85.9%, men 96.9%). No annotation is needed: the honorific is the check.
Table view
Yearpronoun rule agrees with honorific
190093.4% · n=1,388
190392.9% · n=1,248
190692.6% · n=1,266
190991.3% · n=1,361
191290.7% · n=1,318
191590.6% · n=1,150
191892.7% · n=1,117
192190.4% · n=1,444
192490.7% · n=1,155
192788.7% · n=1,102
193089.4% · n=1,108
193387.6% · n=1,118
193689.4% · n=1,127
193992.1% · n=1,295
194290.4% · n=1,035
194585.2% · n=987
194892.7% · n=1,350
195188.4% · n=1,376
195490.1% · n=1,170
195790.4% · n=1,293
196091.7% · n=1,428
196391.0% · n=1,229

AmericanStories (Dell et al. 2023, CC-BY-4.0), built on Chronicling America (Library of Congress). 10,000 sampled articles per year, 22 years.

Roles

Role F1 against reference labels, by methodSampling-weighted to the population. Blank: no positives in the reference sample.
Table view
MethodAUTHORITYPUBLIC_OFFICEMILITARYBUSINESSLABORPROFESSIONALARTS_SPORTSCIVICSOCIALFAMILYCRIME_ACCIDENT
A_lexical0.660.490.440.310.140.580.230.150.180.510.04
B_dependency0.480.550.730.130.320.560.120.130.050.360.05
C_bow@0.50.380.250.220.300.140.480.670.38
D_embed@0.50.590.440.110.450.300.470.350.450.250.28
E_llm0.840.750.590.860.680.840.880.600.690.380.65
MethodRolePrecisionRecallF1Reference positivesReference prevalencePredicted prevalence
E_llmPUBLIC_OFFICE0.630.930.75909.4%14.0%
E_llmBUSINESS0.820.900.864310.1%11.1%
E_llmPROFESSIONAL0.800.890.84589.1%10.0%
E_llmCIVIC0.630.570.608910.1%9.2%
E_llmFAMILY0.660.260.3810414.0%5.7%
E_llmAUTHORITY0.760.930.8418728.2%34.5%
A_lexicalPUBLIC_OFFICE0.450.530.49909.4%11.0%
A_lexicalBUSINESS0.770.190.314310.1%2.6%
A_lexicalPROFESSIONAL0.660.520.58589.1%7.2%
A_lexicalCIVIC0.250.110.158910.1%4.4%
A_lexicalFAMILY0.650.420.5110414.0%9.1%
A_lexicalAUTHORITY0.800.560.6618728.2%19.9%
B_dependencyPUBLIC_OFFICE0.730.440.55909.4%5.8%
B_dependencyBUSINESS0.600.070.134310.1%1.3%
B_dependencyPROFESSIONAL0.950.400.56589.1%3.8%
B_dependencyCIVIC0.540.070.138910.1%1.4%
B_dependencyFAMILY0.830.230.3610414.0%4.0%
B_dependencyAUTHORITY0.870.330.4818728.2%10.7%
C_bow@0.3PUBLIC_OFFICE0.240.250.25909.4%10.1%
C_bow@0.3BUSINESS0.850.130.224310.1%1.5%
C_bow@0.3PROFESSIONAL0.500.190.27589.1%3.4%
C_bow@0.3CIVIC0.520.480.508910.1%9.4%
C_bow@0.3FAMILY0.560.360.4310414.0%9.0%
C_bow@0.3AUTHORITY0.550.280.3718728.2%14.4%
C_bow@0.5PUBLIC_OFFICE0.280.220.25909.4%7.4%
C_bow@0.5BUSINESS0.970.120.224310.1%1.3%
C_bow@0.5PROFESSIONAL0.790.180.30589.1%2.1%
C_bow@0.5CIVIC0.700.370.488910.1%5.3%
C_bow@0.5FAMILY0.590.280.3810414.0%6.7%
C_bow@0.5AUTHORITY0.680.260.3818728.2%10.8%
C_bow@0.7PUBLIC_OFFICE0.260.150.19909.4%5.6%
C_bow@0.7BUSINESS0.970.120.214310.1%1.3%
C_bow@0.7PROFESSIONAL0.860.150.25589.1%1.6%
C_bow@0.7CIVIC0.640.280.398910.1%4.3%
C_bow@0.7FAMILY0.630.250.3610414.0%5.6%
C_bow@0.7AUTHORITY0.680.200.3118728.2%8.4%
D_embed@0.3PUBLIC_OFFICE0.290.620.39909.4%20.5%
D_embed@0.3BUSINESS0.300.580.404310.1%19.7%
D_embed@0.3PROFESSIONAL0.210.350.27589.1%15.1%
D_embed@0.3CIVIC0.330.550.418910.1%17.1%
D_embed@0.3FAMILY0.280.470.3510414.0%23.4%
D_embed@0.3AUTHORITY0.470.710.5618728.2%42.9%
D_embed@0.5PUBLIC_OFFICE0.340.610.44909.4%16.8%
D_embed@0.5BUSINESS0.370.580.454310.1%16.1%
D_embed@0.5PROFESSIONAL0.260.350.30589.1%12.6%
D_embed@0.5CIVIC0.290.440.358910.1%15.6%
D_embed@0.5FAMILY0.220.300.2510414.0%18.9%
D_embed@0.5AUTHORITY0.530.670.5918728.2%35.4%
D_embed@0.7PUBLIC_OFFICE0.330.450.38909.4%13.1%
D_embed@0.7BUSINESS0.410.580.484310.1%14.5%
D_embed@0.7PROFESSIONAL0.290.340.31589.1%10.5%
D_embed@0.7CIVIC0.320.430.378910.1%13.7%
D_embed@0.7FAMILY0.230.300.2610414.0%18.4%
D_embed@0.7AUTHORITY0.500.570.5318728.2%32.2%

Is the error the same for women and men?

Authority-role precision and recall computed separately among people the primary rule signals as women and as men. If a method over-assigns authority to women, or under-finds it for men, the women's share it reports is biased even when overall F1 looks acceptable. Counts are small, so read these as warnings, not corrections. Typical cases are on the failures page.

MethodSignalled genderPrecisionRecallPredicted positive (n)Reference positive (n)
A_lexicalwomen0.270.514112
A_lexicalmen0.920.4584113
B_dependencywomen0.430.242112
B_dependencymen0.960.3575113
C_bow@0.5women0.430.29812
C_bow@0.5men0.620.2243113
D_embed@0.5women0.130.322912
D_embed@0.5men0.500.54109113
E_llmwomen0.600.851912
E_llmmen0.750.79135113