Validation
Every method is checked against a stratified sample of reference labels. The reference labels are not human annotation, and this page says so first.
Who produced the reference labels
v1 AI-annotator reference labels (not human). 600 extracted entities were drawn by stratified random sampling (period × honorific class × role present); 547 of them are real people and carry the gender and role labels used below. The annotator worked blind to every method's output, following research/annotation_protocol.md. Estimates are reweighted to population proportions. Replacing these labels with human annotation is the first open item, and the annotation tool exists for exactly that.
Self-consistency of the reference labels
120 items. Intra-annotator self-consistency: the same AI annotator re-labelled a 20% subset blind to method outputs, in shuffled order, after the full pass, but in the same working session, with first-pass labels still in its context history. Treat it as an UPPER BOUND on self-consistency, not as independent agreement. A near-perfect κ here is expected for that reason and should not be read as evidence that the labels are right.
| Field | Cohen κ | Agreement | Items |
|---|---|---|---|
| is_person | 1.00 | 100% | 120 |
| gender_text | 1.00 | 100% | 120 |
| quoted | 1.00 | 100% | 120 |
| role · PUBLIC_OFFICE | 1.00 | 100% | 120 |
| role · MILITARY | 1.00 | 100% | 120 |
| role · BUSINESS | 1.00 | 100% | 120 |
| role · LABOR | — | 100% | 120 |
| role · PROFESSIONAL | 1.00 | 100% | 120 |
| role · ARTS_SPORTS | 1.00 | 100% | 120 |
| role · CIVIC | 0.96 | 99% | 120 |
| role · SOCIAL | 1.00 | 100% | 120 |
| role · FAMILY | 1.00 | 100% | 120 |
| role · CRIME_ACCIDENT | 1.00 | 100% | 120 |
| role · AUTHORITY | 1.00 | 100% | 120 |
Is the extracted span a person?
Of sampled entities, 69.6% (sampling-weighted) are real references to a person. The rest are OCR fragments, places and organisations tagged as people. Every entity-level estimate carries this noise.
Gender signal
For each gender-evidence rule: accuracy when it assigns a gender, and recall of the genders the text does signal.
| Rule | Accuracy when assigning | Recall of signalled genders | Share assigned | Assigned (n) |
|---|---|---|---|---|
| gender_hp_v2 | 0.92 | 0.72 | 38% | 417 |
| gender_h | 0.98 | 0.57 | 28% | 391 |
| gender_hp | 0.92 | 0.72 | 38% | 415 |
| gender_hpn | 0.93 | 0.78 | 40% | 428 |
| E_llm | 0.52 | 0.97 | 90% | 530 |
| Rule | Gender | Precision | Recall | F1 |
|---|---|---|---|---|
| gender_hp_v2 | F | 0.98 | 0.80 | 0.88 |
| gender_hp_v2 | M | 0.83 | 0.60 | 0.69 |
| gender_h | F | 0.99 | 0.76 | 0.86 |
| gender_h | M | 0.94 | 0.29 | 0.44 |
| gender_hp | F | 0.98 | 0.80 | 0.88 |
| gender_hp | M | 0.83 | 0.60 | 0.69 |
| gender_hpn | F | 0.98 | 0.86 | 0.92 |
| gender_hpn | M | 0.84 | 0.65 | 0.73 |
| E_llm | F | 0.85 | 0.96 | 0.91 |
| E_llm | M | 0.33 | 0.98 | 0.50 |
Genders assigned where the text gives none
This is the failure mode an LLM is most prone to: guessing gender from a first name or from occupational stereotype when the text is silent.
| Method | Share of text-silent people given a gender | Given a gender (n) | Text-silent people (n) |
|---|---|---|---|
| gender_hp_v2 | 4.5% | 28 | 128 |
| gender_h | 0.7% | 20 | 128 |
| gender_hp | 4.5% | 27 | 128 |
| gender_hpn | 4.8% | 28 | 128 |
| E_llm | 83.6% | 115 | 128 |
Table view
| Year | pronoun rule agrees with honorific |
|---|---|
| 1900 | 93.4% · n=1,388 |
| 1903 | 92.9% · n=1,248 |
| 1906 | 92.6% · n=1,266 |
| 1909 | 91.3% · n=1,361 |
| 1912 | 90.7% · n=1,318 |
| 1915 | 90.6% · n=1,150 |
| 1918 | 92.7% · n=1,117 |
| 1921 | 90.4% · n=1,444 |
| 1924 | 90.7% · n=1,155 |
| 1927 | 88.7% · n=1,102 |
| 1930 | 89.4% · n=1,108 |
| 1933 | 87.6% · n=1,118 |
| 1936 | 89.4% · n=1,127 |
| 1939 | 92.1% · n=1,295 |
| 1942 | 90.4% · n=1,035 |
| 1945 | 85.2% · n=987 |
| 1948 | 92.7% · n=1,350 |
| 1951 | 88.4% · n=1,376 |
| 1954 | 90.1% · n=1,170 |
| 1957 | 90.4% · n=1,293 |
| 1960 | 91.7% · n=1,428 |
| 1963 | 91.0% · n=1,229 |
AmericanStories (Dell et al. 2023, CC-BY-4.0), built on Chronicling America (Library of Congress). 10,000 sampled articles per year, 22 years.
Roles
Table view
| Method | AUTHORITY | PUBLIC_OFFICE | MILITARY | BUSINESS | LABOR | PROFESSIONAL | ARTS_SPORTS | CIVIC | SOCIAL | FAMILY | CRIME_ACCIDENT |
|---|---|---|---|---|---|---|---|---|---|---|---|
| A_lexical | 0.66 | 0.49 | 0.44 | 0.31 | 0.14 | 0.58 | 0.23 | 0.15 | 0.18 | 0.51 | 0.04 |
| B_dependency | 0.48 | 0.55 | 0.73 | 0.13 | 0.32 | 0.56 | 0.12 | 0.13 | 0.05 | 0.36 | 0.05 |
| C_bow@0.5 | 0.38 | 0.25 | — | 0.22 | — | 0.30 | 0.14 | 0.48 | 0.67 | 0.38 | — |
| D_embed@0.5 | 0.59 | 0.44 | 0.11 | 0.45 | — | 0.30 | 0.47 | 0.35 | 0.45 | 0.25 | 0.28 |
| E_llm | 0.84 | 0.75 | 0.59 | 0.86 | 0.68 | 0.84 | 0.88 | 0.60 | 0.69 | 0.38 | 0.65 |
| Method | Role | Precision | Recall | F1 | Reference positives | Reference prevalence | Predicted prevalence |
|---|---|---|---|---|---|---|---|
| E_llm | PUBLIC_OFFICE | 0.63 | 0.93 | 0.75 | 90 | 9.4% | 14.0% |
| E_llm | BUSINESS | 0.82 | 0.90 | 0.86 | 43 | 10.1% | 11.1% |
| E_llm | PROFESSIONAL | 0.80 | 0.89 | 0.84 | 58 | 9.1% | 10.0% |
| E_llm | CIVIC | 0.63 | 0.57 | 0.60 | 89 | 10.1% | 9.2% |
| E_llm | FAMILY | 0.66 | 0.26 | 0.38 | 104 | 14.0% | 5.7% |
| E_llm | AUTHORITY | 0.76 | 0.93 | 0.84 | 187 | 28.2% | 34.5% |
| A_lexical | PUBLIC_OFFICE | 0.45 | 0.53 | 0.49 | 90 | 9.4% | 11.0% |
| A_lexical | BUSINESS | 0.77 | 0.19 | 0.31 | 43 | 10.1% | 2.6% |
| A_lexical | PROFESSIONAL | 0.66 | 0.52 | 0.58 | 58 | 9.1% | 7.2% |
| A_lexical | CIVIC | 0.25 | 0.11 | 0.15 | 89 | 10.1% | 4.4% |
| A_lexical | FAMILY | 0.65 | 0.42 | 0.51 | 104 | 14.0% | 9.1% |
| A_lexical | AUTHORITY | 0.80 | 0.56 | 0.66 | 187 | 28.2% | 19.9% |
| B_dependency | PUBLIC_OFFICE | 0.73 | 0.44 | 0.55 | 90 | 9.4% | 5.8% |
| B_dependency | BUSINESS | 0.60 | 0.07 | 0.13 | 43 | 10.1% | 1.3% |
| B_dependency | PROFESSIONAL | 0.95 | 0.40 | 0.56 | 58 | 9.1% | 3.8% |
| B_dependency | CIVIC | 0.54 | 0.07 | 0.13 | 89 | 10.1% | 1.4% |
| B_dependency | FAMILY | 0.83 | 0.23 | 0.36 | 104 | 14.0% | 4.0% |
| B_dependency | AUTHORITY | 0.87 | 0.33 | 0.48 | 187 | 28.2% | 10.7% |
| C_bow@0.3 | PUBLIC_OFFICE | 0.24 | 0.25 | 0.25 | 90 | 9.4% | 10.1% |
| C_bow@0.3 | BUSINESS | 0.85 | 0.13 | 0.22 | 43 | 10.1% | 1.5% |
| C_bow@0.3 | PROFESSIONAL | 0.50 | 0.19 | 0.27 | 58 | 9.1% | 3.4% |
| C_bow@0.3 | CIVIC | 0.52 | 0.48 | 0.50 | 89 | 10.1% | 9.4% |
| C_bow@0.3 | FAMILY | 0.56 | 0.36 | 0.43 | 104 | 14.0% | 9.0% |
| C_bow@0.3 | AUTHORITY | 0.55 | 0.28 | 0.37 | 187 | 28.2% | 14.4% |
| C_bow@0.5 | PUBLIC_OFFICE | 0.28 | 0.22 | 0.25 | 90 | 9.4% | 7.4% |
| C_bow@0.5 | BUSINESS | 0.97 | 0.12 | 0.22 | 43 | 10.1% | 1.3% |
| C_bow@0.5 | PROFESSIONAL | 0.79 | 0.18 | 0.30 | 58 | 9.1% | 2.1% |
| C_bow@0.5 | CIVIC | 0.70 | 0.37 | 0.48 | 89 | 10.1% | 5.3% |
| C_bow@0.5 | FAMILY | 0.59 | 0.28 | 0.38 | 104 | 14.0% | 6.7% |
| C_bow@0.5 | AUTHORITY | 0.68 | 0.26 | 0.38 | 187 | 28.2% | 10.8% |
| C_bow@0.7 | PUBLIC_OFFICE | 0.26 | 0.15 | 0.19 | 90 | 9.4% | 5.6% |
| C_bow@0.7 | BUSINESS | 0.97 | 0.12 | 0.21 | 43 | 10.1% | 1.3% |
| C_bow@0.7 | PROFESSIONAL | 0.86 | 0.15 | 0.25 | 58 | 9.1% | 1.6% |
| C_bow@0.7 | CIVIC | 0.64 | 0.28 | 0.39 | 89 | 10.1% | 4.3% |
| C_bow@0.7 | FAMILY | 0.63 | 0.25 | 0.36 | 104 | 14.0% | 5.6% |
| C_bow@0.7 | AUTHORITY | 0.68 | 0.20 | 0.31 | 187 | 28.2% | 8.4% |
| D_embed@0.3 | PUBLIC_OFFICE | 0.29 | 0.62 | 0.39 | 90 | 9.4% | 20.5% |
| D_embed@0.3 | BUSINESS | 0.30 | 0.58 | 0.40 | 43 | 10.1% | 19.7% |
| D_embed@0.3 | PROFESSIONAL | 0.21 | 0.35 | 0.27 | 58 | 9.1% | 15.1% |
| D_embed@0.3 | CIVIC | 0.33 | 0.55 | 0.41 | 89 | 10.1% | 17.1% |
| D_embed@0.3 | FAMILY | 0.28 | 0.47 | 0.35 | 104 | 14.0% | 23.4% |
| D_embed@0.3 | AUTHORITY | 0.47 | 0.71 | 0.56 | 187 | 28.2% | 42.9% |
| D_embed@0.5 | PUBLIC_OFFICE | 0.34 | 0.61 | 0.44 | 90 | 9.4% | 16.8% |
| D_embed@0.5 | BUSINESS | 0.37 | 0.58 | 0.45 | 43 | 10.1% | 16.1% |
| D_embed@0.5 | PROFESSIONAL | 0.26 | 0.35 | 0.30 | 58 | 9.1% | 12.6% |
| D_embed@0.5 | CIVIC | 0.29 | 0.44 | 0.35 | 89 | 10.1% | 15.6% |
| D_embed@0.5 | FAMILY | 0.22 | 0.30 | 0.25 | 104 | 14.0% | 18.9% |
| D_embed@0.5 | AUTHORITY | 0.53 | 0.67 | 0.59 | 187 | 28.2% | 35.4% |
| D_embed@0.7 | PUBLIC_OFFICE | 0.33 | 0.45 | 0.38 | 90 | 9.4% | 13.1% |
| D_embed@0.7 | BUSINESS | 0.41 | 0.58 | 0.48 | 43 | 10.1% | 14.5% |
| D_embed@0.7 | PROFESSIONAL | 0.29 | 0.34 | 0.31 | 58 | 9.1% | 10.5% |
| D_embed@0.7 | CIVIC | 0.32 | 0.43 | 0.37 | 89 | 10.1% | 13.7% |
| D_embed@0.7 | FAMILY | 0.23 | 0.30 | 0.26 | 104 | 14.0% | 18.4% |
| D_embed@0.7 | AUTHORITY | 0.50 | 0.57 | 0.53 | 187 | 28.2% | 32.2% |
Is the error the same for women and men?
Authority-role precision and recall computed separately among people the primary rule signals as women and as men. If a method over-assigns authority to women, or under-finds it for men, the women's share it reports is biased even when overall F1 looks acceptable. Counts are small, so read these as warnings, not corrections. Typical cases are on the failures page.
| Method | Signalled gender | Precision | Recall | Predicted positive (n) | Reference positive (n) |
|---|---|---|---|---|---|
| A_lexical | women | 0.27 | 0.51 | 41 | 12 |
| A_lexical | men | 0.92 | 0.45 | 84 | 113 |
| B_dependency | women | 0.43 | 0.24 | 21 | 12 |
| B_dependency | men | 0.96 | 0.35 | 75 | 113 |
| C_bow@0.5 | women | 0.43 | 0.29 | 8 | 12 |
| C_bow@0.5 | men | 0.62 | 0.22 | 43 | 113 |
| D_embed@0.5 | women | 0.13 | 0.32 | 29 | 12 |
| D_embed@0.5 | men | 0.50 | 0.54 | 109 | 113 |
| E_llm | women | 0.60 | 0.85 | 19 | 12 |
| E_llm | men | 0.75 | 0.79 | 135 | 113 |