Using 31.6 million scraped administrative and genetic‑test records mapped to ~11,000 U.S. surnames, ancestry proportions from 23andMe pages predict a composite surname socioeconomic score (salaries, occupational prestige, physician licensure, criminal representation) with cross‑validated correlation r ≈ 0.899. The result is robust across modeling approaches and appears driven in part by selective migration, testing uptake, and how surnames traverse populations.
— If replicable, this finding reframes debates about the explanatory power of ancestry data (social vs. biological), informs immigration and inequality narratives, and raises privacy and misuse risks from surname‑level genetic inferences.
Uncorrelated
2026.09.05
100% relevant
Van Pelt & Kirkegaard (2026) paper; scraped 23andMe hidden ancestry fields, 31.6M records, 11,008 surnames, XGBoost out‑of‑sample r = 0.899 for an S‑factor composite.
← Back to all ideas