ABSTRACT
Quantifying intergenerational changes in educational homophily is challenging when education level of marriageable individuals is generation-specific. To address this challenge, the literature proposed various indicators. We argue that suitable indicators should meet several analytical criteria. However, our main criterion is empirical in nature: suitable indicators, when computed from data on American couples born after 1930, should exhibit a U-shaped trend at the national level and for the majority of the US states. A cardinal homophily indicator constructed with the NM-method, based on an old forgotten and recently reinvented ordinal indicator, is one of the few indicators that meets the criteria.
Acknowledgments
The author acknowledges the U.S. Census Bureau, the source of all the underlying data in IPUMS. Also, she acknowledges comments from the participants of the Research Seminar at the University of Turku. Replication: the code and the data allowing the replication of the analysis in the paper are available at Mendeley (see https://data.mendeley.com/datasets/7cm6673vg8
Disclosure statement
No potential conflict of interest was reported by the author(s).
Supplementary material
Supplemental data for this article can be accessed online at https://doi.org/10.1080/0022250X.2026.2662025
Notes
1 Although most of the papers in the empirical assortative mating–inequality literature focus on the reverse causality, both directions are of interest to researchers. For example, Fernandez et al. (Citation2005) wrote that an “increase in inequality increases sorting by making skilled workers less willing to form households with unskilled workers.”
2 Rosenfeld (Citation2008) argues that “One reason the literature on educational assortative mating has produced divergent results is that the changes over time in educational assortative mating are fairly subtle. The underlying marginal distribution of education for both men and women has changed dramatically since the early 20th century, and evaluations of educational assortative mating depend to a great extent on how the rapidly changing educational attainments are controlled for.”
3 Our measure-selection is similar to the approach that Tan et al. (Citation2004) refer to as the “laborious task” of “comparing how well the rankings [of joint distributions of couples] produced by each measure agree with the expectations of domain experts.” As we will see, this task is not “laborious” in our specific case since we can rely on the stylized trend in inequality reflecting the consensual view of multiple studies using a diverse set of data of independent sources.
4 It is worth noting that in practice the condition of is rarely met. However, if it is not met, the Simplified Liu–Lu-indicator and the v-indicator diverge purely due to rounding and their difference is negligible.
5 It is worth noting that the indicator defined by EquationEq. 15(15) (15) in J. Coleman (Citation1958) has been forgotten in the subsequent literature. Surprisingly, it is not discussed even in T. Coleman and Lukina (Citation2025).
6 The possibility that the IPF fails to converge is illustrated with an example in Appendix B.
7 It is easy to see that the odds-ratio is kept fixed by the IPF: in each step of the iterative procedure the denominator and the nominator of the odds ratio are multiplied by the same step-specific constant.
8 Naszodi (Citation2025) presents a comprehensive set of arguments explaining why the IPF is not suitable
for constructing counterfactuals in the context of analyzing assortative mating.
9 Naszodi and Mendonca (Citation2023a) found no sensitivity to any of their empirical results to whether the (Generalized) Liu–Lu-indicator is computed by taking the integer value of in EquationEq. 11(11) (11) , or using the unrounded value. Accordingly, we follow the original formulation of Liu and Lu (Citation2006).
10 While Clark and Cummins (Citation2022) derive that the “greater is assortment, the lower will be social mobility” in a model, and a number of empirical papers document this tendency for many countries, we consider it as a criterion reasonable to be imposed on suitable (relative) measures of homophily.
11 Our empirical condition is binary for each decade and state, with the threshold set to zero. For example, if a given method indicates that the degree of sorting decreased between 1960 and 1970 in California, and income inequality also decreased in California during the 1960s, then the method satisfies the state-decade-specific condition. Conversely, if the method indicates an increase in sorting over the same decade and state, it fails to meet the condition.
12 Such diversity in subnational trends is perhaps not surprising to readers who know that there is a popular saying in the US “If it plays in Peoria (Illinois), it’ll play anywhere,” and that there is no similar saying with Washington (DC), or any city in Alaska.
13 In the 1970 census, the number of young couples (with male partners aged between 30 and 34 years) with unknown or unreported residence were more than 800,000.
14 No census data are available for Delaware, Idaho, Montana, North Dakota, South Dakota, Vermont, and Wyoming for 1970 making it impossible to analyze the trends for the 1960s and 1970s. In addition, we had to eliminate the data for one more state-decade pair, because no couple of a certain type was registered in the census. This elimination was necessary because the IPF algorithm fails to work if there are zero entries in the seed table.
15 As it is rightly pointed out by an anonymous reviewer, “with enough methods tested, some will match the target shape by chance.” Our additional sensitivity analysis shows that the NM-method performs better than its alternatives not by chance, but systematically.
16 Our empirical criterion for selecting the suitable analytical tools is not the classic out-of-sample fit criterion, because none of the analytical tools applied to analyze sorting in the first half of the states are fitted on the data. Thereby, the tools cannot learn anything from the training data (while researchers can learn the relative performance of the tools against the criteria). Not allowing the tools to learn from the training data has the following advantage: we do not have to deal with the well-known problem of over-fitting that complex models with more parameters to be estimated on the training data are more prone to.
17 So, the trend in educational homophily in the US identified by the IPF is most like the atypical trend in income inequality in Alaska.
18 To recall, the dynamics of the top 10% income share in DC also exhibited an atypical W-shaped trend between 1940 and 2015.
19 Kahneman (Citation2011) offers a number of funny examples illustrating this point.
20 In other words, standard robustness analyzes to the choice of the indicator do not qualify as a criterion on the set of criteria (see Athey & Imbens, Citation2017).
21 Unlike this paper, some studies using the IPF in the assortative mating literature motivate their choice by “playing out the authority card”: they appeal to the credibility and expertize of the inventors of the IPF (see e.g. Leesch & Skopek, Citation2023). These papers overlook that Stephan and Deming (Citation1940) proposed the application of the IPF to solve a problem different from constructing counterfactual predictions. Moreover, Stephan and Deming (Citation1940) even warned that their algorithm is “not by itself useful for prediction” (see p. 444).
22 See the web-page of the International Demographic Inequality Lab (https://idil.liIDIL reporting the historical trend in overall inequality in 80 countries computed from data on couples.
Facts Only
* The author uses data on American couples born after 1930.
* Data sources include the U.S. Census Bureau and IPUMS.
* The study focuses on quantifying intergenerational changes in educational homophily.
* The NM-method is a cardinal homophily indicator based on an ordinal indicator from J. Coleman (1958).
* A primary empirical criterion for a suitable indicator is the exhibition of a U-shaped trend at the national level and in most US states.
* Iterative Proportional Fitting (IPF) is used as a method for constructing counterfactuals.
* Analysis excludes census data for Delaware, Idaho, Montana, North Dakota, South Dakota, Vermont, and Wyoming for 1970.
* Code and data for replication are hosted on Mendeley.
* Supplementary material is available via DOI 10.1080/0022250X.2026.2662025.
Executive Summary
Measuring educational homophily—the tendency of individuals to marry those with similar education levels—is complicated by the fact that educational opportunities change across generations. To resolve this, different statistical indicators have been proposed, but many produce divergent results because they fail to account for shifting marginal distributions of educational attainment over time.
A specific cardinal indicator, developed via the NM-method, is proposed as the most reliable tool. Its validity is grounded in an empirical benchmark: a suitable measure should mirror the observed U-shaped trend in income inequality seen at the US national and state levels. While other methods like the Iterative Proportional Fitting (IPF) are common in the literature, they are argued to be unsuitable for constructing the necessary counterfactuals in this context. Despite some data gaps in specific states during the 1970s, the NM-method is presented as systematically superior to alternatives, providing a more consistent alignment with established socio-economic trends.
Full Take
This research operates in ACADEMIC MODE, attempting to standardize the measurement of social sorting.
1. METHODOLOGY CHECK: The author employs a "target-shape" validation strategy, where the tool's validity is judged by its ability to replicate a known stylized fact (the U-shaped trend of inequality). While logically intuitive, a peer reviewer would question if this creates a circularity: is the measure "correct" because it captures the phenomenon, or is it simply tuned to match a pre-existing narrative of inequality? The exclusion of several states for the 1970 census is a noted limitation that could introduce minor geographic bias.
2. CLAIMS vs EVIDENCE: The claim that the NM-method is "one of the few" that meets the criteria is supported by sensitivity analyses showing it outperforms others systematically, rather than by chance. However, the reliance on the "consensual view" of inequality as a benchmark moves the evidence base from purely mathematical to a hybrid of math and sociological consensus.
3. LITERATURE CONTEXT: The work attempts to rescue a "forgotten" 1958 Coleman indicator, positioning itself as a corrective to recent literature (e.g., Leesch & Skopek) that the author suggests relies on the "authority card" regarding the IPF method. It challenges the current status quo of how assortative mating is quantified.
4. REAL-WORLD IMPLICATIONS: If the NM-method is adopted, it may shift our understanding of social mobility. If sorting is more accurately measured as U-shaped, the policy implications for educational equity and income redistribution become more urgent.
5. BRIDGE QUESTIONS: Would the NM-method maintain its superiority if applied to non-US data where inequality trends are not U-shaped? Does the alignment with income inequality prove the measure is capturing homophily, or is it capturing a broader, unobserved latent variable?
COUNTERSTRIKE SCAN: A bad actor would use this to claim "scientific proof" of inevitable social stratification to discourage educational investment. The actual content is a technical methodological debate and does not match this pattern.
