Is Social Mobility a Markov Process?
Introduction
Almost every empirical estimate of intergenerational mobility is a two-generation estimate: a son’s status rank \(z_t\) is regressed on his father’s rank \(z_{t-1}\), and the slope \(b_1\) is reported. Work I am doing on religion and mobility in the nineteenth-century Netherlands follows the same design, and the slope is lower on the Protestant side of the old Archdiocese of Mechelen border, which we interpret as greater mobility.
The design rests on an assumption that is rarely stated, namely that status transmission is a first-order Markov process. Whatever a grandfather transmits to his grandson, he transmits through the father. If that assumption holds, the two-generation slope is a sufficient statistic and the \(k\)-generation association is \(b_1^k\). A genealogical literature spanning England, Germany, and China finds that it does not hold.1 In what follows I describe the test that detects the failure, the latent-factor model that accounts for it, and the identification result that model delivers for comparisons between groups. I then report what three linked generations of Dutch genealogical data show.
The Markov Benchmark
The test requires only three linked generations. Under a first-order Markov process, the directly estimated grandfather–grandson slope \(b_2\) must equal the iterated prediction \(b_1^2\). Estimating both and comparing them is therefore a specification test on the two-generation design.
The two quantities do not match in any setting where they have been compared. In Shiue’s five-generation panel for Tongcheng, China, a father–son rank–rank slope of \(0.579\) implies a grandfather–grandson association of \(0.335\), whereas the directly estimated value is \(0.398\). Braun and Stuhler report a comparable excess in German data, and Keller and Shiue find that extended kin and in-law families carry status information beyond the father altogether. Status persists by more than the two-generation slope allows one to extrapolate. One reading of that fact is a direct grandparental effect, in which the grandfather retains a coefficient conditional on the father. A structural reading is a latent-factor model.
A Latent-Factor Model
Suppose observed status \(y\) is a noisy expression of a latent family endowment \(e\), and that it is the endowment rather than status itself that is transmitted:
\[ y_{t} = \lambda e_{t} + u_{t}, \qquad e_{t} = \rho\, e_{t-1} + v_{t} , \]
with \(u_t\) and \(v_t\) white noise, uncorrelated with everything else and across generations. The parameter \(\lambda\) is the loading, that is, how tightly occupational attainment tracks the endowment, while \(\rho\) is the persistence of the endowment itself. Because the analysis uses within-cohort ranks, \(\mathrm{Var}(y) = \mathrm{Var}(e) = 1\) is a normalisation and slopes can be read as correlations.2
The \(k\)-generation association follows directly. Since the endowment is an AR(1), iterating it \(k\) times gives \(e_t = \rho^k e_{t-k}\) plus noise dated after \(t-k\), and that noise is orthogonal to \(e_{t-k}\), so \(\mathrm{Cov}(e_t, e_{t-k}) = \rho^k\). Substituting the measurement equation and using the fact that the \(u\)’s are uncorrelated across generations:
\[ b_k \;=\; \frac{\mathrm{Cov}(y_t, y_{t-k})}{\mathrm{Var}(y_{t-k})} \;=\; \mathrm{Cov}(\lambda e_t + u_t,\; \lambda e_{t-k} + u_{t-k}) \;=\; \lambda^{2}\,\mathrm{Cov}(e_t, e_{t-k}) \;=\; \lambda^{2}\rho^{k} . \]
The substance of the result lies in the asymmetry between the two exponents. The only channel connecting an ancestor’s observed status to a descendant’s runs through the endowment, and using that channel means paying the loading twice, once to move from the ancestor’s status into his endowment and once to move from the descendant’s endowment out into his status, however many generations lie between them. The persistence \(\rho\), by contrast, is paid once per generation. The term \(\lambda^2\) is thus a fixed toll, while \(\rho^k\) compounds.
Two implications follow. First, \(b_1 = \lambda^2\rho\) lies below \(\rho\) whenever \(\lambda^2 < 1\), so the two-generation slope overstates mobility. Second, \(b_2 = \lambda^2\rho^2\) exceeds \(b_1^2 = \lambda^4\rho^2\) for any \(\lambda < 1\), because squaring the two-generation slope charges the toll twice over. The excess persistence documented in the multigenerational literature is what the model predicts.3
Identification of Latent Persistence
Taking the ratio of the two slopes gives
\[ \frac{b_2}{b_1} = \rho . \]
The loading cancels, so latent persistence is recoverable free of \(\lambda\), and therefore free of however noisily occupation happens to proxy for underlying family status.
This matters because a cross-group comparison of \(b_1\) is not identified as a comparison of mobility. If the persistence slope is lower on the Protestant side, two explanations are consistent with it. Either Protestant families regressed to the mean faster (\(\rho_P < \rho_C\)), or occupational titles tracked latent status more loosely there (\(\lambda_P < \lambda_C\)), which is a measurement difference that presents itself as social fluidity. The two-generation slope cannot separate the two explanations, whereas the ratio can.
Evidence from Dutch Lineages
Linking the Dutch genealogical lineages yields \(38{,}475\) grandfather–father–son triples. Mobility is strongly non-Markov on both sides of the border. The direct slope \(b_2\) is \(0.290\) on the Protestant side and \(0.335\) on the Catholic side, exceeding the iterated predictions \(b_1^2\) of \(0.176\) and \(0.235\) by 65 and 43 percent respectively, magnitudes comparable to the Chinese and German estimates. The excess is statistically indistinguishable across the border (\(-0.015\), s.e. \(0.026\)), so the cross-side contrast is not an artefact of differential higher-order structure. Latent persistence is \(\rho \approx 0.69\) on both sides, which is also the value the ratio \(0.398/0.579\) implies for Tongcheng in Shiue’s panel.
Read strictly through the model, equal \(\rho\) attributes the Protestant advantage entirely to the loading, with \(\lambda^2 = 0.61\) against \(0.70\): occupational position was less tightly determined by family endowment in every generation, which is the allocative channel rather than a measurement artefact. That decomposition should be treated with caution. The ratio test is underpowered on the Catholic side, and the grouped surname estimator, which averages out precisely the \(\lambda\)-type noise, does find higher Catholic persistence.
Conclusion
The Markov assumption becomes testable as soon as three linked generations are observed, and in the Dutch data it fails. It fails symmetrically across the confessional border, however, and the ratio \(b_2/b_1\) supplies a comparison statistic that is unaffected by the measurement problem that would otherwise undermine the cross-group contrast. Re-estimating the differential-persistence IV at the grandparent horizon is consistent with this: the interaction of grandfather status with Protestant share remains negative, between \(-0.093\) and \(-0.119\), and statistically significant. The Protestant advantage compounds across generations rather than dissipating.
Footnotes
Clark, G., & Cummins, N. (2015). Intergenerational wealth mobility in England, 1858–2012: Surnames and social mobility. The Economic Journal, 125(582), 61–85., Braun, S. T., & Stuhler, J. (2018). The transmission of inequality across multiple generations: Testing recent theories with evidence from Germany. The Economic Journal, 128(609), 576–611., Shiue, C. H. (2025). Social mobility in the long run: An analysis of Tongcheng, China, 1300 to 1900. The Journal of Economic History, 85(2), 370–410., Keller, W., & Shiue, C. H. (2023). Intergenerational mobility of daughters and marital sorting: New evidence from Imperial China (NBER Working Paper No. 31695). National Bureau of Economic Research.↩︎
The normalisation is innocuous here but not free in general: it is what permits ignoring secular drift in the variance of status across cohorts, which ranking within cohorts already removes.↩︎
Braun and Stuhler use the reverse letters; their \(\rho\) is my \(\lambda\) and their \(\lambda\) is my \(\rho\), so the two notations should be read across with care.↩︎