Grouped Surnames and Latent Persistence
Introduction
In an earlier post I argued that the two-generation mobility slope is not the parameter of interest. If observed status \(y\) is a noisy expression of a latent family endowment \(e\),
\[ y_t = \lambda e_t + u_t, \qquad e_t = \rho\, e_{t-1} + v_t , \]
then the \(k\)-generation association is \(b_k = \lambda^2 \rho^k\). The loading \(\lambda\) enters twice whatever the number of intervening generations, once to move from an ancestor’s observed status into his endowment and once to move from a descendant’s endowment back out into his status, whereas the persistence \(\rho\) enters once per generation. The two-generation slope therefore understates persistence, and the ratio \(b_2/b_1 = \rho\) recovers it because the loading cancels.
That result identifies \(\rho\) from time depth, and requires three linked generations. A second route to the same parameter requires only two generations and uses the cross-section instead. It is the grouped surname estimator of Clark and Cummins.1 In what follows I describe the estimator, derive the sense in which grouping identifies \(\rho\), and report where the two routes agree and disagree in Dutch genealogical data.
The Grouped Surname Estimator
The procedure assigns every occupation-bearing individual to a surname-by-birth-cohort cell. Each cell is given the mean status of its members, measured here by HISCLASS and oriented so that higher values denote higher status, and cell means are ranked within cohort, which purges the secular drift in the occupational structure. A surname’s rank in one cohort is then regressed on the same surname’s rank in the previous cohort. The slope \(b\) is the persistence parameter, ranging from \(b \approx 0\), where each cohort starts afresh, to \(b \approx 1\), where standing is inherited in full.
The estimator requires no father–son link. Cells in successive cohorts are measured from different individuals, which is why the method remains feasible in records where roughly a tenth of individuals carry a legible occupation. The estimates reported below use Dutch genealogical lineages, cohorts born before 1800, and cells with at least five occupation-bearing members, weighted by precision and by proximity to the border so that the slope reflects municipalities close to the Mechelen boundary.
Why Grouping Identifies Persistence
Averaging the measurement equation over the \(n\) members of a cell gives
\[ \bar y_{g,t} = \lambda \bar e_{g,t} + \bar u_{g,t}, \qquad \mathrm{Var}(\bar u) = \frac{1-\lambda^2}{n}, \]
where \(g\) indexes the surname group and \(t\) the birth cohort, so that \(\bar y_{g,t}\) is the mean observed status of the \(n\) occupation-bearing individuals named \(g\) born in window \(t\), and \(\bar e_{g,t}\) is the endowment those individuals share. The cell rather than the individual is now the unit: \(t\) advances by one 50-year window rather than by one father–son link, and the averaging runs over the \(n\) members within a cell. Normalising \(\mathrm{Var}(y) = \mathrm{Var}(e) = 1\) as before pins down the noise variance, since \(1 = \lambda^2 + \mathrm{Var}(u)\) implies \(\mathrm{Var}(u) = 1 - \lambda^2\), and averaging \(n\) independent draws leaves \((1-\lambda^2)/n\). That is the only term the averaging affects. The endowment is shared by the members of the cell and does not shrink: \(\mathrm{Var}(\bar e_{g,t}) = 1\) still holds, and the cell-level endowment remains an AR(1) in \(\rho\), so that \(\mathrm{Cov}(\bar e_{g,t}, \bar e_{g,t-k}) = \rho^{k}\) exactly as at the individual level.
The cell-level slope of \(\bar y_{g,t}\) on \(\bar y_{g,t-k}\) is a covariance divided by a variance. The covariance removes the noise, because the \(\bar u\) are uncorrelated across cohorts and with the endowment:
\[ \mathrm{Cov}(\bar y_{g,t},\, \bar y_{g,t-k}) = \mathrm{Cov}(\lambda \bar e_{g,t} + \bar u_{g,t},\; \lambda \bar e_{g,t-k} + \bar u_{g,t-k}) = \lambda^{2}\, \mathrm{Cov}(\bar e_{g,t}, \bar e_{g,t-k}) = \lambda^{2}\rho^{k} . \]
The denominator does not. The variance of an observed cell mean carries the surviving noise alongside the signal:
\[ \mathrm{Var}(\bar y_{g,t-k}) = \lambda^{2}\,\mathrm{Var}(\bar e_{g,t-k}) + \mathrm{Var}(\bar u_{g,t-k}) = \lambda^{2} + \frac{1-\lambda^{2}}{n} . \]
Dividing one by the other,
\[ b^{\text{grouped}} \;=\; \frac{\mathrm{Cov}(\bar y_{g,t}, \bar y_{g,t-k})}{\mathrm{Var}(\bar y_{g,t-k})} \;=\; \frac{\lambda^{2}\rho^{k}}{\lambda^{2} + (1-\lambda^{2})/n} \;\xrightarrow[\;n \to \infty\;]{}\; \rho^{k} . \]
The loading thus appears in both numerator and denominator, and cell size determines how much of it cancels. The numerator is the same \(\lambda^2 \rho^k\) as at the individual level, so grouping contributes nothing there. What grouping does is shrink the noise term in the denominator toward zero, and once \((1-\lambda^2)/n\) is negligible the loading in the numerator is divided by the loading in the denominator. At \(n = 1\) nothing cancels and the expression reduces to \(\lambda^2 \rho^k\), the individual slope. The factor \(\lambda^2 / (\lambda^2 + (1-\lambda^2)/n)\) is the reliability of a cell mean as a measure of the endowment it proxies, and as cells fill up the idiosyncratic slippage between endowment and occupational title averages away and that reliability approaches one. The structural content is the same as that of the ratio test: both estimate \(\rho\) free of \(\lambda\), one by iterating the process forward, the other by averaging the noise down.2
Comparing the Two Routes Across the Border
Suppose the Protestant and Catholic sides of the border differed only in \(\lambda\), so that occupational titles tracked family standing more loosely on one side. Grouping would then shrink the measured gap, because the reliability factor tends to one on both sides and \(b^{\text{grouped}}_P / b^{\text{grouped}}_C \to 1\). A pure measurement explanation predicts that the two sides converge as grouping improves.
The estimates diverge instead. The individual father–son slopes are \(0.420\) on the Protestant side against \(0.485\) on the Catholic side, a ratio of \(0.87\). The grouped surname slopes are \(0.290\) against \(0.397\) for cohorts born before 1800, and \(0.347\) against \(0.595\) over the full genealogical window, in which cells are larger, giving ratios of \(0.73\) and \(0.58\). The full-window difference of \(0.248\) is statistically significant. The differential-persistence IV, which interacts a surname’s lagged rank with the instrumented Protestant share, is negative throughout, between \(-0.35\) and \(-0.90\) before 1800.
Read through the model, that is a statement about \(\rho\): Catholic endowments regressed to the mean more slowly. It sits awkwardly beside the ratio test, which yields \(b_2/b_1 = 0.69\) on both sides of the border and attributes the gap to \(\lambda\) instead.
Conclusion
The tension between the two routes cannot be resolved with these data. The ratio test is underpowered on the thinner Catholic side. The grouped estimator has a known weakness of its own: surnames correlate with place, trade and endogamy, so cell-level persistence can absorb group-level advantages that were never transmitted through the family, and there is no reason to expect such contamination to be symmetric across a confessional border.
What survives both routes is the sign of the difference. Whether the Protestant advantage lies in the loading or in the persistence, in allocation or in transmission, every estimator considered here places it on the same side of zero, in the nineteenth century and three centuries earlier.
Footnotes
Clark, G., & Cummins, N. (2015). Intergenerational wealth mobility in England, 1858–2012: Surnames and social mobility. The Economic Journal, 125(582).↩︎
Two approximations are involved. Surname cells pool kin rather than clones, so the shared endowment is a fraction of each member’s, and a 50-year window does not correspond to one generation. Both make the mapping from \(b^{\text{grouped}}\) to \(\rho^k\) approximate rather than exact, which is why I read the estimator comparatively rather than as a point estimate of \(\rho\).↩︎