From omitted variables to Oster’s cubic

How the data, delta, and the long-regression R² jointly restrict beta

Yes: R^2_{\mathrm{long}} is one of the inputs to the identified set. The observed data identify the short and medium regressions. They generally do not identify the coefficient \beta_{\mathrm{long}} in a regression that also includes unobserved controls. Oster’s restrictions give an equation linking a candidate long coefficient b, a selection parameter \delta, and a hypothetical R^2_{\mathrm{long}}. Even after fixing both sensitivity parameters, that equation can leave several admissible values of b.

The useful object is therefore a joint compatibility relation among (b,\delta,R^2_{\mathrm{long}}), conditional on the observed data and the model’s assumptions. Fixing (\delta,R^2_{\mathrm{long}}) takes a slice through that relation; the slice need not be a single point.

This document derives that relation step by step. It explains the starting point of Masten and Poirier (2026), Section I, using their notation and the numerical population in the three-coefficient companion. Everything below concerns population identification. Sampling uncertainty would be an additional issue.

1. Three regressions, only two of which we can observe

Let Y be the outcome, X the treatment, W_1 a vector of observed controls, and W_2 an unobserved control. Work with mean-zero variables so that intercepts disappear from the equations. Write the population regressions as

\begin{aligned} \text{Short:}\quad Y&=\beta_{\mathrm{short}}X+e_{\mathrm{short}},\\ \text{Medium:}\quad Y&=\beta_{\mathrm{med}}X+\gamma_{1,\mathrm{med}}'W_1+u,\\ \text{Long:}\quad Y&=\underbrace{\beta_{\mathrm{long}}}_{\text{target}}X +\gamma_{1,\mathrm{long}}'W_1 +\gamma_{2,\mathrm{long}}W_2+\varepsilon. \end{aligned}

Each residual is uncorrelated with the regressors in its own equation. These orthogonality conditions are properties of population OLS. Calling \beta_{\mathrm{long}} a causal effect additionally requires a causal justification for the long specification; the algebra here concerns its regression coefficient.

Regression Regressors besides the intercept Identified from (Y,X,W_1)?
Short X \beta_{\mathrm{short}}, R^2_{\mathrm{short}}
Medium X,W_1 \beta_{\mathrm{med}}, \gamma_{1,\mathrm{med}}, R^2_{\mathrm{med}}
Long X,W_1,W_2 Generally neither \beta_{\mathrm{long}} nor R^2_{\mathrm{long}}

A fully specified distribution of (Y,X,W_1,W_2) has a unique long-regression coefficient when the regressors have a positive definite covariance matrix. The identification problem arises because we observe only (Y,X,W_1): different distributions involving W_2 can agree on everything we observe and have different long coefficients.

To keep the main derivation in the usual case, suppose \beta_{\mathrm{short}}\ne\beta_{\mathrm{med}} and choose R^2_{\mathrm{long}}>R^2_{\mathrm{med}}. We discuss the zero-bias baseline below.

2. What assumptions are doing the work?

Masten and Poirier maintain Oster’s selection framework. Their Assumptions 1–4 require:

  1. Finite second moments and no relevant perfect collinearity. The covariance matrices of (Y,X,W_1) and of the long regressors (X,W_1,W_2) are positive definite. A perfect long fit, R^2_{\mathrm{long}}=1, is allowed.
  2. A specified proportional-selection parameter \delta. We will define this precisely in the next section.
  3. Nonzero long coefficients on both control groups: \gamma_{1,\mathrm{long}}\ne0 and \gamma_{2,\mathrm{long}}\ne0. This ensures that the two control indices have positive variances.
  4. Orthogonality of the observed and omitted controls: \operatorname{Cov}(W_1,W_2)=0.

The last restriction does not require W_2 to be uncorrelated with X. That correlation is what can generate omitted-variable bias. The phrase “exogenous controls” in this paper refers to the stated covariance restriction between the two control groups; it does not by itself supply a causal identification argument.

We also specify the long-regression fit,

R^2_{\mathrm{long}}\in(R^2_{\mathrm{med}},1].

This value is an assumption about the unobserved regression. Oster calls it R_{\max}, without a superscript 2: her R_{\max} already denotes an R-squared value. It is not an R-squared estimated from the available controls, and the notation does not automatically make it an upper bound. Fixing an exact value and imposing an upper bound are different restrictions, as Section 10 explains.

3. Delta uses the long-regression control weights

Define two scalar indices, both measured in outcome units:

A=\gamma_{1,\mathrm{long}}'W_1, \qquad H=\gamma_{2,\mathrm{long}}W_2.

The two selection coefficients are slopes from separate regressions of X on these indices:

S_{\mathrm{obs}}= \frac{\operatorname{Cov}(X,A)}{\operatorname{Var}(A)}, \qquad S_{\mathrm{unobs}}= \frac{\operatorname{Cov}(X,H)}{\operatorname{Var}(H)}.

Oster’s selection restriction is

\boxed{\delta S_{\mathrm{obs}}=S_{\mathrm{unobs}}.}

When S_{\mathrm{obs}}\ne0, this says \delta=S_{\mathrm{unobs}}/S_{\mathrm{obs}}. The equality formulation is the paper’s formal Assumption 2 and remains meaningful when that selection slope is zero. Delta is signed; it is neither a correlation nor a ratio of R-squared values.

The observable-selection index uses \gamma_{1,\mathrm{long}}, not the estimated medium-regression weights \gamma_{1,\mathrm{med}}. The variables W_1 are observed, but their long-regression weights are not. Consequently, even the observable-selection slope in this definition is not a known constant before we choose a candidate long coefficient. This is the source of the feedback that produces the cubic.

4. Separate the treatment into a predicted part and a residual

Regress X on W_1 and write

X=P+Z, \qquad P=\pi_1'W_1, \qquad Z=X^{\perp W_1}.

Here P is the part of X linearly predicted by the observed controls, while Z is the remaining variation. Thus \operatorname{Cov}(Z,W_1)=0.

Let M=\gamma_{1,\mathrm{med}}'W_1 be the known medium-regression index, and abbreviate the observed moments by

\begin{aligned} v&=\operatorname{Var}(Z)>0, &q&=\operatorname{Var}(P),\\ s&=\operatorname{Var}(M), &c&=\operatorname{Cov}(X,M)=\operatorname{Cov}(P,M). \end{aligned}

For a candidate long coefficient b, define its bias relative to the medium coefficient as

\boxed{B=\beta_{\mathrm{med}}-b.}

This is a candidate value of omitted-variable bias. It can be positive, negative, or zero.

5. A candidate beta determines the observed-control index

Take the long equation at coefficient b,

Y=bX+A+H+\varepsilon.

Because H is uncorrelated with W_1, its projection onto (X,W_1) must be

\operatorname{Proj}(H\mid X,W_1)=BZ=B(X-P).

To see the coefficient B, residualize X on W_1. The omitted-variable bias formula then gives

\beta_{\mathrm{med}}-b =\frac{\operatorname{Cov}(Z,H)}{\operatorname{Var}(Z)}=B.

Substituting the projection of H into the long equation shows what the medium regression sees:

Y=(b+B)X+(A-BP)+u.

Matching its coefficients to the observed medium regression gives

\gamma_{1,\mathrm{long}} =\gamma_{1,\mathrm{med}}+B\pi_1, \qquad \boxed{A=M+BP.}

As we change b, we change B, and therefore change the implied long coefficients on the observed controls. Hence

\boxed{ \begin{aligned} \operatorname{Cov}(X,A)&=c+Bq,\\ \operatorname{Var}(A)&=s+2Bc+B^2q. \end{aligned}}

The covariance is linear in B; the variance is quadratic. The observed-selection slope is therefore

S_{\mathrm{obs}}(B) =\frac{c+Bq}{s+2Bc+B^2q}.

This dependence on B matters even though the medium-regression moments s,c,q are all known.

6. The long R-squared determines the omitted index’s variance

The covariance of the omitted index with X is already pinned down by the candidate bias:

\operatorname{Cov}(X,H) =\operatorname{Cov}(Z,H)=Bv,

because P is a linear combination of W_1, which is uncorrelated with H.

Now define the additional explained variance attributed to going from the medium to the long regression:

\boxed{ D=(R^2_{\mathrm{long}}-R^2_{\mathrm{med}}) \operatorname{Var}(Y)>0. }

The residual variance falls by exactly D:

D=\operatorname{Var}(u)-\operatorname{Var}(\varepsilon).

From the two outcome equations and A=M+BP,

u=H-BZ+\varepsilon.

The long residual \varepsilon is uncorrelated with H and Z. Therefore

\begin{aligned} D &=\operatorname{Var}(H-BZ)\\ &=\operatorname{Var}(H)+B^2v -2B\operatorname{Cov}(Z,H)\\ &=\operatorname{Var}(H)-B^2v. \end{aligned}

Solving gives the second crucial identity:

\boxed{\operatorname{Var}(H)=D+B^2v.}

The R-squared gain is not the entire variance of the omitted index. Its component BZ was already picked up by the medium regression through X. Only the remaining component H-BZ adds explanatory power when we include the omitted control. Its variance is D.

We now have the omitted-selection slope:

S_{\mathrm{unobs}}(B,D)=\frac{Bv}{D+B^2v}.

Thus R^2_{\mathrm{long}} matters directly: it fixes D, which enters the denominator of selection on unobservables.

7. Substitute the two slopes: the cubic appears

Put Sections 5 and 6 into the proportional-selection restriction:

\delta\frac{c+Bq}{s+2Bc+B^2q} =\frac{Bv}{D+B^2v}.

For admissible candidates both variance denominators are positive. Cross-multiply:

\boxed{ -Bv(s+2Bc+B^2q) +\delta(D+B^2v)(c+Bq)=0. }

Now expand each product:

\begin{aligned} 0={}&-Bvs-2B^2vc-B^3vq\\ &+\delta Dc+\delta BDq+\delta B^2vc+\delta B^3vq. \end{aligned}

Collect the powers of B:

\boxed{ vq(\delta-1)B^3 +vc(\delta-2)B^2 +(\delta Dq-vs)B +\delta Dc=0. }

This is equation (4) in Masten and Poirier, written in compact notation:

\begin{aligned} f_0(B)&=-Bv(s+2Bc+B^2q),\\ f_1(B,D)&=(D+B^2v)(c+Bq),\\ f(B,\delta,D)&=f_0(B)+\delta f_1(B,D). \end{aligned}

Why degree three? A candidate bias multiplies a quadratic variance on the first side, and a quadratic variance multiplies a linear covariance on the second. Those products generate B^3. Since B=\beta_{\mathrm{med}}-b is an affine transformation, the same equation is also a polynomial of degree at most three in b.

In our usual case, \beta_{\mathrm{short}}\ne\beta_{\mathrm{med}} implies c\ne0 and q>0. The equation is genuinely cubic for \delta\ne1. At \delta=1 its cubic terms cancel and it is quadratic. Setting \delta=1 alone still need not select a unique coefficient.

8. Exactly what determines the identified set?

The observed distribution supplies all the entries in the first two rows below. The remaining rows are restrictions imposed on possible completions involving the unobserved control.

Input Meaning Where it enters
\beta_{\mathrm{med}},v,s,c,q Medium coefficient; residual treatment variance; variance of each observed index; their treatment covariance B=\beta_{\mathrm{med}}-b and the polynomial coefficients
R^2_{\mathrm{med}},\operatorname{Var}(Y) Medium fit and outcome variance Converts the assumed long fit into D
\delta Assumed proportionality between the two selection slopes Multiplies f_1
R^2_{\mathrm{long}} Assumed fit of the regression including W_2 D=(R^2_{\mathrm{long}}-R^2_{\mathrm{med}})\operatorname{Var}(Y)
Assumptions 1, 3, and 4 Moment regularity, nonzero indices, orthogonal control groups Justifies the derivation and determines admissibility

For fixed observed data, \delta, and R^2_{\mathrm{long}}, first solve the polynomial for all real roots. Then exclude candidates that make the entire observed-control coefficient vector zero:

\mathcal B_{\mathrm{A3fail}} =\left\{b: \gamma_{1,\mathrm{med}}+(\beta_{\mathrm{med}}-b)\pi_1=0 \right\}.

Equivalently, keep only candidates for which s+2Bc+B^2q>0. Masten and Poirier’s sharp identified set is

\boxed{ \mathcal I_\beta(\delta,R^2_{\mathrm{long}}) =\left\{b\in\mathbb R: f(\beta_{\mathrm{med}}-b,\delta,D)=0 \right\}\setminus\mathcal B_{\mathrm{A3fail}}. }

“Sharp” means that every retained value can actually arise in some completion of the same observed population satisfying the restrictions. The polynomial is therefore more than a necessary algebraic check. Their formal sharpness result is Theorem S1 in Supplemental Appendix E.

Under the usual case considered here, this set has at most three elements, and at most two at \delta=1. It can contain a single value; partial identification does not mean that multiplicity must occur for every choice of sensitivity parameters. For example, \delta=0 gives B=0 after inadmissible zero-index roots are removed, so \mathcal I_\beta(0,R^2_{\mathrm{long}})=\{\beta_{\mathrm{med}}\}.

Which regression summaries do I need in practice?

You can recover the scalar inputs without reporting every control coefficient. Let t=\operatorname{Var}(X). Then

\boxed{ \begin{aligned} q&=t-v,\\ c&=(\beta_{\mathrm{short}}-\beta_{\mathrm{med}})t,\\ s&=(R^2_{\mathrm{med}}-R^2_{\mathrm{short}})\operatorname{Var}(Y) +\frac{c^2}{t}. \end{aligned}}

For the second identity, take the covariance of the medium outcome equation with X. For the third, subtract the short fitted variance from the medium fitted variance:

\begin{aligned} R^2_{\mathrm{med}}\operatorname{Var}(Y) &=\beta_{\mathrm{med}}^2t+2\beta_{\mathrm{med}}c+s,\\ R^2_{\mathrm{short}}\operatorname{Var}(Y) &=\beta_{\mathrm{short}}^2t =\beta_{\mathrm{med}}^2t+2\beta_{\mathrm{med}}c+c^2/t. \end{aligned}

So an equivalent list of observed scalar summaries is

\beta_{\mathrm{short}},\ \beta_{\mathrm{med}},\ R^2_{\mathrm{short}},\ R^2_{\mathrm{med}},\ \operatorname{Var}(Y),\ \operatorname{Var}(X),\ \operatorname{Var}(X^{\perp W_1}).

The last quantity can instead be supplied by the R-squared from regressing X on W_1, since v=t(1-R^2_{X\mid W_1}). The treatment coefficients and outcome R-squared values alone are generally insufficient for the unrestricted cubic. Oster highlights this extra residual-treatment-variance input in Section 3.3.

Why a retained root corresponds to a possible omitted variable

Here is an explicit construction, which also shows why fixing the observed population does not force a single b. Let

S=\operatorname{Var}(u)=(1-R^2_{\mathrm{med}})\operatorname{Var}(Y).

Since the long fit is at most one, 0<D\le S. Introduce a mean-zero random variable T, independent of the observed variables, with

\operatorname{Var}(T)=D(1-D/S).

For a candidate B, set

\eta=(D/S)u+T, \qquad H=BZ+\eta, \qquad \varepsilon=u-\eta.

Then \operatorname{Var}(\eta)=\operatorname{Cov}(u,\eta)=D. The medium residual u is uncorrelated with X,W_1, so this construction yields

\begin{aligned} \operatorname{Cov}(W_1,H)&=0,\\ \operatorname{Cov}(X,H)&=Bv,\\ \operatorname{Var}(H)&=B^2v+D,\\ \operatorname{Var}(\varepsilon)&=S-D. \end{aligned}

The residual \varepsilon is uncorrelated with X,W_1,H, and the outcome satisfies

Y=(\beta_{\mathrm{med}}-B)X+(M+BP)+H+\varepsilon.

Choose W_2=H and its long coefficient equal to one. Its residual variance after projecting on (X,W_1) is D>0, so it does not create perfect collinearity. If the observed index is nonzero and B solves the polynomial, this long regression satisfies every stated restriction, including the specified \delta and R^2_{\mathrm{long}}.

Different retained roots give different hypothetical omitted variables while preserving the same distribution of (Y,X,W_1).

9. Beta and delta are linked, but not one-to-one

At a fixed long R-squared, rearranging the compatibility equation gives

\boxed{ \delta(B;D) =\frac{Bv(s+2Bc+B^2q)}{(D+B^2v)(c+Bq)}, \qquad b=\beta_{\mathrm{med}}-B. }

This formula applies where c+Bq\ne0. In our usual case c\ne0, the candidate B=-c/q cannot satisfy the selection equality for any finite \delta: its observed-selection slope is zero while its omitted-selection slope is nonzero.

The function \delta(B;D) need not be monotone. Several different biases can therefore give the same \delta:

\text{fix }(b,R^2_{\mathrm{long}}) \ \Longrightarrow\ \text{at most one compatible finite }\delta,

\text{fix }(\delta,R^2_{\mathrm{long}}) \ \Longrightarrow\ \text{possibly several compatible }b\text{'s}.

The first statement concerns a feasible candidate under the regularity conditions above; it does not guarantee that every b has a compatible finite delta.

Thus it is useful to think of beta and delta as jointly constrained. The observed data do not estimate an unrestricted pair (\beta,\delta) uniquely. Nor does choosing delta generally identify one beta. The relation also depends on the assumed long R-squared.

The same observed population, different slices

Use the population from the companion document. Let U,V,Z,E be independent and mean-zero, with variances 1,3,1,12, respectively, and define

X=U+Z,\qquad Y=2U+V+Z+E,\qquad W_1=(U,V)'.

Its observed regressions give

\beta_{\mathrm{short}}=3/2,\quad \beta_{\mathrm{med}}=1,\quad R^2_{\mathrm{short}}=9/40,\quad R^2_{\mathrm{med}}=2/5,

\operatorname{Var}(Y)=20,\quad v=1,\quad s=4,\quad c=1,\quad q=1.

Here B=1-b, D=20(R^2_{\mathrm{long}}-2/5), and

-B(4+2B+B^2)+\delta(D+B^2)(1+B)=0.

The long observed-control vector is (1+B,1)', which can never be the zero vector, so there are no zero-index exclusions.

At R^2_{\mathrm{long}}=1 and \delta=1/2, the equation is

(B+3)(B+2)(B-2)=0, \qquad \mathcal I_\beta(1/2,1)=\{-1,3,4\}.

Now keep \delta=1/2 and change only the long fit. At R^2_{\mathrm{long}}=0.7, D=6, and the polynomial becomes

B^3+3B^2+2B-6 =(B-1)(B^2+4B+6)=0.

The quadratic factor has no real roots, so B=1 is the only real solution and b=0.

Fixed delta Assumed R^2_{\mathrm{long}} D Identified set for beta
1/2 0.52 2.4 \{0.6463\ldots\}
1/2 0.70 6 \{0\}
1/2 1.00 12 \{-1,3,4\}

These rows have exactly the same observed data and delta. Changing the assumed explanatory power of the omitted control changes both the values and the number of compatible long coefficients.

Figure 1: Each curve is a compatibility relation at one assumed long R-squared. A vertical slice at delta = 1/2 contains every marked beta, rather than a chosen branch.

10. What if delta and the long R-squared are bounded?

For clarity, make the data dependence explicit and write the joint feasible relation as

\mathcal J =\left\{(b,\delta,r): r\in(R^2_{\mathrm{med}},1],\quad b\in\mathcal I_\beta(\delta,r) \right\},

where r denotes the R-squared value itself. The observed moments and Assumptions 1, 3, and 4 are held fixed.

If your assumptions permit \delta\in\Delta and R^2_{\mathrm{long}}\in\mathcal R, the beta set is

\boxed{ \mathcal I_\beta(\Delta,\mathcal R) =\bigcup_{\delta\in\Delta}\ \bigcup_{r\in\mathcal R} \mathcal I_\beta(\delta,r). }

Examples include \Delta=[-\bar\delta,\bar\delta] for a bound on |\delta|, or [0,\bar\delta] if you additionally require nonnegative relative selection. A bound R^2_{\mathrm{long}}\le\bar r means allowing all permitted long fits up to \bar r, not just solving at R^2_{\mathrm{long}}=\bar r.

The upper bound on the long fit must remain at most one. The endpoint R^2_{\mathrm{long}}=R^2_{\mathrm{med}} has zero omitted-variable bias under the regular regression conditions, so it is handled separately rather than by the positive-D derivation above.

The fixed-parameter set may have only a few points, while a union over sensitivity values can contain intervals, gaps, and unbounded components. It need not be the interval between a baseline estimate and one selected adjusted estimate.

An additional substantive assumption, such as |B|\le L, restricts the set further by intersecting it with [\beta_{\mathrm{med}}-L,\beta_{\mathrm{med}}+L]. Selecting a root merely because it is closest to the medium coefficient is not implied by the four assumptions we used.

11. Why a familiar single adjustment does not replace this relation

Oster also gives a simpler formula under stronger conditions:

\beta_{\mathrm{long}} =\beta_{\mathrm{med}} -(\beta_{\mathrm{short}}-\beta_{\mathrm{med}}) \frac{R^2_{\mathrm{long}}-R^2_{\mathrm{med}}} {R^2_{\mathrm{med}}-R^2_{\mathrm{short}}}.

Its point-identification result requires both \delta=1 and proportional control coefficients: the outcome-control coefficient vector must be proportional to \pi_1, the vector predicting X. Under the orthogonal-controls assumption, this can be expressed as \gamma_{1,\mathrm{med}}=C\pi_1 for a scalar C. This is Oster’s restricted analysis in Section 3.2, rather than the unrestricted cubic in Section 3.3.

In our numerical population, \gamma_{1,\mathrm{med}}=(1,1)' while \pi_1=(1,0)', so proportionality fails. At \delta=1 and R^2_{\mathrm{long}}=1, the unrestricted equation is

-B^2+8B+12=0,

giving

\mathcal I_\beta(1,1) =\{-3-\sqrt{28},\,-3+\sqrt{28}\} \approx\{-8.292,\,2.292\}.

The simpler formula instead returns -5/7, which is not in this set. Multiplying that simpler formula’s bias correction by an arbitrary delta also does not reproduce the unrestricted compatibility relation. The polynomial derivation tells us precisely which restrictions support a proposed adjustment.

12. How this sets up Masten and Poirier’s question

Their starting point is this entire identified relation. At a fixed R^2_{\mathrm{long}}, a plot with delta on the horizontal axis and beta on the vertical axis shows all compatible pairs. A vertical line can intersect several branches, as the figure above illustrates.

Asking “which delta is compatible with b=0?” fixes beta and solves for delta. This gives the explain-away calculation. Asking “which delta first permits a beta of the opposite sign?” searches across all compatible branches. Those questions can have different answers because the relation is nonlinear and can have disconnected branches; there is no general requirement that the first opposite-sign coefficient be reached by continuously moving one chosen beta through zero.

For this method, the precise statement to keep in mind is:

The observed data identify the short and medium regressions. Given those data and the maintained restrictions, candidate long coefficients are jointly constrained with delta and the long-regression R-squared. Fixing both sensitivity parameters generally identifies a set of possible coefficients, and sometimes a single coefficient.

Sources and notation

The derivation uses Oster’s supplied 2017 text, Sections 3.1–3.3, especially Proposition 2, and Masten and Poirier’s supplied 2026 text, Sections I.A–I.D, especially equations (2)–(5). Their paper cites Oster as 2019a; the supplied Oster file is the earlier online publication of the article with DOI 10.1080/07350015.2016.1227711. We use Masten and Poirier’s short/medium/long labels throughout.

The numerical population is constructed for the three-coefficient companion; it is not an empirical example from either paper. The explicit completion in Section 8 is a construction for this explanation. The figure code checks the numerical long regressions, their R-squared values, and their selection restrictions using exact covariance matrices.

Render with quarto render oster-cubic-and-identification.qmd. The executable figure uses base R and requires Quarto’s knitr and rmarkdown packages.