09: Endogeneity

Endogeneity

Transcript

This slide gives you the definition that ties the rest of the lecture together. Look at the first expression. It says that the expected value of the error, called u, conditional on values of all k explanatory variables, x one through x k, is not zero. “Conditional on” means that we compare observations with the same included x values. In a correctly specified causal OLS model, once those x values are fixed, the remaining unobserved determinants in u should not be systematically positive or negative. Endogeneity is the failure of that zero-conditional-mean requirement.

The red phrase “for whatever reason” is important. We call a particular explanatory variable endogenous whenever it is correlated with the error term. The definition does not depend on how that correlation arose. Once x k and u move together, OLS cannot hold the unobserved component fixed while changing x k. Its coefficient mixes the effect of x k with the effects of factors hidden in u, so the usual causal interpretation is not justified. The issue is systematic correlation, not simply that the error is large or that an observation is unusual.

At the bottom you see the four mechanisms we will study. An omitted variable affects the outcome, is left in u, and is related to an included regressor. Selection occurs when people, firms, doctors, or governments choose treatment using factors that also affect the outcome. Reverse causality means the proposed outcome also helps determine the explanatory variable. Measurement error can put the same recording error into a regressor and the composite disturbance. These categories can overlap. Selection is often an omitted-variable problem, and reverse causality is one way a regressor becomes correlated with the error. As you move through the deck, keep asking one question: what exactly creates the relationship between an included x and u?

Endogeneity

E[u\mid x_1,\dots,x_k] \ne 0

The zero-conditional-mean assumption fails: the error term is systematically related to at least one explanatory variable.


Endogenous independent variable

If an explanatory variable x_k is, for whatever reason, correlated with the error term, we call x_k an endogenous explanatory variable.


Sources of endogeneity

  • Omitted variable
  • Selection
  • Reverse causality
  • Measurement error

Omitted Variable

Transcript

Use these two equations to see exactly how an omitted variable creates endogeneity. The top line is the true population model. The outcome is log wage, so beta one is the change in log wage associated with one more unit of education, holding experience and ability fixed. Beta zero is the intercept. Beta two is the partial effect of experience, beta three is the partial effect of ability, and u contains the remaining determinants of log wages. For OLS to recover these causal coefficients, u must have conditional mean zero given education, experience, and ability.

Now look at the model we can actually estimate if ability is unavailable. It includes education and experience but omits ability. Omitting it does not make its wage contribution disappear. On the right of the second line, the new composite error v equals the original u plus beta three times ability. That algebra is why an omitted determinant belongs in the new error term.

The education regressor is endogenous in this shorter equation if education is related to ability. In that case, education is correlated with a component of v. OLS cannot compare people who differ in education while holding ability fixed, because ability is unobserved and tends to change with education. The estimated coefficient on education then combines the return to education with some of ability’s association with wages.

You need two relationships to sign the omitted-variable bias. Ask first whether ability raises or lowers wages, which is the sign of beta three. Then ask whether ability and education are positively or negatively related after accounting for experience. If ability raises wages and is positively related to education, the education coefficient is biased upward. If one of those relationships is negative and the other positive, the bias is downward. Merely identifying an omitted variable tells you that bias is possible; it does not by itself tell you the direction. We now use sensor adoption to show why purposeful selection creates exactly this structure.

True Model

\log(wage) = \beta_0 + \beta_1 educ + \beta_2 exper + \beta_3 ability + u


Incorrectly specified (your) model

\log(wage) = \beta_0 + \beta_1 educ + \beta_2 exper + v, \qquad v=u+\beta_3 ability

Bias from self-selection

Transcript

This example turns the definition of endogeneity into a research-design question. We want to know whether a soil-moisture sensor reduces a farmer’s irrigation. The data are observational, not experimental, so no researcher randomly assigns sensors. Each farmer decides whether to adopt one. That distinction matters because adopters and non-adopters may differ before the sensor has any effect.

Look at the model of interest. Irrigation is the amount of water applied by the farmer. Beta zero is the expected irrigation amount for the reference group, farmers with sensor equal to zero, subject to the interpretation of the error. Sensor is an indicator: it equals one for an adopter and zero for a non-adopter. Beta one is therefore the difference in expected irrigation associated with adoption. If lower irrigation means water savings, a beneficial sensor effect would make beta one negative. The error u collects every other determinant of irrigation omitted from this very simple equation, such as soil, crop choice, weather, water prices, farm management, and conservation preferences.

For beta one to be a causal sensor effect, sensor status must be unrelated to those unobserved determinants. That is the question in red at the bottom: is sensor correlated with u? A raw regression can calculate an association even when the answer is yes, but the association then mixes the technology’s effect with pre-existing differences between adopters and non-adopters. Before estimating anything, describe the process that generated the explanatory variable. Here that means opening up the farmer’s adoption decision, which we do on the next tab.

Research Question

Does a soil moisture sensor reduce water use for farmers?


Data

Observational (non-experimental) data on soil moisture sensor adoption and irrigation amount


Model of interest

irrigation = \beta_0 + \beta_1 sensor + u

  • irrigation: the amount of irrigation by the farmer
  • sensor: indicator equal to 1 if the farmer adopted a soil-moisture sensor


Question

Is sensor endogenous (is sensor correlated with the error term)?

Transcript

Treat adoption as an outcome of a separate decision, not as if it were assigned by a coin flip. The displayed selection equation says sensor equals alpha zero plus alpha one times x one, continuing through alpha k times x k, plus v. Alpha zero is a baseline adoption propensity. The x variables represent factors farmers use when deciding whether a sensor is worthwhile, and each alpha describes how one of those factors changes adoption. Because sensor is actually a zero-one variable, this linear equation is a stylized way to organize the determinants of adoption, not a claim that the decision is literally continuous. The disturbance v contains additional influences on adoption that are not listed.

For the first question, plausible x variables include soil type, field size, crop mix, irrigation technology, water price, prior drought experience, access to credit, technical ability, and conservation preferences. Farmers can rationally use all of this information. The econometric problem appears when you answer the second question. Many of the same variables also determine irrigation demand. Sandy soil may increase both the value of timely moisture information and the amount of water a field needs. High water prices may encourage sensor adoption while independently reducing water use. Conservation-minded farmers may adopt and may also irrigate less even without a sensor.

Those shared causes make adopters and non-adopters systematically different before treatment. If we observe and accurately model every common cause, we may condition on it. If one remains unobserved, it enters the outcome error and links sensor status to that error. The next tab takes one common factor, soil type, and traces that link explicitly.

Farmers do not adopt soil-moisture sensors randomly. They use available information to decide whether adoption would benefit them.


Adoption (selection) equation

sensor = \alpha_0 + \alpha_1 x_1 + \dots + \alpha_k x_k + v


Question

What variables might farmers consider when deciding whether to adopt a soil-moisture sensor?


Question

Do any of the variables listed above also affect irrigation demand?

Transcript

This tab isolates soil quality or soil type as one concrete common factor. The first bullet says farmers with sandy fields may be more likely to adopt a soil-moisture sensor. Sandy soil drains quickly, so knowing current moisture can be especially useful for timing irrigation. The second bullet says sandy fields may also require more irrigation because they retain less water. Soil type therefore belongs in both behavioral stories: it helps determine the adoption decision and it directly affects the irrigation outcome.

Now follow the two paths summarized under “Key.” Sensor is a function of soil type, meaning the distribution of sensor status changes across soil types. Irrigation is also a function of soil type. If we do not accurately control for soil type in the irrigation regression, its effect is absorbed by the error term. Sensor status is then correlated with a factor inside that error, which is precisely endogeneity.

This matters even if farmers decide rationally and all recorded data are entered correctly. Self-selection itself produces the problem. It also shows why comparing average irrigation for adopters and non-adopters is not enough. If adopters disproportionately farm sandy soil, they might irrigate more despite a water-saving sensor. The raw comparison could understate the reduction, show no reduction, or even have the opposite sign. Conversely, a different common factor could make adopters irrigate less before adoption and exaggerate the sensor’s benefit. Flip to the next tab to see the soil path written inside the regression error.

Example

Soil quality/type (hard to accurately measure)

  • Farmers whose fields are sandy may be more likely to adopt a soil-moisture sensor.
  • Sandy fields may require more irrigation.


Key

Soil quality or type affects both sensor adoption and irrigation.

  • sensor is a function of soil quality/type
  • irrigation is also a function of soil quality/type, which enters the error term when it is not controlled for
Transcript

Here is the common-factor argument written directly in the outcome equation. Read the line as irrigation equals beta zero, plus beta one times sensor, plus a parenthetical composite error. Beta one is the causal sensor effect we want. Inside the parentheses, beta sub s is the effect of soil type on irrigation, “soil type” is the omitted common factor, and v contains all other unobserved determinants of irrigation. The entire parenthetical expression is the error term in a regression that includes sensor but leaves soil type out.

Now connect this equation to the previous tab. Sensor adoption varies with soil type. Soil type also appears inside the composite error through beta sub s times soil type. Therefore sensor and the regression error vary together. OLS cannot separate irrigation changed by the sensor from irrigation associated with the different soils on which sensors are disproportionately adopted.

The callout’s conclusion is bias in the estimated impact of a sensor. Its direction is not determined by the word “selection.” It depends on how soil affects adoption and on the sign of beta sub s. In the sandy-soil story, sandy fields are more likely to adopt and require more irrigation, so the omitted factor pushes the sensor coefficient upward. If the true sensor effect is negative, that upward bias makes the estimate less negative and can even make it positive. With different relationships, the bias could go the other way. The decisive point is that adopters differ from non-adopters along a factor that also affects irrigation. The final tab asks when including that factor can remove this source of bias.

\begin{aligned} irrigation = \beta_0 + \beta_1 sensor + (\beta_s\,\text{soil type}+v) \end{aligned} where v includes all other unobserved determinants of irrigation.

Note

So, sensor and the error term in the irrigation model are correlated through soil type, leading to biased estimation of the impact of a sensor.

Transcript

The callout states the main lesson from this example: selection bias is a form of omitted-variable bias when the factors driving selection also affect the outcome and are excluded from the outcome equation. That observation suggests a possible remedy. If you can accurately measure the common factors, include them explicitly rather than leaving them in the error.

Look at the revised equation. Irrigation equals beta zero, plus beta one times sensor, plus beta sub s times soil type, plus v. Once soil type is included, beta one compares adopters and non-adopters at the same included value of soil type. Beta sub s captures systematic irrigation differences associated with soil, and v now contains the other unobserved determinants. Soil’s contribution is no longer automatically part of the error, so this particular path from sensor to the error is closed.

The qualifications beneath the equation do real work. Soil type must be measured accurately. A crude category may leave relevant soil variation in v, while measurement error in the control may fail to remove the confounding completely. The model must also represent the relationship adequately, for example if soil affects irrigation nonlinearly or changes the sensor effect. Most importantly, there must be no other common causes of adoption and irrigation left in v. Controlling for soil does not solve confounding from crop choice, management skill, conservation preferences, or water price if those still link sensor to the error.

So “add controls” is not a mechanical guarantee of causality. You need a substantive conditional-independence argument: after conditioning on the included common factors, sensor adoption is as good as randomly assigned with respect to the remaining potential irrigation outcomes. If that argument is credible, beta one can be interpreted as the sensor effect conditional on those controls. If important common factors are unmeasured, we need a different identification strategy. The next section turns to another source of endogeneity, reverse causality.

Note

Selection bias is a form of omitted variable bias.


If you accurately measure the common factors in the two equations, you can simply include them explicitly in the main model.

For example,

\begin{aligned} irrigation = \beta_0 + \beta_1 sensor + \beta_s\,\text{soil type} + v \end{aligned}

If soil type is measured accurately and there are no other confounders, including it removes this source of correlation between sensor adoption and the error term.

Reverse Causality

Transcript

Now apply the same diagnostic habit to medical treatment. The research question is causal: does this treatment improve health? But the red word “cross-sectional” tells you what the data can directly compare. We observe different patients’ current treatment status and current health at one point in time. We do not observe a before-and-after change for each patient, and doctors, rather than a randomized experiment, choose who receives treatment.

In the displayed model, health is the patient’s current health measure. Treatment is an indicator equal to one for a treated patient and zero for an untreated patient. Beta zero describes expected health for the untreated reference group under the model, and beta one is the treated-versus-untreated difference we would like to interpret as the treatment’s effect. The error u contains all other determinants of current health not included in this simple equation, such as underlying severity, prognosis, age, behavior, access to care, and other medical conditions.

The endogeneity question at the bottom asks whether treatment status is correlated with those omitted health determinants. That is highly plausible because doctors use health information when prescribing treatment. If sicker patients are more likely to be treated, the treated group can have worse current health even when the treatment helps them. A negative estimated beta one could then reflect who was treated, not harm caused by treatment. Likewise, a favorable raw comparison need not isolate the treatment effect if healthier or better-connected patients obtain treatment more often.

Cross-sectional OLS gives an association between currently treated and untreated people. It does not reconstruct what the same treated patients’ health would have been without treatment. To assess whether the association is causal, we must model or otherwise address the treatment decision. Flip to the next tab to make that selection process explicit.

Research Question

Does a particular type of medical treatment improve health?


Data

Observational (non-experimental) cross-sectional data on a particular type of medical treatment and current health. Treatment is chosen by doctors rather than randomly assigned.


Model

health = \beta_0 + \beta_1 treatment + u

  • health: measure of a patient’s current health
  • treatment: indicator equal to 1 if the patient receives treatment

This model compares the health of treated and untreated patients at a point in time; it does not make a before-and-after comparison.


Question

Is treatment endogenous? (Is treatment correlated with the error term?)

Transcript

The first question asks how doctors decide whether to place patients under treatment. Open the answer and you see the essential input: their patients’ health conditions. In practice doctors may use symptoms, test results, medical history, expected prognosis, contraindications, and many other variables, but the key point is that assignment responds to information about health. Treatment is therefore selected, not randomly assigned.

The stylized selection equation says treatment equals alpha zero plus alpha one times health plus v. Alpha zero is a baseline treatment propensity, alpha one describes how the health measure changes the probability or level of treatment, and v contains other determinants of the doctor’s decision. Since treatment is a zero-one indicator, think of this equation as a compact representation of selection. The sign of alpha one depends on how health is coded. If larger health values mean better health and less healthy people are more likely to be treated, alpha one is negative. If the variable were severity, with larger values meaning worse illness, the corresponding coefficient would be positive.

Treatment status consequently carries information about health. It can also carry information about health determinants that the analyst does not observe but the doctor does. That makes treated and untreated patients systematically incomparable in the simple outcome regression.

Be careful about timing because it determines the most precise label. If doctors choose current treatment in response to health measured before treatment, this is selection on baseline health, which is an omitted-confounder problem in a regression of later health. If the slide’s health and treatment are contemporaneous and each helps determine the other, the two variables form a simultaneous feedback system. The next tab writes that contemporaneous case and shows why it is called reverse causality.

Question

How do doctors decide whether to place their patients under medical treatment?

Answer Their patients’ health conditions.


Selection (treatment decision) model

treatment = \alpha_0 + \alpha_1 health + v

Less healthy people are more likely to be treated.

Transcript

Look at the two equations as a system. The first is the health outcome equation: health equals beta zero plus beta one times treatment plus u. Beta one represents the effect of treatment on health, and u is an unobserved health shock or other omitted health determinant. The second is the treatment-decision equation: treatment equals alpha zero plus alpha one times health plus v. Alpha one represents the response of treatment assignment to health, and v contains other determinants of treatment.

Now trace the consequence line by line. Suppose u changes, perhaps because an unobserved illness worsens. That shock directly changes health through the first equation. Health then changes treatment through the second equation. Treatment status therefore responds, indirectly, to a component of u. In probabilistic terms, treatment and u are correlated. That violates zero conditional mean in the health equation, so an OLS regression of health on treatment does not isolate beta one.

This is reverse causality because the causal arrow proposed in the research question runs from treatment to health, while another arrow runs from health back to treatment. The feedback can make an effective treatment look harmful when poor health triggers treatment. Solving the system also makes clear that changing treatment and observing health are not separate one-way events when both are determined together.

The final sentence on screen gives an important timing qualification. If doctors choose treatment using pre-treatment health, then later treatment cannot literally cause that earlier health. The problem is more precisely selection on baseline health, and baseline health belongs among the confounders. If current health and current treatment are jointly determined, it is genuine simultaneity or reverse causality. Under either description, the single health equation does not make treatment exogenous, and beta one is not automatically causal. The nested example now shows the same policy-response logic in environmental enforcement.

Consequence

\begin{aligned} health &= \beta_0 + \beta_1 treatment + u,\\ treatment &= \alpha_0 + \alpha_1 health + v. \end{aligned}

A shock in u changes health, and health in turn affects treatment. Consequently, treatment is correlated with u and is endogenous in the health equation.


Reverse Causality

This endogeneity problem is called reverse causality: the proposed outcome also affects the explanatory variable of interest. If treatment is chosen using pre-treatment health instead, the problem is more precisely described as selection on baseline health.

Transcript

This context illustrates reverse causality when the explanatory variable is a policy response. Under the Clean Water Act, facilities that discharge waste into water, such as oil refineries, must satisfy applicable discharge requirements. The EPA can respond to violations with enforcement actions, including financial penalties. Those two bullets establish both the regulated outcome and the institution choosing the treatment.

The causal research question is whether enforcement actions improve the quality of wastewater discharges. The annual facility-level data contain two types of information listed at the bottom: measures of each facility’s wastewater-discharge quality and the enforcement actions the EPA took against that facility. Repeated annual observations can be valuable, but panel structure alone does not make enforcement exogenous. We still need to understand why enforcement changes across facilities and years.

Regulators observe discharge performance and tend to direct inspections, penalties, or other actions toward facilities with violations. Thus, the policy is often applied precisely where the outcome is poor or deteriorating. If we simply compare facilities with more and less enforcement, the high-enforcement group can start with worse water quality. The same logic applies within a facility if an adverse discharge shock triggers an enforcement action that year.

The institutional targeting rule is therefore part of the econometric model, not background detail. A raw association between enforcement and quality combines the effect of the policy with the EPA’s response to violations. Flip to the next tab to see the two directions written as outcome and selection equations.

Context

  • Under the Clean Water Act, facilities that discharge waste into water (e.g., oil refineries) must comply with applicable discharge requirements.

  • The Environmental Protection Agency (EPA) can take enforcement actions (e.g., financial penalties) against facilities that violate those requirements.

Research Question

Are enforcement actions effective in improving the quality of wastewater discharges?


Data

Annual data on

  • wastewater-discharge quality measures for individual facilities
  • enforcement actions taken against facilities by the EPA
Transcript

The top equation is the model of interest. Water quality equals beta zero plus beta one times enforcement actions plus u. Beta one is the policy effect we want, and u includes other observed or unobserved shocks to discharge quality that are not in this simple specification. If a larger water-quality measure means cleaner discharge, effective enforcement would imply a positive beta one. If the measure instead records pollution or violations, effective enforcement would imply a negative beta one, so always check how the outcome is coded.

The second equation represents the EPA’s selection decision. Enforcement actions equal alpha zero plus alpha one times water quality plus v. Alpha one describes how enforcement responds to measured quality, and v contains other influences on enforcement. When higher outcome values mean cleaner water, targeting poor quality would make alpha one negative. With a pollution measure, the sign would be positive.

The two bullets summarize the feedback: enforcement can affect water quality, and water quality can affect enforcement. More specifically, an adverse factor in u can worsen water quality and prompt an enforcement response. Enforcement is then correlated with u in the outcome equation and is endogenous. A negative raw association between enforcement and a clean-water measure need not mean enforcement causes worse quality. It may mean regulators correctly target facilities experiencing the worst discharge problems. Even a positive association does not automatically reveal beta one because policy impact and policy selection are mixed together.

That is the consequence printed at the bottom: enforcement actions are endogenous because they respond to the outcome itself. Estimating a causal effect requires variation in enforcement that is not driven by those water-quality shocks, or a design that credibly accounts for the targeting process. Examples might include an exogenous enforcement rule, a defensible timing strategy, or a panel design with assumptions strong enough to separate policy response from policy impact. The next major section turns from purposeful decisions to inaccurate measurement as another route to endogeneity.

Model of Interest

\text{water quality} = \beta_0 + \beta_1 \text{enforcement actions} + u


Selection (enforcement decision) model

\text{enforcement actions} = \alpha_0 + \alpha_1 \text{water quality} + v

  • water quality is affected by enforcement actions
  • enforcement actions are affected by water quality

Consequence

Enforcement actions are endogenous because they respond to water quality itself.

Measurement Error

Transcript

Measurement error means the value recorded in the dataset differs from the true quantity the regression is intended to use. “True” here means the underlying economic or physical variable in the model, not necessarily a value that researchers can ever observe perfectly. The gap may come from a respondent, an instrument, or an estimation procedure.

The examples show several mechanisms. In a household survey, respondents may forget, round, or deliberately misreport income and savings. Farmers in developing countries may have to recall rice yield without complete records, so the reported yield can differ from actual production. Some variables are estimated rather than observed at the relevant location. Spatially interpolated precipitation uses readings from surrounding stations to predict weather at a field, and imputed irrigation cost fills in a value using other information. Those constructed values can be useful, but they are not identical to the underlying field-level quantities.

The question at the bottom has no single answer for all measurement error. The consequence depends first on which side of the regression is measured inaccurately. Error in the dependent variable often adds noise without bias under a conditional-mean assumption. Error in an explanatory variable can make that observed regressor endogenous because the recording error becomes part of both the regressor and the composite disturbance. It also matters whether the error is mean zero, whether it is related to the true value, whether it is related to other regressors or the structural error, and whether one or several variables are mismeasured.

So do not memorize the blanket statement that measurement error always biases coefficients toward zero. Attenuation toward zero is a specific result for the classical errors-in-variables case we derive later. First, flip to the dependent-variable tab, where substituting the measured outcome shows why the condition and consequence are different.

Definition

The observed value of a variable differs from its true value.


Examples

  • reporting errors (any survey has the potential for misreporting)
    • household survey on income and savings
    • survey on rice yield by farmers in developing countries
  • the use of estimated values
    • spatially interpolated weather conditions (precipitation)
    • imputed irrigation costs

Question

What are the consequences of measurement error in variables used in a regression?

Transcript

Begin with the true model in the first line. Y star is the true dependent variable. Beta zero is the intercept; beta one through beta k are the coefficients on explanatory variables x one through x k; and u is the original structural disturbance. The slide states that the relevant OLS assumptions hold, including that the expected value of u, conditional on all the x variables, is zero.

We do not observe y star perfectly. The measurement-error definition says e equals observed y minus true y star. Rearranging that definition gives observed y equals y star plus e. A positive e means the recorded outcome is above its true value, and a negative e means it is below.

Now substitute the true model for y star. Observed y equals beta zero, plus beta one x one, continuing through beta k x k, plus u plus e. The estimable line groups the last two terms into a new composite error v, where v equals u plus e. Notice that none of the slope coefficients changed algebraically. Measurement error in y is added to the disturbance; it is not built into an explanatory variable.

For OLS slopes to remain unbiased, the new error must have conditional mean zero. The answer panel supplies the additional requirement: the expected value of e, conditional on x one through x k, must equal zero. Together with the original condition on u, this gives expected v conditional on all the x variables equal to expected u conditional on x plus expected e conditional on x, which is zero plus zero. Under suitable regularity conditions, the estimators are also consistent.

Intuitively, conditionally mean-zero outcome error spreads observations above and below the same regression function. When e adds noise that is uncorrelated with u, as in the usual classical case, it increases the disturbance variance, leading to less precise coefficient estimates and larger standard errors, but it does not systematically tilt the fitted relationship. This conclusion is conditional, not automatic. If high-x observations systematically overreport or underreport y, then expected e given x is not zero. For example, if high-income households underreport savings more severely, the recording error can be related to income and bias the slope. We now move the measurement error from y to x and see why even classical error creates a correlation by construction.

True Model

y^*= \beta_0 + \beta_1 x_1 + \dots + \beta_k x_k + u

with the relevant OLS assumptions satisfied, including E[u\mid x_1,\dots,x_k]=0.


Measurement error

Define the difference between the observed value y and the true value y^* as

e = y-y^*


Estimable Model

Using y=y^*+e, the estimable model is

y = \beta_0 + \beta_1 x_1 + \dots + \beta_k x_k + v, \qquad v=u+e


Question

What are the conditions under which OLS estimators are unbiased?


Answer

E[e\mid x_1, \dots, x_k] = 0

Together with E[u\mid x_1,\dots,x_k]=0, this implies E[v\mid x_1,\dots,x_k]=0. Measurement error in the dependent variable then increases noise but does not bias the OLS coefficient estimators.
Transcript

Now the measurement error is in an explanatory variable. The true model says y equals beta zero plus beta one times x one star plus u. X one star is the true regressor the economic relationship calls for, beta one is its true slope effect, and u is the structural disturbance. The statement that the relevant OLS assumptions hold includes zero conditional mean for u given x one star.

The researcher observes x one rather than x one star. The definition e one equals x one minus x one star says that measurement error is the recorded value minus the true value. Rearrange it carefully: x one star equals x one minus e one. Substitute that expression into the true equation. You get y equals beta zero plus beta one times the quantity x one minus e one, plus u. Distributing beta one gives beta zero plus beta one x one plus u minus beta one e one.

The estimable model groups the last two terms into v. Thus y equals beta zero plus beta one x one plus v, where v equals u minus beta one times e one. This algebra shows the crucial contrast with dependent-variable error. The observed regressor x one contains e one because x one equals x one star plus e one. At the same time, the new regression error v contains negative beta one times that same e one. The regressor and error can therefore share a component.

OLS applied to the observed x requires the expected value of v conditional on x one to be zero. Even if u is exogenous with respect to the true regressor, that does not guarantee the new condition. We need to ask whether observed x one is related to its own measurement error e one. The next tab imposes the favorable classical assumptions and calculates that covariance. The surprising answer is yes, because the observed variable literally includes the error.

True Model

Consider the following general model

y = \beta_0 + \beta_1 x_1^* + u

with the relevant OLS assumptions satisfied.


Measurement error

Define the difference between the observed value x_1 and the true value x_1^* as

e_1 = x_1-x_1^*


Estimable Model

Using x_1^*=x_1-e_1, the estimable model is

y = \beta_0 + \beta_1 x_1 + v, \qquad v=u-\beta_1 e_1


Question

OLS requires E[v\mid x_1]=0. Even when u is exogenous, this condition fails if the observed regressor x_1 is correlated with its measurement error e_1. Does classical measurement error create such a correlation?

Transcript

The definition callout lists the classical errors-in-variables assumptions. First, the expected value of e one is zero, so the measurement error has no systematic positive or negative mean. Second, the covariance between true x one star and e one is zero, so the size and direction of the recording error are unrelated to the true regressor. Third, the covariance between u and e one is zero, so measurement error is unrelated to the structural disturbance. These are strong, favorable assumptions, but they do not make the observed regressor exogenous.

Work through the three displayed covariance lines. The first replaces observed x one with its definition, x one star plus e one. We are therefore taking the covariance of x one star plus e one with e one. The second line uses covariance’s additivity: this equals the covariance of x one star with e one, plus the covariance of e one with itself. A variable’s covariance with itself is its variance, so the second term is variance of e one. The classical assumption makes the first covariance zero. The result is sigma squared sub e one, the variance of the measurement error, which is strictly positive whenever there is nondegenerate measurement error.

This is not a contradiction. E one is uncorrelated with the true x one star, but observed x one is true x one star plus e one. A high positive measurement error mechanically raises the recorded x, and a negative error lowers it. Observed x and e must therefore move together.

Finally, recall from the setup tab that v equals u minus beta one e one. Under the displayed classical assumptions and the original exogeneity assumptions, the covariance between observed x one and v includes negative beta one times the variance of e one and is generally not zero. Thus observed x one is endogenous in the estimable equation. OLS is inconsistent even though the measurement process is mean zero and unrelated to the true regressor. The next tab derives the special direction of that inconsistency in this simple classical case.

Definition

Under classical errors-in-variables assumptions, the measurement error has mean zero and is uncorrelated with both the true regressor and the structural error:

E[e_1]=0,\qquad \operatorname{Cov}(x_1^*,e_1)=0,\qquad \operatorname{Cov}(u,e_1)=0.

Even though e_1 is uncorrelated with x_1^*, it is correlated with the observed regressor x_1=x_1^*+e_1:

\begin{align*} \operatorname{Cov}(x_1,e_1) &=\operatorname{Cov}(x_1^*+e_1,e_1)\\ &=\operatorname{Cov}(x_1^*,e_1)+\operatorname{Var}(e_1)\\ &=\sigma_{e_1}^2>0. \end{align*}

Thus, under the classical assumptions, the mismeasured regressor is correlated with the composite error v=u-\beta_1e_1.

Transcript

The displayed probability-limit formula gives the direction of the inconsistency for this simple-regression, classical-measurement-error case. “The probability limit of beta one hat” means the value toward which the OLS slope estimate converges as the sample grows. It equals the true beta one times a fraction. The numerator is the variance of true x one star, which is the signal variation. The denominator is that same true-signal variance plus the variance of e one, the measurement-error noise. Under the classical assumption that x one star and e one are uncorrelated, that denominator is also the variance of observed x one.

The slide names this fraction lambda. Lambda is often called the reliability ratio: signal variance divided by total observed variance. When both the true regressor varies and measurement error has positive variance, lambda lies strictly between zero and one. Therefore beta one hat converges to lambda times beta one rather than to beta one itself. OLS is inconsistent, and multiplying by a positive number below one pulls the coefficient toward zero.

Use the two sign bullets to interpret the asymptotic bias, which is the probability limit of beta one hat minus beta one. If beta one is positive, lambda beta one is smaller than beta one, so the bias is negative. If beta one is negative, lambda beta one is less negative than beta one, so the bias is positive. In both cases the estimated magnitude is attenuated. “Toward zero” does not mean the estimate necessarily equals zero, and it does not mean a finite-sample estimate can never cross zero. It describes the probability limit relative to the true coefficient.

The callout’s second bullet follows directly from the ratio. Holding signal variance fixed, increasing measurement-error variance enlarges the denominator, reduces lambda, and produces stronger attenuation. With no measurement-error variance, lambda equals one and this source of inconsistency disappears. If noise dominates signal, lambda approaches zero and the probability limit is severely attenuated.

Keep the scope of this result attached to the formula. It relies on classical error, one mismeasured explanatory variable in this simple regression, and the stated exogeneity conditions. With nonclassical error, multiple regressors, or several mismeasured variables, other coefficients can also be affected and the bias need not point toward zero. What carries forward generally is the endogeneity diagnosis from the previous tab; attenuation is the special directional result under these assumptions.

Question

So, what is the direction of the bias?


Direction of bias

Under classical measurement error,

\operatorname{plim}(\widehat{\beta}_1) =\beta_1\frac{\operatorname{Var}(x_1^*)} {\operatorname{Var}(x_1^*)+\operatorname{Var}(e_1)} =\lambda\beta_1, \qquad 0<\lambda<1.

  • If \beta_1>0, the asymptotic bias is negative.
  • If \beta_1<0, the asymptotic bias is positive.


Attenuation bias

  • The coefficient estimator is inconsistent and converges toward zero relative to the true coefficient.

  • More measurement-error variance makes the attenuation factor \lambda smaller.