Transcript
We have established that finding an association among variables is usually easy, while isolating a causal effect, especially in economics, is hard. Now we give the central reason a name: endogeneity. That word will appear in nearly every lecture from here on, and much of econometrics is organized around either preventing it, diagnosing its consequences, or finding credible variation that gets around it.
At an intuitive level, endogeneity means that the explanatory variable whose effect you want is entangled with other determinants of the outcome. When that happens, a comparison between high and low values of the explanatory variable also compares different values of those other determinants. The regression cannot tell which part of the outcome difference belongs to which cause. In the glasses story, wearing glasses was entangled with time spent studying. In the next example, the number of firefighters is entangled with the scale of a fire.
I am not going to make you memorize the fully formal definition on this tab. Lecture nine defines it precisely as correlation between an explanatory variable and the econometric error term, and explains what that does to an estimator. For now, build the intuition through the example. Students who meet the definition before the intuition can often recite it perfectly and still fail to spot the problem in their own work. Flip to the Example tab and look for the variable that determines both the response and the decision.
It is super easy to find an association of multiple variables, but it is incredibly hard to find a causal effect (at least in Economics)!!
That is due to the problem called endogeneity , which is going to be defined formally later.
Transcript
Suppose you want to know the causal impact of firefighter deployment on the number of deaths in fire events. Be precise about the variables. The outcome is the death toll. The variable of interest is the number of firefighters deployed. The causal question asks how the death toll would change if we sent more firefighters to a fire while holding the fire’s relevant circumstances fixed.
Now read the five rows in the table. Event one has five deaths and twenty firefighters. Event two has zero deaths and three firefighters. Event three has ten deaths and ten firefighters. Event four has three deaths and five firefighters. Event five has fifty deaths and fifty firefighters. Across these events, the overall association is strongly positive: the fires receiving the largest deployments tend to have the largest death tolls. Event two is at the low end of both columns, and event five is at the high end of both. Stating that association is easy.
Answer the second question at the bottom honestly. Can we infer that deploying additional firefighters causes deaths, or recommend sending fewer firefighters? Obviously not. Fire departments do not choose crew size by rolling a die. They respond to information about the emergency. A severe fire can both kill more people and trigger a larger deployment. That creates a positive association even if the causal effect of an additional firefighter is to save lives.
This is not a logical technicality. If you confuse the raw association with a causal effect, you produce the dangerous recommendation to withhold firefighters from the worst fires. Hold on to the discomfort you feel about that conclusion. Your economic knowledge tells you that the observations being compared differ in an important way. The next tab adds that missing variable and shows exactly what happened.
You are interested in the causal impact of fire fighters on the number of death tolls in fire events
fire event | death toll | # of firefighters deployed |
|---|
1 | 5 | 20 |
2 | 0 | 3 |
3 | 10 | 10 |
4 | 3 | 5 |
5 | 50 | 50 |
Questions
- How are they associated ?
- Can you say anything about the causal effect of fire fighters deployment on the number of death tolls?
Transcript
Here is the same five-event table with one additional column: the scale of each fire. Read across the rows. Event one has scale twenty, event two scale five, event three scale twenty, event four scale ten, and event five scale one hundred. Once that column is visible, the positive raw association has a clear explanation. Larger fires tend to receive more firefighters, and larger fires also tend to cause more deaths. Scale was driving both variables. It is the C from our fourth arrow diagram, and in the first table it was invisible to the researcher.
Now make the comparison highlighted beneath the table. Events one and three both have scale twenty, so on this recorded dimension they are comparable fires. Event one received twenty firefighters and had five deaths. Event three received ten firefighters and had ten deaths. Within that same-scale pair, the fire with more firefighters has fewer deaths. The direction of the comparison reverses after we hold scale fixed. This is a small illustration, not enough observations by itself to establish a general causal effect, but it shows why the uncontrolled association can have the wrong sign.
The phrase “you ignored an important variable” is doing real work. If scale were observed for every event and measured well, we could include it in a regression or compare events within sufficiently similar scale categories. If scale is missing or inherently difficult to measure, then it lives in the error term, and ordinary comparisons cannot directly hold it fixed. Every causal method in this course is trying, in a different way, to create the comparison we actually want: fire events with the same underlying severity but different firefighter deployments. The next tab turns that intuition into a definition.
You ignored an important variable!!
fire event | death toll | # of firefighters deployed | scale of fire |
|---|
1 | 5 | 20 | 20 |
2 | 0 | 3 | 5 |
3 | 10 | 10 | 20 |
4 | 3 | 5 | 10 |
5 | 50 | 50 | 100 |
Now compare fire events 1 and 3, which are of the same scale : the one with more firefighters (20 vs 10) had fewer deaths (5 vs 10).
Transcript
Now read the definition in the important callout. A variable of interest is endogenous when it is correlated with the error term. The error term is everything that affects the outcome we want to explain but is not included in the model. That is the general statement, and it is deliberately general: several different stories can produce that correlation, and the next slide lists them. Whatever the story, the problem is the same correlation.
The bold line beneath the definition gives the most common of those stories, and it is the one we just walked through. An unobservable, meaning a variable we cannot observe or that is missing from the data, is correlated with the variable of interest and has a nonzero effect on the outcome. Because it is missing from the model, it sits inside the error term, so a correlation with it is a correlation with the error term. In the firefighter example, deployment is the variable of interest, fire scale is missing from the first table, deployment rises with scale, and scale affects deaths. A comparison across deployment levels therefore also changes scale. Such an unobserved variable is called a confounder or confounding factor because it confounds, or mixes together, the effect we want with another effect.
Read the note carefully, because this is a distinction students routinely lose. Not every unobservable is a confounder. There are thousands of details we never measure. To threaten the effect of a particular variable of interest, an omitted factor must satisfy both conditions shown here: it must be associated with that variable of interest, and it must affect the outcome. An omitted variable that affects deaths but is unrelated to firefighter deployment can add unexplained variation without systematically mixing its effect into the deployment comparison. A variable related to deployment but with no effect on deaths does not distort the death comparison through this channel either.
The two conditions are specific to the research question. A factor that confounds one explanatory variable may not confound another, and a variable irrelevant for one outcome may matter for a different outcome. That is why identifying likely confounders requires economic reasoning about both the decision process and the outcome process, not just scanning a correlation matrix. The next tab maps every piece of this definition into an equation.
A variable of interest is endogenous when it is correlated with the error term , which collects everything that affects the variable you want to explain but is not in the model
The most common source: an unobservable (a variable that cannot be observed or is missing) that is correlated with the variable of interest and has a non-zero impact on the variable you want to explain. Because it is missing from the model, it sits in the error term, so the variable of interest ends up correlated with the error term. Such unobserved variables are also called confounder/confounding factor .
Note that not every unobservable is a confounder. To create endogeneity, it has to be correlated with the variable of interest and affect the variable you want to explain. We will come back to this point in later lectures.
Transcript
Let’s map the definition back onto the fire example term by term, because I want you fluent in translating a story into a model. The red labels identify the three roles. The variable of interest is the number of firefighters deployed. The unobservable or confounder is the scale of the fire, along with any other missing determinants that influence both deployment and deaths. The variable to explain, which we also call the outcome or dependent variable, is the death toll.
Now look at the first equation. Read it aloud as: death toll equals alpha, plus beta times the number of firefighters, plus u. Alpha is the intercept. Beta is the parameter we would like to interpret as the change in death toll caused by one additional firefighter, if the assumptions needed for that causal interpretation hold. U is the error term, the collection of determinants of deaths that the displayed regressor does not include.
The second line opens that container. It says u equals gamma times fire scale plus v. Gamma measures how scale affects the death toll. V collects the other unobserved determinants left after singling out scale. This decomposition is conceptual. We may not have scale in the data, but writing it inside u lets us see the source of the problem.
The endogeneity statement at the bottom now follows step by step. Firefighter deployment is correlated with scale because larger fires receive larger crews. Scale is a component of u because it affects deaths and was omitted from the first equation. Therefore the number of firefighters is correlated with u. When deployment rises, the omitted severity component tends to rise too, so a regression that uses only deployment cannot separate beta from gamma’s contribution. The error term is not harmless noise here. It contains a systematic reason that high-deployment fires have high death tolls. That is endogeneity written down.
In the firefighter example,
- variable of interest : the number of firefighters
- unobservables/confounder : the scale of fire events (and other factors)
- variable to explain : death toll
The model
\begin{align*}
\text{death toll} & = \alpha + \beta\; \text{\# of fire fighters} + u\\
u & = (\gamma\; \text{scale} + v) \text{ is the error term (collection of unobservables)}
\end{align*}
Endogeneity Problem
# of fire fighters is correlated with scale, which we ignored
Transcript
Now return to the wage equation from the start of the lecture and run the same argument yourself. Read the equation as: wage equals beta zero, plus beta one times education, plus beta two times experience, plus beta three times training, plus u. Wage is the outcome. Education, experience, and training are included explanatory variables. Beta one is the education slope we would like to interpret, and u contains every omitted determinant of wage.
The question on the screen focuses on education: what is inside u that is likely to be correlated with years of education? The classic answer written below is innate ability. Follow both arrows in the bullets. Innate ability can affect wage directly because, for a given level of schooling, a more able worker may be more productive and receive a higher wage. Innate ability can also affect education because people with greater academic ability may find school less costly or receive stronger encouragement to continue. Ability is then associated with education and has a nonzero effect on wage, so it satisfies both conditions for a confounder.
If ability is omitted, workers with more schooling differ not only in schooling but also, on average, in ability. A regression may credit education for wage differences partly produced by ability. Under the usual positive relationships described on the slide, that produces an upward omitted-variable bias in the education coefficient. The exact sign depends on the direction of both relationships, so it is the economic story, not the word “omitted,” that determines the direction.
Experience and training can have their own confounders as well. The model is causal only if the relevant included variables are uncorrelated with the remaining error term after we condition on the other regressors. Lecture eleven uses instrumental variables to address this kind of education example. Before we get there, the next slide organizes the main ways endogeneity enters and the two stages at which we can respond.
wage = \beta_0 + \beta_1 educ + \beta_2 exper + \beta_3 training + u
What are unobservables in u that are likely to be correlated with educ?
An important unobservable
- innate ability \Rightarrow wage
- innate ability \Rightarrow education