J <- 50
n <- 100
area_mean_x <- rnorm(J, 0, 1)
people <- data.frame(area = rep(1:J, each = n)) |>
mutate(X = area_mean_x[area] + rnorm(J * n, 0, 2),
Y = X + rnorm(J * n, 0, sqrt(20)))Mindless statistics and sophisticated bigotry
Class
This is the student version of the class sheet. Parts marked (class) are for the class; the rest are practice. Answers are revealed in class and are not released afterwards.
To start: two studies of infants
In groups of three or four, with people who skimmed the same study as you if you can find them.
- Which study did you skim? Would its result be reproduced? Say what in the paper made you think so.
- Find one comparison the abstract reports, and one comparison the same tables allow that the paper does not report.
Problems
Before you start, from memory: when you average people into areas and regress the area averages, what happens to the slope, and what happens to \(R^2\)?
Everything can be done by hand with \(\Phi(2) = 0.977\), \(\Phi(1.82) = 0.966\) and \(e^{-0.46} = 0.63\).
Problems 1 to 4 are for the class. The problems headed “Optional, can be done at home” are for practice afterward. Answers are revealed in class. Attempt each part first.
1. A \(p\)-value of 0.01
A report of a trial says: “\(p = 0.01\), so there is a 1% chance that the treatment does nothing, and a repeat of the trial would be significant 99 times in 100.”
Say what the \(p\)-value is the probability of. Which two of the six statements from the lecture has the report made?
Of the effects a field tests for, 10% are real. Tests are run at level \(0.01\) with power \(0.5\). Compute the share of rejections that are false.
What does the chance that a repeat of the trial is significant depend on? Suppose the repeat is run at level \(0.01\), with power \(0.5\) against the effect the trial was designed to detect. Give the chance if the treatment has that effect, and if it does nothing.
2. Aggregation and \(R^2\)
People \(i\) live in areas \(j\), \(n\) in each. \(X_{ij} = B_j + W_{ij}\), where \(B_j\) is shared by everyone in area \(j\) and has variance \(\sigma_B^2\) across areas, and \(W_{ij}\) has mean zero and variance \(\sigma_W^2\) and is uncorrelated with \(B_j\) and across people. \(Y_{ij} = \alpha + \beta X_{ij} + \varepsilon_{ij}\), where the errors have variance \(\sigma^2\), are uncorrelated across people, and have mean zero given the \(X\) values. Take \(\beta = 1\), \(\sigma_B^2 = 1\), \(\sigma_W^2 = 4\) and \(\sigma^2 = 20\).
Compute the population \(R^2\) of the regression of \(Y\) on \(X\) across people, and of \(\bar Y_j\) on \(\bar X_j\) across areas, for \(n = 25\) and for \(n = 400\).
Is the difference between the two \(R^2\) values a matter of bias or of variance? Say what happens to the slope.
3. Group membership and \(R^2\)
Two groups of equal size. Within each group an outcome has standard deviation 1, and the two group means differ by \(d\).
Show that the variance of the outcome across everyone is \(1 + d^2/4\), and that group membership accounts for a share \(R^2 = \dfrac{d^2}{d^2 + 4}\) of it. Compute the share for \(d = 0.2\) and for \(d = 1\).
A chart shows the two group means and nothing else. What is the \(R^2\) of the outcome on group membership when the data are those two numbers? What has the chart left out?
4. Ecological regression and constancy
In each area \(i\) (areas are indexed by \(i\) from here on), a fraction \(x_i\) of the people belong to a group. The group’s rate of some outcome is \(p_i\) and everyone else’s rate is \(q_i\), so the area’s rate is \(y_i = p_i x_i + q_i(1 - x_i)\). Ecological regression fits \(y_i = a + b x_i\) and reads \(\hat a\) as everyone else’s rate and \(\hat a + \hat b\) as the group’s rate.
Two areas of equal size, \(x = 0.1\) and \(x = 0.3\). The group’s rate is \(p = 0.2\) in both; everyone else’s rate is \(q = 0.30\) in the first and \(q = 0.40\) in the second. Compute the two area rates, fit the line through the two points, and read off what ecological regression reports for the group and for everyone else. Compare with the truth over the two areas together.
Optional, can be done at home
Do not open a hint unless you have been stuck for at least a few minutes, and open them in order.
5. Aggregation and \(R^2\) with an area effect
The model of problem 2.
Now let each area have an effect of its own, \(\varepsilon_{ij} = u_j + e_{ij}\) with \(\operatorname{Var}(u_j) = 1\) and \(\operatorname{Var}(e_{ij}) = 20\). \(u_j\) and \(e_{ij}\) are uncorrelated with each other and with the \(X\) values, and \(e_{ij}\) is uncorrelated across people. Compute the \(R^2\) across people, the \(R^2\) across areas at \(n = 100\), and its limit as \(n\) grows.
A report shows a regression across 50 regions with \(R^2 = 0.9\). What does that tell you about the \(R^2\) of the same relationship across people? What would you need to know?
The area effect \(u_j\) is the same for everyone in the area. What does averaging do to it?
6. Ecological regression in general
The setting of problem 4.
Now in general: \(p_i = p\) and \(q_i = q_0 + \gamma x_i\). Substitute into the accounting identity and show that when the shares are small the fitted slope is approximately \(p - q_0 + \gamma\). What does the regression then report as the group’s rate, and which way is it wrong when \(\gamma < 0\)? In problem 4, \(\gamma = 0.5\). How good is the approximation there?
Freedman’s neighborhood model puts \(\hat p_i = \hat q_i = y_i\). Apply it to the two areas in problem 4. Which model is closer to the truth here, and what would you need to know to say which is closer in a case where the truth is unknown?
Substitute \(q_i = q_0 + \gamma x_i\) into \(y_i = p x_i + q_i(1 - x_i)\) and collect the powers of \(x_i\).
7. Two slopes
Split the covariance and the variance into between-area and within-area parts: \(\operatorname{Cov}(X, Y) = \operatorname{Cov}_B + \operatorname{Cov}_W\) and \(\operatorname{Var}(X) = \operatorname{Var}_B + \operatorname{Var}_W\). Suppose \(\operatorname{Cov}_B = -1\), \(\operatorname{Var}_B = 2\), \(\operatorname{Cov}_W = 4\) and \(\operatorname{Var}_W = 8\).
Compute the slope across area averages, \(b_B\), the pooled slope within areas, \(b_W\), and the slope across people, \(b_{\text{all}}\). Check that \(b_{\text{all}} = \lambda b_B + (1 - \lambda) b_W\) with \(\lambda = \operatorname{Var}_B/(\operatorname{Var}_B + \operatorname{Var}_W)\).
Keep the two variances and the within-area covariance. How negative does the slope across area averages have to be for the slope across people to be negative?
In Gelman, Shor, Bafumi and Park (2007), within states richer voters were more likely to vote Republican, while richer states voted Democratic. From those two facts alone, can you say which way income and vote were related across all voters? The authors add that income varies far more within states than average income does between states. What does that settle, and what does it leave open?
A slope is a covariance divided by a variance. Compute it between areas, within areas, and for the two together.
8. The method of bounds
This is Freedman’s section 4, optional in the reading list; the notes state the bounds.
In a district, 8% of residents belong to a group and 1% of all residents have some outcome. Write the accounting identity in the group’s rate \(p\) and everyone else’s rate \(q\), and find the range of values each can take.
In Freedman’s example (Washington, 1995) the group was 7.9% of adults and 34.4% of adults had the outcome. Find the bounds on the group’s rate there, compare them with yours in (a), and explain what you find. State a rule for when the bounds on a group’s rate are the whole of \([0, 1]\).
A precinct has \(x = 0.5\) and \(y = 0.6\). Bound both rates.
Say, in one sentence each, what the bounds settle and what they leave open.
Set the other rate to 0 and then to 1, and check whether the rate you want stays between 0 and 1.
9. The largest of \(n\)
\(M_n\) is the largest of \(n\) independent draws from a distribution with distribution function \(F\), so \(P(M_n \le x) = F(x)^n\). For an exponential with mean 1, \(E[M_n] = 1 + \tfrac12 + \cdots + \tfrac1n\); the notes derive it.
Compute \(E[M_n]\) for four draws from Uniform\((0, 1)\), and for four draws from an exponential with mean 1. Then compute the probability that the largest of 20 standard normals exceeds 2.
Derive \(E[M_n] = n/(n + 1)\) for the uniform.
Two groups post items, and each item’s extremity is an independent draw from the same continuous distribution. One group posts 200 items a day and the other 1,800. A feed shows the most extreme item from each. (i) With what probability is the larger group’s item the more extreme of the two? (ii) In the notes the two groups posted 100 and 10,000 items. The expected largest of 100 standard normals is \(2.51\) and of 10,000 is \(3.85\). Say what those two numbers do and do not say about the two groups. (iii) Now both groups post 1,000 items a day, and extremity is normally distributed with the same mean in both groups and a standard deviation 10% larger in the second. The expected largest of 1,000 standard normals is \(3.24\). What are the two expected shown items, in the first group’s standard deviations above the mean?
(Harder.) A journal publishes an estimate only when its \(z\)-value exceeds \(1.96\); a feed shows the largest of 100 \(z\)-values. All the \(z\)-values are true nulls. Compare the expected published value with the expected shown value, using \(E[Z \mid Z > 1.96] = 2.34\) from the significance filter, and say in one sentence what the two selections have in common.
The largest of the draws is below a value when all of them are.
10. Spreads and tails (harder)
Two groups have normally distributed scores with the same mean. Group 2’s standard deviation is 10% larger than group 1’s. A prize goes to everyone more than two of group 1’s standard deviations above the mean.
Using \(\Phi(2) = 0.977\) and \(\Phi(1.82) = 0.966\), compute the fraction of each group who win a prize, and the ratio. If the groups are the same size, what share of the prize-winners are from group 2?
Hyde and Mertz (2009) report, from US state mathematics tests, a mean difference \(d\) between boys and girls of \(-0.02\) to \(0.06\) and a variance ratio of \(1.11\) to \(1.21\) (their Table 1); from one state’s test, a boys-to-girls ratio above the 99th percentile of \(2.06\) among white students and \(0.91\) among Asian-American students; and from the 2003 PISA test, variance ratios from \(0.95\) to \(1.24\) across countries (their Table 2). Which of these are claims about averages, which about spreads, and which about tails? Using part (a) and Table 2, say in two sentences how far a tail ratio can move when the variance ratio moves.
A commentator says: group 2 is 60% of the prize-winners, so group 2 should be expected to hold 60% of the jobs that need a score that high. List what has to be assumed (i) to compute the 60% and (ii) to read it as a share of jobs. Then say what further step would be needed to conclude anything about one applicant.
Express the cutoff in each group’s own standard deviations.
11. Ecological inference from a table of boroughs
A newspaper prints, for each of the 32 London boroughs, the share of adults with a degree and one party’s share of the votes cast, notes that the correlation across boroughs is 0.6, and headlines it as a fact about how graduates vote.
Say what the table identifies and what it does not.
The identity \(y = p x + q(1 - x)\) needs \(x\) and \(y\) to be shares of the same people. Say why they are not here, and what would have to be assumed about who votes for the identity to be used.
Granting that, state the assumption under which a regression across boroughs gives graduates’ support for the party, and say what happens to the estimate if non-graduates’ support rises with the graduate share of the borough.
Say what the Duncan–Davis bounds would add, and what you would need to know about a borough to say whether its bounds on graduates’ support are informative.
Ask what the units of the table are, and whether \(x\) and \(y\) are shares of the same people.
Discussion break
In groups of three or four, then to the room:
How do you become a good critic of a statistical claim about a group? Have you seen bad statistical criticism, and what made it bad? Try to say which of your points the problems above give you a tool for.
Coding
Open a fresh R script. You will need ggplot2 and dplyr. In class: A.1 to A.3, then B.1 and B.2.
A. Aggregation and \(R^2\)
Fifty areas, one hundred people in each. Areas differ in \(X\); people differ more.
- Regress
YonXacross people. Record the slope and \(R^2\). - Average
XandYwithin each area, regress the area averages, and record the slope and \(R^2\). Compare with what the formula from problem 2 gives at \(n = 100\). - Plot the people as faint points and the area averages as large points, with both regression lines. Describe in one sentence what a reader shown only the area averages would conclude.
Optional, can be done at home
- Add an area effect:
Y + rep(rnorm(J, 0, 1), each = n). Repeat 1 and 2. Compare with problem 5(a). - Make the within-area covariance positive and the between-area covariance negative: generate
Yas-0.5 * area_mean_x[area] + 0.5 * (X - area_mean_x[area]) + noise. Check that the two slopes have opposite signs.
B. The largest of \(n\)
- Simulate the largest of \(n\) standard normals, 2000 times, for \(n = 1, 10, 100, 1000\). Compute the mean for each \(n\) and compare with \(2.51\) at \(n = 100\) and with \(\sqrt{2 \ln n}\).
- Two groups, 100 posts and 10,000 posts a day, same distribution. Find the most extreme post from each. Repeat 200 times. In what fraction of the repeats is the larger group’s post the more extreme? Compare with the answer in the notes.
Optional, can be done at home
- Plot a histogram of the 2000 maxima at \(n = 100\). Where is it relative to zero?
- Check \(E[M_n] = n/(n+1)\) for the uniform at \(n = 4\) and the harmonic sum for the exponential at \(n = 4\) by simulation.
- A group has 1000 posts a day. Estimate by simulation the fraction of days on which its most extreme post exceeds 3, and compare with \(1 - 0.99865^{1000}\).
C. Ecological regression on simulated precincts
Optional, can be done at home
Fifty precincts of equal size. The group’s share \(x_i\) runs from 0.05 to 0.5. The group’s rate is \(p = 0.3\) everywhere; everyone else’s rate is \(q_i = 0.4 + 0.4\,x_i\).
- Compute the precinct rates \(y_i\) and fit
lm(y ~ x). Read off the two rates the regression reports and compare with the truth over all precincts. - Compute the Duncan–Davis bounds for \(p_i\) and \(q_i\) in each precinct and plot them against \(x_i\). For which group and which precincts are they informative?
- Set \(q_i = 0.4\) (constancy holds) and repeat 1.
Discussion
More prompts than we will get through. Groups of three or four, then to the room; every group’s answer goes on the shared document.
- Stark and Saltelli end with what statisticians can do. Pick one cause of cargo-cult statistics from the lecture (courses, software, editorial policies, careers). What would you change, and who would have to agree?
- Is an aggregate claim about a group ever a claim about a member of it? You gave one case each way in the lecture. Now take three: an insurer pricing one driver from the claims of drivers of the same age; a doctor quoting a five-year survival rate to one patient; an employer screening applicants by the average of their school. For each: is it statistical discrimination, and on whom do its errors fall?
- Three claims from this week: the last sentence of the abstract of the study you skimmed; “the foreign-born have higher incomes,” read off Freedman’s line; and Summers’s sentence about the available pool. For each: is it about averages, spreads or tails? About people or about places? Who was in the data, and how did they get there? (No verdict on whether the study you skimmed holds up; that is not the question.)
- Each group takes one setting and writes three sentences: what is selected; which way the numbers you see are biased and what makes the bias larger or smaller; who benefits from the bias being invisible.
- A league table of the boroughs whose exam results improved most this year.
- A social media feed showing the ten posts from one political side that got the most reactions today.
- A national newspaper’s map of the regions with the highest rate of some crime, colored by an aggregate correlation with one demographic variable.
- A dating app’s “top picks” for each user.
- A ranking of the best-performing fund managers over the last year.
- Freedman’s Table 2 (in the notes): for the same aggregate data, ecological regression puts a group’s high-income rate at 85% and the neighborhood model at 36%; the truth was 28%. You are advising someone in a case like this one, with only the aggregate data, who has to publish one number for the group’s rate. Which model would you put forward, what would you say it assumes, and to whom do you owe the explanation if it is wrong?
To finish
One sentence, individually: a regression across regions has \(R^2 = 0.9\). What have you learned about people?
For next time
Bring one argument from the news that justifies a rule or a policy by its consequences. Write the argument in one sentence, and list the cause-and-effect claims it needs to be true. Every group’s examples will go on the screen.