---
title: "Mindless statistics and sophisticated bigotry"
subtitle: "Class"
author: "Joshua Loftus"
format:
  html:
    toc: true
execute:
  message: false
  warning: false
---

::: {.callout-note}
This is the student version of the class sheet. Parts marked **(class)** are for the class; the rest are practice. Answers are revealed in class and are not released afterwards.
:::

```{r setup, include=FALSE}
library(ggplot2)
library(dplyr)
theme_set(theme_minimal())
set.seed(313)
```

## In class

1. A discussion of the study of infants you skimmed.
2. Problem 1, by hand.
3. Problem 2, by hand, then a discussion.
4. One simulation in R, then a discussion.
5. Two more discussions.

Discussions are in groups of three or four, then with the room. Each
problem, and the simulation, is gone through as soon as the room has
tried it. Everything under "Optional, can be done at home" is practice
for afterward.

## Discussion 1: two studies of infants

Sit with people who skimmed the same study as you, if you can find
them.

How convincing is the study's main claim?

## Problems

Before you start, from memory: in the lecture's simulation model, when
you average people into areas and regress the area averages, what
happens to the slope, and what happens to $R^2$?

### 1. False rejections

From last week. Of the effects a field tests for, 10% are real. Tests
are run at level $0.01$ with power $0.5$. Compute the share of
rejections that are false.

### 2. Group membership and $R^2$

Two groups of equal size. Within each group an outcome has standard
deviation 1, and the two group means differ by $d$.

(a) Show that the variance of the outcome across everyone is
$1 + d^2/4$, and that group membership accounts for a share
$R^2 = \dfrac{d^2}{d^2 + 4}$ of it. Compute the share for $d = 0.2$ and
for $d = 1$.

(b) A chart shows the two group means and nothing else. What is the
$R^2$ of the outcome on group membership when the data are those two
numbers? What has the chart left out?

## Discussion 2: decisions and groups

When is it reasonable to use information about someone's group to make
a decision about them?

## Simulation: the most extreme post

Two groups post on the same day. The extremity of each post is a draw
from the same distribution, whichever group it comes from. One group
posts 100 times and the other 10,000 times. A feed shows the most
extreme post from each group.

```{r feed_starter, eval=FALSE}
# one day
small <- rnorm(100)        # the smaller group's posts
large <- rnorm(10000)      # the larger group's posts, same distribution
max(small)                 # the post the feed shows from the smaller group
max(large)                 # the post the feed shows from the larger group

# 200 days: on each day, is the larger group's shown post the more extreme?
days <- replicate(200, max(rnorm(10000)) > max(rnorm(100)))
mean(days)                 # the fraction of days on which it is
```

1. Run the first block a few times. A typical post is at 0. How far
   from it are the two posts the feed shows?
2. Before you run the second block, write down the fraction you
   expect. Then run it.
3. Give the larger group 3,000 posts and the smaller group 1,000.
   Write down the fraction you expect, then run it.

## Discussion 3: reading a feed

A feed shows the most extreme post from each group. How should readers
interpret what they see?

## Discussion 4: a disagreement

A junior analyst disagrees with the conclusion their manager wants to
publish. What should they do?

## Discussion 5: AI and cargo-cult statistics

Could AI make cargo-cult statistics worse, or help solve the problem?

## If there is time

- How do you become a good critic of a statistical claim about a
  group? Have you seen bad statistical criticism, and what made it
  bad?
- Pick one of these. What was selected before you saw it, and does
  that matter?
  - A league table of the boroughs whose exam results improved most
    this year.
  - A social media feed showing the ten posts from one political side
    that got the most reactions today.
  - A national newspaper's map of the regions with the highest rate of
    some crime, colored by an aggregate correlation with one
    demographic variable.
  - A dating app's "top picks" for each user.
  - A ranking of the best-performing fund managers over the last year.
- For the same aggregate data, ecological regression puts a group's
  high-income rate at 85% and the neighborhood model puts it at 36%.
  You have only the aggregate data and are asked for the group's
  rate. What do you report?

## To finish

One sentence, individually: a regression across regions has
$R^2 = 0.9$. What have you learned about people?

## For next time

Bring one argument from the news that justifies a rule or a policy by
its consequences. Write the argument in one sentence, and list the
cause-and-effect claims it needs to be true. Every group's examples
will go on the screen.

## Optional, can be done at home

Practice for afterward: problems 3 to 13, then coding A, B and C.
Everything can be done by hand with $\Phi(2) = 0.977$,
$\Phi(1.82) = 0.966$ and $e^{-0.46} = 0.63$. Attempt each part before
you open a hint.

::: {.callout-warning title="Hints"}
Do not open a hint unless you have been stuck for at least a few
minutes, and open them in order.
:::

### 3. A $p$-value of 0.01

A report of a trial says: "$p = 0.01$, so there is a 1% chance that the
treatment does nothing, and a repeat of the trial would be significant
99 times in 100."

(a) Say what the $p$-value is the probability of. Which two of the
six statements from the lecture has the report made?

(b) What does the chance that a repeat of the trial is significant
depend on? Suppose the repeat is run at level $0.01$, with power $0.5$
against the effect the trial was designed to detect. Give the chance
if the treatment has that effect, and if it does nothing.

::: {.callout-tip collapse="true" title="Hint 1"}
A $p$-value is a probability computed on the assumption that the null hypothesis is true. The probability of what event?
:::

### 4. Aggregation and $R^2$

People $i$ live in areas $j$, $n$ in each. $X_{ij} = B_j + W_{ij}$,
where $B_j$ is shared by everyone in area $j$ and has variance
$\sigma_B^2$ across areas, and $W_{ij}$ has mean zero and variance
$\sigma_W^2$ and is uncorrelated with $B_j$ and across people.
$Y_{ij} = \alpha + \beta X_{ij} + \varepsilon_{ij}$, where the errors
have variance $\sigma^2$, are uncorrelated across people, and have
mean zero given the $X$ values. Take $\beta = 1$, $\sigma_B^2 = 1$,
$\sigma_W^2 = 4$ and $\sigma^2 = 20$.

(a) Compute the population $R^2$ of the regression of $Y$
on $X$ across people, and of $\bar Y_j$ on $\bar X_j$ across areas, for
$n = 25$ and for $n = 400$.

(b) Is the difference between the two $R^2$ values a
matter of bias or of variance? Say what happens to the slope.

::: {.callout-tip collapse="true" title="Hint 1"}
Write the area average as $\bar X_j = B_j + \bar W_j$. Which of the two terms does averaging shrink?
:::

::: {.callout-tip collapse="true" title="Hint 2"}
The variance of an average of $n$ uncorrelated terms, each with variance $\sigma^2$, is $\sigma^2/n$. Use it for $\bar W_j$ and for $\bar\varepsilon_j$.
:::

### 5. Ecological regression and constancy

In each area $i$, a fraction $x_i$ of the people belong to a group. The
group's rate of some outcome is $p_i$ and everyone else's rate is
$q_i$, so the area's rate is $y_i = p_i x_i + q_i(1 - x_i)$. Ecological
regression fits $y_i = a + b x_i$ and reads $\hat a$ as everyone else's
rate and $\hat a + \hat b$ as the group's rate.

Two areas of equal size, $x = 0.1$ and $x = 0.3$. The
group's rate is $p = 0.2$ in both; everyone else's rate is $q = 0.30$
in the first and $q = 0.40$ in the second. Over the two areas
together, the group's rate is $0.2$ and everyone else's is $0.344$.

Compute the two area rates, fit the line through the two points, and
read off what ecological regression reports for the group and for
everyone else. Compare with $0.2$ and $0.344$.

::: {.callout-tip collapse="true" title="Hint 1"}
Use the accounting identity $y = p x + q(1 - x)$ in each area.
:::

::: {.callout-tip collapse="true" title="Hint 2"}
Two points fix a line: the slope is $(y_2 - y_1)/(x_2 - x_1)$, then the intercept. The regression's rate for the group is the height of the line at $x = 1$.
:::

### 6. Aggregation and $R^2$ with an area effect

The model of problem 4.

(a) Now let each area have an effect of its own, $\varepsilon_{ij} = u_j + e_{ij}$
with $\operatorname{Var}(u_j) = 1$ and $\operatorname{Var}(e_{ij}) = 20$.
$u_j$ and $e_{ij}$ are uncorrelated with each other and with the $X$
values, and $e_{ij}$ is uncorrelated across people. Compute the $R^2$
across people, the $R^2$ across areas at $n = 100$, and its limit as
$n$ grows.

(b) A report shows a regression across 50 regions with $R^2 = 0.9$.
What does that tell you about the $R^2$ of the same relationship across
people? What would you need to know?

::: {.callout-tip collapse="true" title="Hint 1"}
The area effect $u_j$ is the same for everyone in the area. What does averaging do to it?
:::

### 7. Ecological regression in general

The setting of problem 5.

(a) Now in general: $p_i = p$ and $q_i = q_0 + \gamma x_i$. Substitute
into the accounting identity and show that when the shares are small
the fitted slope is approximately $p - q_0 + \gamma$. What does the
regression then report as the group's rate, and which way is it wrong
when $\gamma < 0$? In problem 5, $\gamma = 0.5$. How good is the
approximation there?

(b) Freedman's neighborhood model puts $\hat p_i = \hat q_i = y_i$.
Apply it to the two areas in problem 5. Which model is closer to the truth
here, and what would you need to know to say which is closer in a case
where the truth is unknown?

::: {.callout-tip collapse="true" title="Hint 1"}
Substitute $q_i = q_0 + \gamma x_i$ into $y_i = p x_i + q_i(1 - x_i)$ and collect the powers of $x_i$.
:::

### 8. Two slopes

Split the covariance and the variance into between-area and within-area
parts: $\operatorname{Cov}(X, Y) = \operatorname{Cov}_B + \operatorname{Cov}_W$
and $\operatorname{Var}(X) = \operatorname{Var}_B + \operatorname{Var}_W$.
Suppose $\operatorname{Cov}_B = -1$, $\operatorname{Var}_B = 2$,
$\operatorname{Cov}_W = 4$ and $\operatorname{Var}_W = 8$.

(a) Compute the slope across area averages, $b_B$, the
pooled slope within areas, $b_W$, and the slope across people,
$b_{\text{all}}$. Check that $b_{\text{all}} = \lambda b_B + (1 - \lambda) b_W$
with $\lambda = \operatorname{Var}_B/(\operatorname{Var}_B + \operatorname{Var}_W)$.

(b) Keep the two variances and the within-area covariance. How
negative does the slope across area averages have to be for the slope
across people to be negative?

(c) In Gelman, Shor, Bafumi and Park (2007), within states richer
voters were more likely to vote Republican, while richer states voted
Democratic. From those two facts alone, can you say which way income
and vote were related across all voters? The authors add that income
varies far more within states than average income does between
states. What does that settle, and what does it leave open?

::: {.callout-tip collapse="true" title="Hint 1"}
A slope is a covariance divided by a variance. Compute it between areas, within areas, and for the two together.
:::

### 9. The method of bounds

This is Freedman's section 4, optional in the reading list; the notes
state the bounds.

(a) In a district, 8% of residents belong to a group and
1% of all residents have some outcome. Write the accounting identity
in the group's rate $p$ and everyone else's rate $q$, and find the
range of values each can take.

(b) In Freedman's example (Washington, 1995) the group was 7.9% of
adults and 34.4% of adults had the outcome. Find the bounds on the
group's rate there, compare them with yours in (a), and explain what
you find. State a rule for when the bounds on a group's rate are the
whole of $[0, 1]$.

(c) A precinct has $x = 0.5$ and $y = 0.6$. Bound both rates.

(d) Say, in one sentence each, what the bounds settle and what they
leave open.

::: {.callout-tip collapse="true" title="Hint 1"}
Set the other rate to 0 and then to 1, and check whether the rate you want stays between 0 and 1.
:::

### 10. The largest of $n$

$M_n$ is the largest of $n$ independent draws from a distribution with
distribution function $F$, so $P(M_n \le x) = F(x)^n$. For an
exponential with mean 1, $E[M_n] = 1 + \tfrac12 + \cdots + \tfrac1n$;
the notes derive it.

(a) Compute $E[M_n]$ for four draws from Uniform$(0, 1)$,
and for four draws from an exponential with mean 1. Then compute the
probability that the largest of 20 standard normals exceeds 2.

(b) Derive $E[M_n] = n/(n + 1)$ for the uniform.

(c) Two groups post items, and each item's extremity is an
independent draw from the same continuous distribution. One group
posts 200 items a day and the other 1,800. A feed shows the most
extreme item from each. (i) With what probability is the larger
group's item the more extreme of the two? (ii) In the notes the two
groups posted 100 and 10,000 items. The expected largest of 100
standard normals is $2.51$ and of 10,000 is $3.85$. Say what those two
numbers do and do not say about the two groups. (iii) Now both groups
post 1,000 items a day, and extremity is normally distributed with
the same mean in both groups and a standard deviation 10% larger in
the second. The expected largest of 1,000 standard normals is $3.24$.
What are the two expected shown items, in the first group's standard
deviations above the mean?

(d) (Harder.) A journal publishes an estimate only when its $z$-value
exceeds $1.96$; a feed shows the largest of 100 $z$-values. All the
$z$-values are true nulls. Compare the expected published value with
the expected shown value, using $E[Z \mid Z > 1.96] = 2.34$ from the
significance filter, and say in one sentence what the two selections
have in common.

::: {.callout-tip collapse="true" title="Hint 1"}
The largest of the draws is below a value when all of them are.
:::

### 11. Spreads and tails (harder)

Two groups have normally distributed scores with the same mean. Group
2's standard deviation is 10% larger than group 1's. A prize goes to
everyone more than two of group 1's standard deviations above the mean.

(a) Using $\Phi(2) = 0.977$ and $\Phi(1.82) = 0.966$, compute the
fraction of each group who win a prize, and the ratio. If the groups
are the same size, what share of the prize-winners are from group 2?

(b) Hyde and Mertz (2009) report, from US state mathematics tests, a
mean difference $d$ between boys and girls of $-0.02$ to $0.06$ and a
variance ratio of $1.11$ to $1.21$ (their Table 1); from one state's
test, a boys-to-girls ratio above the 99th percentile of $2.06$ among
white students and $0.91$ among Asian-American students; and from the
2003 PISA test, variance ratios from $0.95$ to $1.24$ across countries
(their Table 2). Which of these are claims about averages, which about
spreads, and which about tails? Using part (a) and Table 2, say in two
sentences how far a tail ratio can move when the variance ratio
moves.

(c) A commentator says: group 2 is 60% of the prize-winners, so group
2 should be expected to hold 60% of the jobs that need a score that
high. List what has to be assumed (i) to compute the 60% and (ii) to
read it as a share of jobs. Then say what further step would be needed
to conclude anything about one applicant.

::: {.callout-tip collapse="true" title="Hint 1"}
Express the cutoff in each group's own standard deviations.
:::

### 12. Ecological inference from a table of boroughs

A newspaper prints, for each of the 32 London boroughs, the share of
adults with a degree and one party's share of the votes cast, notes
that the correlation across boroughs is 0.6, and headlines it as a
fact about how graduates vote.

(a) Say what the table identifies and what it does not.

(b) The identity $y = p x + q(1 - x)$ needs $x$ and $y$ to be shares of
the same people. Say why they are not here, and what would have to be
assumed about who votes for the identity to be used.

(c) Granting that, state the assumption under which a regression
across boroughs gives graduates' support for the party, and say what
happens to the estimate if non-graduates' support rises with the
graduate share of the borough.

(d) Say what the Duncan--Davis bounds would add, and what you would
need to know about a borough to say whether its bounds on graduates'
support are informative.

::: {.callout-tip collapse="true" title="Hint 1"}
Ask what the units of the table are, and whether $x$ and $y$ are shares of the same people.
:::

### 13. Three claims

Three claims from this week: the main claim of the study you skimmed;
"the foreign-born have higher incomes," read off Freedman's line; and
Summers's sentence about the available pool. For each: is it about
averages, spreads or tails? About people or about places? Who was in
the data, and how did they get there? (No verdict on whether the study
you skimmed holds up; that is not the question.)

### Coding

Open a fresh R script. You will need `ggplot2` and `dplyr`.

#### A. Aggregation and $R^2$

Fifty areas, one hundred people in each. Areas differ in $X$; people
differ more.

```{r coding_a_setup}
J <- 50
n <- 100
area_mean_x <- rnorm(J, 0, 1)
people <- data.frame(area = rep(1:J, each = n)) |>
  mutate(X = area_mean_x[area] + rnorm(J * n, 0, 2),
         Y = X + rnorm(J * n, 0, sqrt(20)))
```

1. Regress `Y` on `X` across people. Record the slope and $R^2$.
2. Average `X` and `Y` within each area, regress the area averages,
   and record the slope and $R^2$. Compare with what the formula from
   problem 4 gives at $n = 100$.
3. Plot the people as faint points and the area averages as large
   points, with both regression lines. Describe in one sentence what a
   reader shown only the area averages would conclude.

::: {.callout-tip collapse="true" title="Hint 1"}
For item 2: `areas <- people |> group_by(area) |> summarize(X = mean(X), Y = mean(Y))`,
then `summary(lm(Y ~ X, areas))$r.squared`.
:::

4. Add an area effect: `Y + rep(rnorm(J, 0, 1), each = n)`. Repeat 1
   and 2. Compare with problem 6(a).
5. Make the within-area covariance positive and the between-area
   covariance negative: generate `Y` as
   `-0.5 * area_mean_x[area] + 0.5 * (X - area_mean_x[area]) + noise`.
   Check that the two slopes have opposite signs.

#### B. The largest of $n$

1. Simulate the largest of $n$ standard normals, 2000 times, for
   $n = 1, 10, 100, 1000$. Compute the mean for each $n$ and compare
   with $2.51$ at $n = 100$ and with $\sqrt{2 \ln n}$.
2. Plot a histogram of the 2000 maxima at $n = 100$. Where is it
   relative to zero?
3. Check $E[M_n] = n/(n+1)$ for the uniform at $n = 4$ and the harmonic
   sum for the exponential at $n = 4$ by simulation.
4. A group has 1000 posts a day. Estimate by simulation the fraction of
   days on which its most extreme post exceeds 3, and compare with
   $1 - 0.99865^{1000}$.

::: {.callout-tip collapse="true" title="Hint 1"}
`replicate(2000, max(rnorm(100)))` gives 2000 maxima at $n = 100$.
:::

#### C. Ecological regression on simulated precincts

Fifty precincts of equal size. The group's share $x_i$ runs from 0.05
to 0.5. The group's rate is $p = 0.3$ everywhere; everyone else's rate
is $q_i = 0.4 + 0.4\,x_i$.

1. Compute the precinct rates $y_i$ and fit `lm(y ~ x)`. Read off the
   two rates the regression reports and compare with the truth over
   all precincts.
2. Compute the Duncan--Davis bounds for $p_i$ and $q_i$ in each precinct
   and plot them against $x_i$. For which group and which precincts are
   they informative?
3. Set $q_i = 0.4$ (constancy holds) and repeat 1.

