ST313 course page
Ethics for data science: the statistical choices hidden inside technical steps, and who they affect.
Each week has one lecture and one class (80 minutes each, lecture first). The notes are the main reference; the slides are what is shown in the lecture; the class sheet is what you work through in class, and its .qmd link downloads the source to run the R code. Bring a laptop with R and RStudio to the class.
Week 1: Selection and the replication crisis
Why a literature of honest tests can still be wrong on average: power, the significance filter, and what a published estimate is conditioned on. The standing question, with two cases.
Reading: ms4ds, chapter 3: Ethical data science, sections 3.1 to 3.6 · Wikipedia: Replication crisis, skim
Class: class sheet (.qmd)
Week 2: Mindless statistics and sophisticated bigotry
Cargo cult science, cargo-cult statistics and the null ritual, and Feynman’s first principle: you must not fool yourself. Then numbers that are computed correctly and read as answers to other questions: a correlation across fifty states with the opposite sign among the people who live in them, what \(R^2\) measures and what averaging does to it, decisions about a person made from the average of their group, what a small difference in spread does far out in a tail, and what the largest of \(n\) items shows you of a group.
Required reading: Stark and Saltelli, Cargo-cult statistics and scientific crisis (in full) · Hyde and Mertz, Gender, culture, and mathematics performance (first paragraph of “Do Gender Differences Exist Among the Mathematically Talented?”, Figure 1, Table 2; free full text) · Freedman, Ecological Inference and the Ecological Fallacy (sections 1 to 3) · one of two psychology studies, as announced (Connellan and colleagues 2000, or Alexander and colleagues 2009), skimmed for its data and analysis
Optional: Rathje, Van Bavel and van der Linden, Out-group animosity drives engagement on social media · Gigerenzer and Marewski, Surrogate Science
Class: class sheet (.qmd)