Why not use a
standard linear
regression?
Normal distribution has a support of (-∞,∞), but we know the outcome variable takes on only two values.
Under some circumstances, results can be interpreted as proportions or probabilities, but this can lead to predicted values less than zero or more than one.
Why not use a
standard linear
regression?
Ronald Fisher
statistician and eugenicist
Joseph Berkson
statistician and tobacco apologist
Inverse Logit transformation
Takes values between 0 and 1, and turns
them into values between -∞ and ∞.
Takes values between -∞ and ∞, and turns
them into values between 0 and 1.
⇔
| x | logit–1(x) |
|---|---|
| -4 | 0.018 |
| -2 | 0.119 |
| -1 | 0.269 |
| 0 | 0.500 |
| 1 | 0.731 |
| 2 | 0.881 |
| 4 | 0.982 |
Why this model instead of the model we built in the first week of class?
Logistic regression allows us to include explanatory covariates.
Pr(α)
logit–1(Pr(α))
| Median | 95% C.I. | |
|---|---|---|
|
|
-3.34 | (-3.48, -3.20) |
|
|
0.036 | (0.031, 0.041) |
|
|
0.034 | (0.030, 0.039) |
Figures by Peter McMahan (source code)
Promotional image for Pretty in Pink (1987)
Ronald Fisher, via The Science History Institute
Joseph Berkson, via The Rochester Epidemiology Project
James Spader in Pretty in Pink
James Spader in Pretty in Pink