All four datasets share the same mean, standard deviation, correlation, and regression line. Drag any point and watch how far the statistics move from the originals.
| Point | x | y |
|---|---|---|
| 1 | 10.00 | 8.04 |
| 2 | 8.00 | 6.95 |
| 3 | 13.00 | 7.58 |
| 4 | 9.00 | 8.81 |
| 5 | 11.00 | 8.33 |
| 6 | 14.00 | 9.96 |
| 7 | 6.00 | 7.24 |
| 8 | 4.00 | 4.26 |
| 9 | 12.00 | 10.84 |
| 10 | 7.00 | 4.82 |
| 11 | 5.00 | 5.68 |
A plain linear relationship with ordinary scatter.
| Mean y | SD y | r | Slope |
|---|---|---|---|
| 7.50 | 2.03 | 0.82 | 0.50 |
| Point | x | y |
|---|---|---|
| 1 | 10.00 | 9.14 |
| 2 | 8.00 | 8.14 |
| 3 | 13.00 | 8.74 |
| 4 | 9.00 | 8.77 |
| 5 | 11.00 | 9.26 |
| 6 | 14.00 | 8.10 |
| 7 | 6.00 | 6.13 |
| 8 | 4.00 | 3.10 |
| 9 | 12.00 | 9.13 |
| 10 | 7.00 | 7.26 |
| 11 | 5.00 | 4.74 |
A clean curve. A straight line is the wrong model entirely.
| Mean y | SD y | r | Slope |
|---|---|---|---|
| 7.50 | 2.03 | 0.82 | 0.50 |
| Point | x | y |
|---|---|---|
| 1 | 10.00 | 7.46 |
| 2 | 8.00 | 6.77 |
| 3 | 13.00 | 12.74 |
| 4 | 9.00 | 7.11 |
| 5 | 11.00 | 7.81 |
| 6 | 14.00 | 8.84 |
| 7 | 6.00 | 6.08 |
| 8 | 4.00 | 5.39 |
| 9 | 12.00 | 8.15 |
| 10 | 7.00 | 6.42 |
| 11 | 5.00 | 5.73 |
A perfect line plus one outlier that drags the fit off it.
| Mean y | SD y | r | Slope |
|---|---|---|---|
| 7.50 | 2.03 | 0.82 | 0.50 |
| Point | x | y |
|---|---|---|
| 1 | 8.00 | 6.58 |
| 2 | 8.00 | 5.76 |
| 3 | 8.00 | 7.71 |
| 4 | 8.00 | 8.84 |
| 5 | 8.00 | 8.47 |
| 6 | 8.00 | 7.04 |
| 7 | 8.00 | 5.25 |
| 8 | 19.00 | 12.50 |
| 9 | 8.00 | 5.56 |
| 10 | 8.00 | 7.91 |
| 11 | 8.00 | 6.89 |
Every x is identical except one. That single point sets the slope by itself.
| Mean y | SD y | r | Slope |
|---|---|---|---|
| 7.50 | 2.03 | 0.82 | 0.50 |
Francis Anscombe published these four datasets in 1973 to make a point about statistical practice, not about arithmetic. Numerical summaries compress, and compression discards. Each dataset here has the same mean, the same standard deviation, the same correlation, and the same fitted line, yet one is linear, one is a curve, one is a line with a single outlier, and one is a vertical stack whose slope is decided by one point. A table of coefficients cannot tell them apart. A ten second glance at a plot can.
In 2017 Matejka and Fitzmaurice generalized the idea with the Datasaurus Dozen, using simulated annealing to move points toward arbitrary target shapes while holding the summary statistics fixed to two decimals. One of the datasets is a dinosaur. The lesson is the same and the demonstration is harder to dismiss as a curiosity of hand-picked numbers.