This sounds like a prank pulled on a scientific journal. It was actually a demonstration of a serious statistical problem.
In 2009, researchers Craig Bennett, Abigail Baird, Michael Miller and George Wolford reported an intentionally absurd experiment. A mature Atlantic salmon — about 18 inches long and already dead — went into an fMRI scanner. The researchers presented it with photographs of humans in social situations and, with a straight face worthy of the setup, described the task as asking the fish to determine what emotion the people in the pictures were experiencing.
Then they analyzed the scan.
And there it was: a small cluster that appeared to show a statistically significant response in the salmon’s brain cavity.
The machine was not detecting a zombie fish
The point was never that the salmon had somehow resumed thinking. The point was that researchers had asked the data an enormous number of questions at once.
An fMRI image is divided into many tiny three-dimensional elements called voxels. In the salmon demonstration, the analysis involved roughly 130,000 voxel-wise comparisons. If you test enough locations separately, some can cross a conventional statistical threshold purely by chance — even when there is no genuine effect at all.
This is the multiple comparisons problem. The more independent chances you give random noise to look interesting, the more likely it is that something will eventually look interesting.
Run 100 completely meaningless tests
Each square below gets a random number between 0 and 1. We highlight anything below 0.05 — the familiar 5% threshold. There is deliberately no real effect in this toy simulation.
This is only a visual demonstration of the logic, not a model of fMRI statistics. Real neuroimaging analyses have dependencies, spatial structure and specialized correction methods.
Why one “significant” result can be a trap
Suppose you flip a fair coin a few times. A strange-looking streak can happen by luck. Now imagine giving yourself thousands of different ways to declare a streak interesting. Even if every test begins with random noise, the collection of tests creates many opportunities for a false positive.
The salmon paper made the same idea impossible to ignore. Using an uncorrected statistical threshold, the researchers found apparent activation. When they applied formal corrections for the large number of comparisons, the apparent brain activity disappeared.
That is exactly what should happen. A dead salmon is useful here because everyone agrees on the ground truth before seeing the statistics: there should be no task-driven neural response.
The funniest part was also the most useful part
The experiment was deliberately ridiculous, but it was not an attack on brain imaging itself. Multiple-comparison corrections already existed. The demonstration was a warning about what can happen when an analysis ignores the number of statistical tests being performed.
That distinction matters. “An fMRI once detected activity in a dead fish” is a great headline. “Therefore fMRI is nonsense” would be the wrong conclusion.
The real lesson is narrower and more powerful: sophisticated equipment does not rescue weak statistical reasoning. A beautiful color map can still contain a false positive. A tiny p-value can still be misleading if you forget how many chances you gave yourself to find one.
Why the story escaped the lab
The dead-salmon example became famous because it compresses an abstract statistical warning into one image your brain refuses to forget.
You can explain family-wise error rates for twenty minutes and lose half the room. Or you can say, “They found brain activity in a dead salmon until they corrected the statistics.” Suddenly everyone wants to know what went wrong.
The work later became associated with the 2012 Ig Nobel Prize in Neuroscience, awarded to Bennett, Baird, Miller and Wolford. Ig Nobel prizes celebrate research that first makes people laugh and then makes them think. Few examples fit that description better.
This problem is much bigger than brain scans
The same logic can appear anywhere people search a large dataset for something that looks unusual: finance, genetics, A/B testing, sports statistics, psychology, medicine, marketing or machine learning.
If you slice data into enough groups, try enough outcomes, test enough dates, inspect enough variables or repeatedly change the analysis until a pattern appears, chance can start wearing a convincing costume.
That does not mean every surprising result is fake. It means the number of opportunities to be surprised matters.
A useful question to steal from a dead fish
The next time you see a chart claiming a shocking connection, ask one extra question:
“How many things did they test before finding this one?”
You may never look at a “statistically significant” headline quite the same way again.
Sources & further reading
- Bennett, Miller & Wolford — NeuroImage conference abstract, “Neural correlates of interspecies perspective taking in the post-mortem Atlantic Salmon” (2009)
- Bennett, Wolford & Miller — “The principled control of false positives in neuroimaging” (SCAN, 2009)
- Multiple comparison correction methods for magnetic resonance imaging — review
- Improbable Research — 2012 Ig Nobel Prize winners