The world is too loud. Read what matters.

Very Bad Wizards

Ranking Psychology's Classic Experiments: Many Conclusions Don't Survive Replication

The Milgram shock experiment still replicates, which is itself absurd; and the Good Samaritan experiment had only 40 subjects total, so even if they got it right, they got it right by accident.

PsychologyReplication CrisisResearch EthicsBehavioral ScienceMethodology
Two philosophers plus a psychologist rank classic studies from D to S, and along the way settle accounts with the replication crisis, Zimbardo's fraudulent original notes, and small samples.

The argument · tap a timestamp to hear it

10:08

Milgram still replicating is itself absurd

Milgram recruited New Haven workers in a Yale basement, got them to believe they were shocking another person, with the voltage labeled up to 300 volts and marked dangerous; the person being "shocked" was actually a trained actor, and about 65% of them pressed all the way to the end. He ran a large number of variants, manipulating whether the experimenter wore a white coat, the distance to the person being shocked, and which university's name was on the door. The study directly rewrote psychology's ethics review system; today this kind of experiment can't be done. Dave says that a social psychology study from the 50s and 60s that still replicates is itself "[ __ ] banana," which is why it goes in the S tier — he's using replication as the bar, not fame.

— Dave Pizarro
29:20

The Good Samaritan experiment had only 40 people

Darley and Batson had seminary students go to the other side of campus to speak on the parable of the Good Samaritan, and arranged for a visibly injured, groaning person asking for help along the way; the manipulated variable was the time pressure they were under. It turned out time pressure predicted whether they stopped to help better than what topic they were going to speak on. But Dave points out that the whole experiment had only 40 subjects total, and after splitting into groups the samples were tiny. He groups this kind of study with the telephone-booth coin experiment as a common disease of 60s and 70s social psychology: severely underpowered, "even if they got it right, they got it right by accident." Paul says it's great as material for making people reflect on human nature, but as a scientific finding he doesn't trust it, and gives it a D.

— Dave Pizarro
48:38

On the cloth mother, mothers already knew

Harlow separated infant monkeys from their mothers and built two surrogate mothers: a wire frame with a bottle attached, and a soft, huggable cloth one that provided no milk. Except when they had to nurse, the infants stayed on the cloth mother the rest of the time. This refuted the behaviorist explanation that "infants love their mothers because the mother provides food," and supported the role of contact comfort. The episode clarifies that Harlow did not actually shock the monkeys; what's hard to watch is the footage of frightened infants rushing to the cloth mother and clinging on for dear life. He also called the isolation apparatus the "pit of despair" and the restraint device the "rape rack" in published articles, and the reaction on the show was "he's just a psychopath." Tamler argues for the lowest score, not because of cruelty, but because these studies didn't show anything any mother would know by intuition the moment you asked her.

— Tamler Sommers
1:10:52

The prison experiment's author was coaching guards to abuse

The original notes and recordings dug up on the Stanford prison experiment show that Zimbardo himself was coaching the people randomly assigned to be guards to engage in abusive behavior. If the intended story of the study is "put people in this role and they spontaneously become sadists," then on the available evidence that conclusion is highly suspect — he didn't report how much he pushed them. David thinks this is close to scientific fraud, groups it with "the dog torturer," and no longer teaches it. But David also cautions: across the Milgram and Asch series, people differ, and some said "screw you, I'm not doing it" at the lowest shock level. Tamler adds that when he interviewed Zimbardo, the latter talked about Milgram, saying Milgram did no psychological screening and some of the people who showed up might have been sociopaths, while Zimbardo claimed he did screen, using that to prove these were normal people.

— David
1:15:59

Even adults fail the false belief task

The false belief task comes from Premack and Woodruff's paper "Does the Chimpanzee Have a Theory of Mind"; Dennett proposed this design in his commentary: Sally puts a toy in a box and leaves, Anne moves it to a basket, and you ask where Sally will look for it when she returns. Children before age three or four say the place where the toy actually is, and around age five they start getting it right; this shift from failure to success is one of the most robust findings in all of developmental psychology. But Paul and Susan Birch tested adults and found adults find it hard too — there's an immediate intuition pointing to "where the thing actually is" that has to be inhibited; he demonstrates it live in his intro course, and about a third of the students get it wrong. What bothers Tamler is that once the paradigm got hot, the whole field shifted from that interesting question to the methodology itself, and the test became the construct.

— Paul Bloom
1:24:06

To deny the methods are all broken, you must accept ESP

Daryl Bem published a paper in a top social psychology journal in 2011 claiming to have found precognition: subjects first rated their preference for unfamiliar Chinese characters, and only afterward were shown emotional images as primes; the characters that were positively primed afterward had been rated higher beforehand. He made all his materials and data public, and Paul thinks he genuinely believed the result. Bem's defense is a dilemma: if you say my research isn't good research, then almost all research in social psychology isn't good research — so either admit your methods are reliable and accept ESP, or admit the methods are unreliable. Paul says he was furious at the time: this would overturn our entire understanding of space and time, and you want to accept it on a p less than 0.05. This is also the driving force behind the Bayesian countermovement in psychology: results should be compared against prior plausibility, not against 0.05.

— Paul Bloom
1:29:09

Delayed gratification may measure trust, not self-control

Mischel and Ebbesen had preschoolers choose between taking a small reward immediately and waiting for two. What's more interesting in the original study is that some children spontaneously used strategies — thinking of the marshmallow as a white cloud, singing, distracting themselves; Tamler says they were "tying themselves to the mast." The problem comes later: following these children and claiming the wait time predicts SAT scores, and Tamler just read that the SAT correlation used only 35 children, yet schools started teaching delayed gratification on that basis. He dismantles the premises that "delayed gratification causes success" would need: it has to be teachable, rather than merely capturing some natural difference that already predicts on its own; self-control has to be the active ingredient, rather than the willingness to please the teacher; and there's the question of why the child believes the other person will really give another one — if you're in an impoverished environment, why wait.

— Tamler Sommers
1:32:12

Social priming collapsed, and conceptual priming was buried with it

Bargh's classic study: subjects were randomly assigned a set of words to unscramble, and if those words were related to concepts like elderly and retirement, the speed at which they walked from room A to room B after the task was measured, and it turned out those primed with the elderly concept walked more slowly. It claims that stimuli entirely outside consciousness can influence behavior — you don't even know you were assigned elderly words. Tamler recalls that when he had just finished his degree, this kind of research was at its peak. Paul says social priming has since been shown to be wrong, but what really bothers him is the consequence: its collapse took down with it conceptual priming (for example, flashing "doctor" to prime "nurse") and other very robust, easily replicated findings — that's what the architecture of the mind actually looks like; the problem was extrapolating it directly to effect sizes in behavior and everyday life.

— Paul Bloom

In their own words · checked verbatim

for a social psych study from the 50s or 60s to replicate is already [ __ ] banana

Dave Pizarro11:08

if it's supposed to prove something like the situation is much more influential than character that seems like a stretch that you could do that with these experiments

Paul Bloom32:22

I think this does border on scientific fraud, not providing this information about the extent to which we were coached to doing this.

David1:12:54

So the test has become the construct.

Tamler Sommers1:18:00

if you think my studies aren't good studies, then that's true for just about everything else in social psychology.

Paul Bloom1:24:06

this would upend our knowledge of what spaceime is, you know, like this would literally shake the core understanding of of all of our models of how the universe works. And you're going to take a like P of less than 0.05

Paul Bloom1:25:06

I just read that the the SAT correlation was done with 35 kids.

Tamler Sommers1:29:09

one of the things that kind of upsets me about the whole priming, social priming work is it undermined the very real and robust findings of conceptual priming.

Paul Bloom1:34:13

Figures

Share who pressed all the way to the end in the Milgram experiment65%10:08
Total subjects in the Good Samaritan experiment4030:21
Error rate for Asch conformity in a high-powered replication33% (210 subjects), versus 37% in Asch's original study55:42
Infants willing to keep crawling over the edge in the original visual cliff study3 out of 271:07:50
Share of adults in the intro course who got the false belief task wrongabout one third1:15:59
Year Bem's ESP paper was published20111:20:02
Probability above chance for Bem under one condition52%1:25:06
Sample size for the marshmallow test's SAT correlation35 children1:29:09
Final S-tier studiesthe Milgram shock study, the invisible gorilla, the Stroop effect1:36:13

Glossary

replication crisis
A crisis of confidence in which a large number of published psychology results cannot be reproduced.
false belief task
Tests whether a person can understand that someone else holds a belief that conflicts with the facts one knows.
theory of mind
The ability to infer others' mental states and predict their behavior on that basis.
social priming
The hypothesis that unconscious stimuli (such as elderly-related words) can change actual behavior.
precognition
The claim that future events can influence judgments in the present; Bem used it to explain his ESP results.
construct
An abstract psychological variable operationalized in research, such as "theory of mind."

How to listen

Who it's for

People who need to cite classic psychology findings in articles, slides or decisions, and anyone who wants to know which classics can no longer be trusted.

Skip

The Derren Brown envelope-switch and mentalist segment (from 40:33) can be skipped.