Twelve radiologists looked the gorilla in the face and none of them saw it

We accept readily enough that a distracted spectator can miss a man in a gorilla suit walking across a basketball game. We accept it far less readily when the observer is a professional, trained for years, examining an image from his own field. The intuition is a solid one: expertise is supposed to build a wider gaze, one that catches the anomaly even when it is not looking for it. In 2013, three researchers at the Visual Attention Lab, attached to Harvard Medical School and the Brigham and Women's Hospital, put that intuition to the test by dropping a gorilla into a chest CT scan. Twenty radiologists out of twenty-four failed to report it. And the most unsettling measurement is not that failure rate: it is that twelve of them had looked the gorilla straight in the face.
A gorilla the size of a matchbox
The study is signed by Trafton Drew, Melissa Võ and Jeremy Wolfe, under the title "The Invisible Gorilla Strikes Again: Sustained Inattentional Blindness in Expert Observers", posted online on 17 July 2013 and then published in the September issue of Psychological Science, volume 24, number 9, pages 1848 to 1853.
The protocol does not cheat. Twenty-four radiologists, average age 48, ranging from 28 to 70: nine tested at the Brigham and Women's Hospital in Boston, fifteen American Board of Radiology examiners tested at their meeting in Louisville, Kentucky. Each one gets a maximum of three minutes per case to scroll freely through five chest CT scans and click on the lung nodules, roughly ten nodules per case on average. This is exactly their job: hunting for small bright discs in a stack of axial slices.
On the last case, the authors insert an image of a gorilla. It measures 29 by 50 millimetres, it carries a white outline, and its opacity climbs from 50 to 100 percent before dropping back to 50 percent across five 2 millimetre slices, so as to avoid an abrupt appearance that would have drawn the eye on its own. The full case runs to 239 slices, within the usual range of 100 to 500 images in a chest stack. The volume of the rectangular box that would contain the gorilla exceeds 7,400 mm³, roughly a matchbox. In these images, taken from the Lung Image Database Consortium, the average volume of a nodule is 153 mm³. The gorilla therefore takes up more than 48 times the volume of the target the radiologists are busy hunting. On screen, it covers 0.9 by 0.5 degrees of visual angle in Boston, 1.3 by 0.65 in Louisville.
Twenty out of twenty-four
At the end of the last case, the authors ask three escalating questions: did this case seem different from the others, did you notice anything unusual, did you see a gorilla. Twenty radiologists out of twenty-four report it at none of the three. That is 83 percent.
The obvious objection would be a visibility problem. It collapses immediately: when the image is shown to them after the experiment, all twenty-four see the gorilla. The formal control comes in experiment 3, where twelve naive observers watch the same stack roll past as a film, with the gorilla present half the time, and a circular cue marking its possible location. At 35 or 70 milliseconds per image, the gorilla is visible for only 175 or 350 milliseconds. They reach 88 percent correct answers, with no measurable effect of the frame rate. A fifth of a second is enough when you know you are looking for it.
The radiologists had far more than that. They scrolled through the layer containing the gorilla 4.3 times on average. Those who missed it spent 5.8 seconds on average on the five slices concerned, with durations ranging from 1.1 to 12 seconds. Time was not the limiting factor.
Looking is not seeing
The eye tracking is the centrepiece of the case file. Among the twenty radiologists who did not report the gorilla, twelve fixated directly on its location while it was visible. In those twelve, the gaze stayed on the zone for 547 milliseconds on average. More than half a second of gaze resting in the right place, on a shape whose volume is more than 48 times that of the target being sought, and nothing came through to awareness.
That is the operational definition of inattentional blindness, and it draws a clean line between two things ordinary language runs together: the image lands on the retina, the gaze goes to the right place, and yet the object is not perceived. What the study establishes is that neither acuity, nor exposure time, nor size is enough to account for the failure. What it does not measure, on the other hand, is the exact weight of what the observer was looking for. The authors put forward a hypothesis there without testing it: the radiologists were hunting small bright nodules, and an albino gorilla might have been spotted more readily because it would have matched the light polarity of the targets more closely. They push the logic all the way to its counterintuitive flip side: a smaller gorilla might have been detected more often, because it would have looked more like a nodule. Neither of those two variants was ever put to the experiment.
The rival explanation the data does not support
In radiology there is a well known rival candidate for explaining this kind of oversight: "satisfaction of search", the idea that having found one abnormality inhibits the detection of the ones that follow. The gorilla happened to sit on a slice carrying a nodule, detected by 71 percent of the radiologists. The explanation looks ready made.
The authors put it up against their data, and its two predictions do not hold. If satisfaction of search were driving the effect, the radiologists who missed the neighbouring nodule should have seen the gorilla more often, and those who had found the nodule should have seen it less often. Yet of the seven radiologists who missed the nodule, not one saw the gorilla. And the four who did see the gorilla had all seen the nodule on that same slice. The relationship runs the opposite way from the one the hypothesis would require.
This refutation has to be read for what it is. It rests on eleven observers split into two handfuls, seven on one side, four on the other. And the authors write in plain terms that, in the absence of a further experiment measuring gorilla detection with no neighbouring nodule, it is hard to establish with certainty the role played by that nodule. The rival explanation is weakened on its own predictions; it is not ruled out.
The experts do better, not worse
This is where the popular version of the story turns over on itself. Experiment 2 runs the protocol again with naive observers who have no medical training, average age 33.7, after about ten minutes of learning to recognise a nodule. Not one reports the gorilla: 100 percent failure, against 83 percent among the experts. The paper recruits 25 people and counts 24 in its results sentence, an internal inconsistency that does not shift the finding, since the detection rate is zero either way.
So the radiologists do better than the novices on the gorilla, but the margin is thin and deserves to be announced as such: Fisher exact test, p = 0.0497, the conventional threshold grazed from just underneath. The whole gap comes down to four detections against zero, and the sample size that test puts in play is precisely the one the paper gives at two different values. The result is worth citing, not worth loading up.
On the task they are paid for, on the other hand, the gap is massive: 55 percent of nodules detected against 12 percent. The novices had nevertheless spent a comparable amount of time in the zone, 4.9 seconds on the images containing the gorilla, and nine of the twenty-five fixated on its location. On those two measures, the authors find no statistically significant difference from the radiologists.
Put another way, the study does not show that expertise blinds. It shows that expertise, which massively improves performance on the intended task and which, at worst, does not degrade the detection of the unexpected, is not enough to cancel the limit. The authors say so without ambiguity: it would be a mistake to read these results as an indictment of radiologists, whose work is an unusually demanding visual search task. The message is that this level of expertise grants no immunity against the inherent limits of human attention and perception.
What 1999 already said
The most famous experiment in the field is the one by Daniel Simons and Christopher Chabris, "Gorillas in Our Midst: Sustained Inattentional Blindness for Dynamic Events", Perception, volume 28, number 9, 1999, pages 1059 to 1074. Famous does not mean founding, and that is the first thing the popular version quietly drops. The basketball game setup comes from Ulric Neisser and his collaborators, in the 1970s. The woman with the umbrella who walks across the court is their find, not the find of Simons and Chabris, who borrow it and say so. The term inattentional blindness itself is credited in the article to Mack and Rock (1998). What Simons and Chabris contribute is a controlled version of the paradigm: sixteen conditions, and a gorilla.
The figure that circulates, "half of people do not see the gorilla", is an approximation. The real data is more interesting. Of 228 participants, 36 are set aside, most of them because they already knew about the phenomenon or had lost count of the passes. The remaining 192 are split evenly across sixteen conditions, twelve per condition. Across the whole set, 54 percent noticed the unexpected event and 46 percent missed it, even though it lasted five seconds in a 75 second video.
The detail the popular version erased lies elsewhere. There were two unexpected events, the gorilla and the woman with the open umbrella. The woman with the umbrella was noticed markedly more often than the gorilla: 65 percent against 44 percent. The authors do not settle the cause and leave three avenues open, her visual salience, her consistency with what one expects from a basketball scene, and her semantic proximity to the events being tracked.
The genuinely instructive contrast lies elsewhere again. When observers were tracking the team in black, they saw the black gorilla in 58 percent of cases, against 27 percent when they were tracking the team in white. Same gorilla, same video; only the instruction changes. For the woman with the umbrella, dressed in light colours that matched neither team, the team being tracked changed almost nothing: 62 percent against 69 percent. Simons and Chabris draw from this a conclusion that flatly contradicted an earlier result from Neisser: an unexpected object is noticed more readily when it shares basic visual features, here colour, with whatever attention is currently tracking.
The popular version kept the gorilla because it is funny and it makes a good picture. It lost what makes the result useful, which is not the gorilla but the colour. Even so, the bridge between the two studies is a hypothesis, not a demonstration. The 2013 gorilla was dark in an image where bright spots were being sought, exactly the unfavourable configuration identified in 1999; but the albino gorilla was never tested, and the authors are careful not to conclude. What is established, on the other hand, fits in a single line: twelve gazes landed in the right place, for more than half a second on average, and nothing happened.
