Back to the blog
Psychology

A 22 year old student redid the Ebbinghaus experiment, and the forgetting curve held

June 1, 2026 11 min read

Illustration generated with GPT Image.
Illustration generated with GPT Image.

The forgetting curve has the status of a textbook ornament: a line that plunges and then flattens out, credited to a nineteenth century German said to have worked on a single subject, himself. The implicit reading is that a result like that belongs to the history of ideas more than to verifiable science.

Two details sharpen the suspicion. Hermann Ebbinghaus really is both the experimenter and the only subject. And the first state of his work, dated 1880, is not a publication: it is the manuscript he submitted for his habilitation, the Urmanuskript, which stayed unpublished and appeared in German only in 1983, with Passavia Universitätsverlag in Passau. The book, Über das Gedächtnis, came out in Leipzig with Duncker in 1885. The English translation, Memory: A Contribution to Experimental Psychology, arrived in 1913, signed by Henry A. Ruger and Clara E. Bussenius for Teachers College, Columbia.

In late 2011, a researcher at the University of Amsterdam and his student did the one thing that settles a debate of this kind. They ran the experiment again.

What Ebbinghaus measured, exactly

The 1913 text sets out the protocol without ambiguity. The investigations into forgetting fall in the year 1879 to 1880 and comprise 163 double tests. Each double test consists of learning eight series of 13 syllables, 104 syllables in all, then relearning them after a fixed delay. Learning continues until two consecutive errorless recitations. Thirty eight of those double tests, run between 11 in the morning and noon, had only six series.

The material is the control device that makes the measurement possible: roughly 2,300 nonsense syllables, built by placing a vowel or a diphthong between two consonants, shuffled and then drawn at random. Ebbinghaus did try other materials, but his reasons for setting them aside are not the ones usually attributed to him. Series of numbers, he writes, turned out to be impractical because their basic elements were too few and too quickly exhausted. As for the cantos of Byron's Don Juan, which he learned by heart, they gave him no wider spread of measurements than the syllables did. What he claims for his material is therefore not the absence of variability, it is homogeneity, an inexhaustible stock of fresh combinations, and the ability to vary length cleanly.

The seven intervals are stated on page 66: about a third of an hour, one hour, nine hours, one day, two days, six days, thirty one days. The values actually measured drift a little from those labels, the final table carrying 0.33 hour, 1 hour, 8.8 hours, then 24, 48, 6 times 24 and 31 times 24 hours. The measurement is not a recall score but the saving in work at relearning, expressed as a percentage of the time taken for first learning, recitation time deducted. The table gives 58.2, then 44.2, 35.8, 33.7, 27.8, 25.4 and 21.1 percent, with probable errors of the mean of 1 on the first three points, then 1.2, 1.4, 1.3 and 0.8. The number of tests per point is not constant: 12, 16, 12, 26, 26, 26, and 45 for the thirty one day point. Ebbinghaus also corrects his times according to the hour of the session, late morning values being cut by about 5 percent and those from 6 to 8 in the evening by about 12 percent, percentages he estimates on other series and spells out in his text. Murre and Dros, who read the 1880 manuscript, give 13 percent for the evening correction. They also point out that he was twenty nine years old.

Amsterdam, winter 2011

Jaap Murre and Joeri Dros submitted their replication to PLOS ONE on November 22, 2014. It was accepted on January 25, 2015 and published on July 6, 2015 under the reference PLoS ONE 10(7): e0120644. The second author, J. Dros, 22 years old, is the only subject. As with Ebbinghaus, the subject is also an author, with the difference that the protocol here is held by a third party. The experiment was reviewed and approved by the ethics committee of the psychology department at the University of Amsterdam, file 2014-BC-3879.

The training phase ran from November 8 to 29, 2011: 14 lists learned then relearned after twenty minutes, followed by 19 lists learned without relearning, to absorb the familiarization effect. The experimental phase ran from December 1, 2011 to February 13, 2012, which the authors count as 75 days. Sixty nine lists were learned and relearned in it, ten per interval except nine for the nine hour interval. Total collection time comes to about 70 hours.

The material takes over the original structure: 70 lists of 104 syllables, eight rows of 13. The syllables are consonant vowel consonant, built on Dutch phonotactics, with the removal of any that carried too much meaning, a filter Ebbinghaus did not apply. Rows are recited at 150 beats per minute, a tempo taken from Ebbinghaus, one repetition lasting 5.2 seconds on average, with fifteen seconds of pause between two rows. The stopping criterion changes: one correct recitation rather than two, the choice Ebbinghaus himself adopted in his second campaign, in 1883 to 1884, judging that the first method served its purpose only incompletely. That is not the only departure, and the authors list them: ten lists per interval instead of 12 to 45, no fixed session hour where Ebbinghaus tested his four longest intervals at three times of day, a recitation rhythm in iambs rather than in three beats, and a learning schedule whose shape, in Ebbinghaus, is unknown beyond nine hours. A check confirms that the draw did not bias the distribution: the mean number of repetitions at first learning does not differ significantly across intervals, F(6, 69) = 0.691, p = 0.658. The raw data are deposited on the Open Science Framework, at osf.io/6kfrp.

The four curves

The comparison is not between two curves but four. It includes an earlier replication, by Heller, Mack and Seitz, published in 1991 in the Zeitschrift für Psychologie, volume 199, pages 3 to 18, with two subjects whose data are carried over separately. The savings scores set side by side give, for Ebbinghaus: 0.582, 0.442, 0.358, 0.337, 0.278, 0.254, 0.211. For Mack: 0.544, 0.432, 0.285, 0.316, 0.365, 0.309, 0.258. For Seitz: 0.442, 0.325, 0.270, 0.270, 0.286, 0.205, 0.201. For Dros: 0.472, 0.373, 0.276, 0.317, 0.230, 0.168, 0.041.

The authors' verdict is explicit: given the individual differences to be expected, they find the resemblance of the four graphs remarkable. More than one hundred and thirty years separate the first series from the last, the language of the material is not the same, and once the curves are normalized on their first point, the traces largely overlap.

The one clear departure is the Dros point at thirty one days: 0.041 against 0.211 for Ebbinghaus. The authors do not hide it, but they do not settle it either. They write that they can only speculate, and offer three lines. This subject may simply forget more over the long run. The lists for the long interval were all learned at the very start of the period, for want of available time, so under massed learning where the other intervals were spread out. And learning time drifted upward across the 75 days, by 2.67 seconds per day per list, a linear regression accounting for 56.18 percent of the variance. Corrected for that drift, the thirty one day score climbs back to 0.137, which still sits well below the other three. A fourth hypothesis, the number of lists interposed between learning and relearning, is tested and set aside: it accounts for only about 8.4 percent of the variance.

The twenty four hour point

The replication does not merely confirm. It goes back over a point Ebbinghaus had seen and had not managed to place.

On his own curve, savings go from 35.8 percent at 8.8 hours to 33.7 percent at twenty four hours. The fall almost stops. Ebbinghaus explains himself at length. He puts the drop at 2.5 points over the stretch from 9 to 24 hours and at 6.1 points over the one from 24 to 48 (his own table gives rather 2.1 and 5.9), and judges it not credible that roughly three times as much should be lost over the second. Above all, he explicitly considers sleep as a cause and rules it out: even supposing the night markedly slows the fading, the shape he obtains does not seem to him to hold up. He concludes that one of the three values must be affected by chance, supposes the 33.7 is one to two points too high, then changes his mind, because observations he presents next confirm it, and declares himself in doubt.

Those observations come from the 1883 to 1884 campaign: seven tests of nine series of 12 syllables, relearned after twenty four hours, give 33.4 percent. Against 33.7. Ebbinghaus himself stresses that the two figures come from distant periods and very different investigations, which makes the agreement all the more notable. The point holds. Murre and Dros read that sequence as distrust strong enough to warrant a counter test, which Ebbinghaus, for his part, does not say in those terms.

In 1924, Jenkins and Dallenbach read the anomaly as an effect of sleep, and that is what pushed them to set up their study of forgetting during sleep and during waking, published in the American Journal of Psychology. Murre and Dros formalize the intuition: they add a constant boost term to the power function, applied to intervals of one day and longer only. Across the four data sets, mean explained variance goes from 98.7 to 99.1 percent, the sum of squared errors from 0.00932 to 0.00628, the Akaike information criterion from minus 22 to minus 23.6. The size of the jump is 0.030 for Ebbinghaus and 0.131 for Mack.

The authors stay cautious, and they deserve to be followed there. They call these results mixed themselves, and say that an AIC gap of 1.6 on average can at most be called a trend toward a difference that would matter, with large differences between subjects. The Dros data show no jump at all, their fits probably being pulled by that very low thirty one day point. And attributing the jump to sleep would require, they write, further experiments. The published conclusion says exactly that: the curve has indeed been replicated, it is not perfectly smooth, and it very probably shows an upward jump starting at the twenty four hour point.

Why the decorative version held

The timing is almost insolent. The replication appeared on July 6, 2015. Less than two months later, on August 28, Science published the Open Science Collaboration report, reference 349(6251): aac4716. Across 100 psychology studies rerun, experimental and correlational, published in 2008 in three journals, 97 percent of the originals showed a statistically significant result, against 36 percent of the replications, and the replicated effects came out at half the size of the original ones.

The parallel is tempting and it calls for a caveat, because the two exams do not measure the same thing. The Open Science Collaboration judges a reproduction mainly on the statistical significance of an effect, where Murre and Dros compare the shape of a descriptive curve, point by point, with no central hypothesis test. A measurement from 1880 obtained by one man on himself therefore does not sit the same exam as the hundred studies from 2008. It sits a different one, and it passes.

Still, Murre and Dros credit Ebbinghaus, in their closing remark, with a list of practices psychology took a century to make standard: controlled stimulus material, counterbalancing for time of day effects, protection against optional stopping, statistical analysis of the data and modeling. The counterbalancing deserves a qualification. Ebbinghaus did test his four longest intervals at three times of day, but for the short intervals he corrected after the fact, with percentages estimated on other series. Heller and his coauthors, for their part, were able to fix their schedule, found no time of day effect and applied no correction.

That leaves the question of why the curve stayed a decoration. A replication had existed since 1991, the one by Heller, Mack and Seitz. It was in German, with no English abstract, absent from the web at the time Murre and Dros were writing, and never cited in an international English language journal. The evidence was available and invisible. Ebbinghaus himself urged caution: this statement and the formula it rests on, he writes, have no value beyond that of a summary of results obtained once and under the circumstances described, and he cannot say, at this stage, whether other circumstances or other individuals would give them different constants. It took 70 hours of syllables with neither head nor tail, a Dutch winter and one hundred and thirty one years to begin answering him. Three more subjects do not make a population, but they do better than a textbook legend.

A project like this one?

I design and deploy products like this. Let's talk.

Let's talk