Monday, July 20, 2026

The Groundless Memory Research Boasts of the Late Susumu Tonegawa

Today's neuroscience is guilty of promoting many a groundless triumphal legend. One of those groundless socially-constructed triumphal legends is the clam that researcher Susumu Tonegawa did something to show a physical basis for memory. Tonegawa recently died, and the Transmitter magazine has a worshipful article repeating some of these groundless legends. 

The article starts out by quoting false boasts about the very low-quality 2015 paper "Engram cells retain memory under retrograde amnesia" co-authored by Tonegawa. The boasts are made by a co-author of the paper. When we look at the end of the supplemental material, and look at figure s13, we find that the experimenters were using a number of mice that was equal to only 8 in one study group, and 7 in another study group.  Such a paltry sample size does not result in any decent statistical power, and we should have no confidence in any paper using such way-too-small study groups sizes.  The paper failed to use a blinding protocol, an essential for a paper like this to be taken seriously. The paper had a thorough reliance on an utterly reliable technique for trying to judge recall in rodents: the worthless method of trying to judge "freezing behavior." All research papers relying on that method are examples of junk science, for reasons I thoroughly explain in my post here

neuroscience false alarms

Google Gemini infographic (pardon its spelling errors)

Next the Transmitter article makes this untrue claim: "In a series of papers in the 2010s, Tonegawa and his team showed that simply activating a subset of cells via optogenetics could reactivate a memory, change its valence and even create a false memory." The claim has no basis in fact. 

Let's take a look at some of the schlock work that Tonegawa and his collaborators produced on this topic, all of which is very low-quality research work utterly unworthy of praise:

  • The reference in the quote above to reactivating a memory is a reference to the very low-quality 2012 paper "Optogenetic stimulation of a hippocampal engram activates fear memory recall." Figure 2 tells us that in one of the groups of mice there were only 5 mice, and that in another group there were only 3 mice. Figure 3 tells us that in two other groups of mice there were only 12 mice. Figure 4 tells us that in some other group there was only 5 mice. Such  paltry sample sizes does not result in any decent statistical power, and the results are no good evidence of anything, because the study group sizes are way-too-small for any reliable result to be claimed.  The paper has a very big reliance on the use of "freezing behavior" judgments, not a reliable method for measuring fear or recall in rodents. No convincing evidence has been provided of artificially activating a fear memory by the use of optogenetics.
  • The reference in the quote above to creating a false memory is a reference to the very low-quality science 2013 paper "Creating a False Memory in the Hippocampus" co-authored by Tonegawa, which you can read hereWhen we look at Figure 2 and Figure 3, we see that the sample sizes used were paltry: the different groups of mice had only about 8 or 9 mice per group. Such a way-too-small sample size does not result in any decent statistical power, so the results have no weight, because the study group size is way too small for any reliable result to be honestly claimed.  The paper had a thorough reliance on an utterly reliable technique for trying to judge recall in rodents: the worthless method of trying to judge "freezing behavior." The paper makes no use of a blinding protocol, an essential for a paper like this to provide serious evidence of an effect. No convincing evidence has been provided of creating a false memory.
  • The claim in the quote above about switching a valence is a reference to the 2015 low-quality science study "Bidirectional switch of the valence associated with a hippocampal contextual memory engram" co-authored by Tonegawa.  We see in that paper 5 or 6 results reported with a borderline statistical significance of only "< 0.05," so this paper is  guilty of p-hacking. No detailed description is given of how an effective blinding protocol was achieved, and only the skimpiest mention is made of blinding, so this paper fails to convince us effective blinding was achieve.  The study used only "freezing behavior" to try to measure fear or recall, without corroborating such a thing by measuring heart rates.  So the paper has done nothing to reliably measure fear or recall in the mice it studies.  The study involved stimulating certain cells in the brains of mice, with something called optogenetic stimulation. The authors have assumed that when mice "freeze" after stimulation, that this is a sign that they are recalling some fear memory stored in the part of the brain being stimulated. What the authors neglect to tell us is that stimulation of quite a few regions of a rodent brain will produce freezing behavior. So there is actually no reason for assuming that a fear memory is being recalled when the stimulation occurs. Some of the study group sizes are good, but others are too-small. Because of all of these problems, no reliable evidence has been produced of a brain storage of memory in mice. 
  • Tonegawa co-authored the 2016 paper "Memory retrieval by activating engram cells in mouse models of early Alzheimer’s disease."  This very low-quality paper states that “No statistical methods were used to predetermine sample size.” That means the authors did not do what they should have done to make sure their sample size was large enough. When we look at page 8 of the paper, we find that the sample sizes used were merely 8 mice in one group and 9 mice in another group. On page 2 we hear about a group with only 4 mice per group, and on page 4 we hear about a group with only 4 mice per group. Such a paltry sample size does not result in any decent statistical power, and we should have no confidence in any paper using such way-too-small study groups sizes. The paper had a thorough reliance on an utterly reliable technique for trying to judge recall in rodents: the worthless method of trying to judge "freezing behavior." The study therefore provides no convincing evidence for any of its main claims. 
A close look at the memory research work of Susumu Tonegawa  will fail to find any studies that provide any good evidence for the grander boasts made in the titles of the papers. Although he apparently did some good work in another field (immunology), Susumu Tonegawa was a crappy experimenter in the field of cognitive neuroscience.  The praising quotes about him in the Transmitter article are mostly from scientists often guilty of the same type of bungling and poor experimental design that Tonegawa was so often guilty of. 

The Transmitter article also links to these very low-quality papers co-authored by Tonegawa:
  • A Tonegawa study claimed to have “identified engram cells” in the prefrontal cortex. It was a study entitled “Engrams and circuits crucial for systems consolidation of a memory.”  In Figure 1 (containing multiple graphs), we learn that the number of animals used in different study groups or experimental activities were 10, 10, 8, 10, 10, 12, 8, and 8, for an average of 9.5. In Figure 3 (also containing multiple subgraphs), we have even smaller numbers. The numbers of animals mentioned in that figure are 4, 4, 5, 5, 5, 10, 8, 5, 6, 5 and 5. None of these numbers are anything like what would be needed for a moderately convincing result, which would be a minimum of 15 or 20 animals per study group. No detailed description is given of an effective blinding protocol.  The study relies on judgments of freezing behavior of rodents, which is not a reliable way of measuring fear or recall in rodents. 
  • Tonegawa co-authored a study "Brain-wide mapping reveals that engrams for a single memory are distributed across multiple brain regions." It is a very low-quality piece of work using way-too-small sizes such as only 7 mice and only 9 mice. The study relies on judgments of freezing behavior of rodents, which is not a reliable way of measuring fear or recall in rodents. 
Tonegawa got a Nobel Prize not for any memory research, but for work in an entirely different area: immunology. 

Tonegawa founded a memory research lab at MIT. Every time I study the rodent memory research results of that lab I find research as low-quality as Tonegawa's memory research papers. For a look at some bogus boasts and very bad experimental methods employed by researchers at that lab, see my post here

I might try to put myself in the shoes of the person writing the gushing Transmitter article, and ask: what is going through the mind of such a person when you get that person falling "hook, line and sinker" for such junk science studies? Trying to imagine the person's train of thought, I can imagine the person thinking something like this:

The studies were published in major journals such as Nature and Cell. And the researchers worked at big prestigious universities such as MIT. And the papers were peer-reviewed. So the claims of the studies were probably true. If the authors had made untrue claims, the peer reviewers would have prevented publication. 

But the actual situation is this:
  • The church-like neuroscientist belief community (in which professors serve like priests) is a community that has long been addicted to very low quality methods of neuroscience research, largely so that it can maintain the illusion that its cherished but easily debunked belief dogmas are true.  
  • Junk cognitive neuroscience research is more the rule than the exception these days, even at laboratories of major universities such as Harvard and MIT. 
  • Leading journals such as Nature and Cell are routinely publishing very low-quality research in cognitive neuroscience. Such journals do not have published research standards guaranteeing high-quality research work in neuroscience. 
  • Peer-reviewers of submitted papers in cognitive neuroscience tend to be other researchers following research practices as bad as the methods of papers they are asked to review. Such peer-reviewers don't like to reject papers for being guilty of the same research methodology sins that the peer reviewers themselves are guilty of. So peer review does very little to prevent the publication of junk low-quality studies making false claims in their titles and abstracts. 
I can give an analogy for the type of memory research Tonegawa typically did. Imagine that you are testing whether people can get more "heads" flips than "tails" flips when flipping coins, if the people try to use psychokinetic "mind over matter" power to cause a "heads" flip. Imagine if you tested several small groups, each consisting of only 8 or 9 people.  There would be a good chance that one of the groups would report a better-than-50% number of "heads" flips. But that would be mere chance at work, and the effect would disappear if you used larger test groups such as 40 subjects per group. Imagine you also did not verify what each coin flip was, but relied on self-reports by the coin flippers of whether a coin on the ground viewed 10 meters away had landed "heads" or "tails." That would not be a reliable way of observing what the coin flips were, because some of the coin flippers might lie, or might be more prone to say that some not-very-clearly-seen coin had landed "heads."  Such an experiment would be  similar to a typical Tonegawa experiment. He typically used way-too-small study groups much smaller than 15, and reported claimed effects that chance could very easily have produced, effects that would disappear if a larger study group (such as 30 subjects) had been used. And the "freezing behavior" method he typically used to try measure fear or recall in rodents was as unreliable as self-reports from coin flippers viewing a coin from ten meters away. 

Today's cognitive neuroscience research landscape is a swampland of junk research, groundless legends, sleazy shortcuts, poor study design, bad methods and irreproducible results. When people who did frequently bungling memory-related research as bad as Tonegawa's are lionized as "giants," it helps show how deceptive a hall-of-mirrors echo chamber legend machine the world of today's neuroscience literature is. Bogus boasts and unjustified lionization are the enemies of the quest to establish scientific truth. Science goes astray when people put on pedestals scientists who made false boasts of doing things they did not do. 

scientist lionization

Appendix: A Short Look at the Folly of "Freezing Behavior" Estimations

When "freezing behavior" estimations go on, things typically work like this. A rodent will be trained to fear some thing such as a shock plate that gives the rodent a shock when the rodent steps on it. Then later the rodent will be placed in a cage that includes the fear stimulus such as the shock plate. The researchers will attempt to record what percentage of some time (say, a minute or 3 minutes) that the rodent was immobile when placed in such a case. This will be called a "freezing percentage," and will be claimed as a measure of how well the rodent remembered the fear stimulus. 

The technique makes no sense. In the real world, rodents don't usually freeze and become immobile when they are afraid. They are much more likely to flee. I know that from years of observing how mice act in the presence of shrieking humans, in an apartment where mice would occasionally appear. So trying to judge recall of a fearful stimulus by judging how much time a rodent was immobile in a cage makes no sense as a way of measuring fear or recall. The thing that utterly destroys the credibility of all "freezing behavior" graphs is that they can be produced in any of more than a dozen ways. A researcher can put a rodent in the cage for three minutes and graph the whole three minutes. Or he can graph only the first 30 seconds, or only the first minute, or only the first two minutes. In each ten seconds of such a three minutes, the researcher can count it as "moving" if the rodent moves one second during that period; or the researcher can count two seconds of movement as being mobility; or the researcher can use three seconds, or four seconds, or five seconds. 

There are no prevailing standards for how "freezing behavior" is judged. With there being a dozen different possibilities of how "freezing behavior" can be judged and graphed, with each having a possibility of success of about 50% (and with pre-registration -- a commitment to an exact methodology before gathering data -- being rare in neuroscience, as a neuroscientist recently confessed), it will be almost certain that the researcher will be able to choose some analysis method that will show the desired difference in "freezing behavior," even if the memory intervention being tested had no real effect. This is a large part of the reason why "freezing behavior" judgments are worthless as evidence for an increase or decrease in memory in rodents. 


A rodent may sometimes "freeze" or become immobile when seeing some fearful stimulus. But there was never any sound basis for assuming that you could reliably measure how well a rodent recalled something by measuring (over a timespan such as a minute) how immobile a rodent was in a cage that contained a fearful stimulus. No one ever did a study with a large sample size establishing the truth of such an assumption. 

Below is a depiction of a reliable method for measuring recall in rodents. 


Using this technique, a mouse is trained to avoid a fear stimulus -- the red shock plate shown in the center of the diagram. At some later date the mouse (in a hungry state) is put into the cage. If the mouse does not remember that the shock plate will cause pain, the mouse will take the direct route to the cheese, which requires crossing over the shock plate. If the mouse does remember that the shock plate will cause pain, the mouse will take an indirect and harder route, requiring it to jump up and down a set of stairs.  This is an easy and foolproof method of testing memory recall in rodents. Here we have a nice binary result -- either the mouse touches the shock plate, or it doesn't. There's no subjective element at all. 

You could use this fear stimulus avoidance technique with a setup even simpler than the one above, a setup with no stairs. You simply put a hungry mouse in a special cage with only one route to the cheese, a route that requires walking over the shock plate. If the mouse avoids the cheese, and fails to touch the shock plate, that would count as remembering that the shock plate will shock; but any touching of the shock plate would count as forgetting that the shock plate will shock. 


Instead of using good reliable objective methods such as the ones above,  today's cognitive neuroscience researchers tend to use the utterly unreliable and subjective method of trying to judge recall by trying to judge how immobile a mouse was in a cage during some arbitrary time interval such as one minute or three minutes.  This method is preferred because it is a "see whatever you want to see" method that maximizes the chance that some researcher will be able to claim some desired effect supposedly involving an increase or decrease in recall.  If peer reviewers were doing their job well, they would reject all papers using the worthless "freezing behavior" method of trying to judge recall. 

No comments:

Post a Comment