Showing posts with label low statistical power in neuroscience experiments. Show all posts
Showing posts with label low statistical power in neuroscience experiments. Show all posts

Sunday, September 28, 2025

Irredeemable: Reproducibility and Power Size in Neuroscience Are Very Bad, and Not Getting Any Better

 A recent study offers some encouraging news about psychology research. The paper is entitled "Increasing Sample Sizes in Psychology Over Time." The paper reports this:

"We collected data from 3176 studies across six journals over three years. Results show a significant increase in sample sizes over time (b=44.83, t(6.25)=4.48, p=.004, 95%CI[25.23,64.43]), with median sample sizes being 40 in 1995, 56.5 in 2006, and 122.5 in 2019. This growth appears to be a response to the credibility crisis....The increase in sample sizes is a promising development for the replicability and credibility of psychological science."

The credibility crisis referred to is the widely reported reproducibility crisis in fields such as psychology and neuroscience.  For decades it has been reported that experimental studies in psychology and neuroscience tend to be unreliable and poorly reproducible, largely because the sample sizes used were way too small.  This was commonly called a "reproducibility crisis in psychology," although it was very much a reproducibility crisis in both psychology and neuroscience. A tendency to produce studies with too-small sample sizes was just as prevalent in neuroscience as psychology. 

Psychology experiments typically involve humans, and advances in internet technology may have been a factor helping to lead to increased study group sizes in psychology. Decades ago a scientist might have found it necessary to recruit subjects to come into some laboratory where an experiment can be done. But now there are online platforms that allow people to sign up to be subjects in psychology experiments, while being paid for their efforts. This provides a very large pool of potential test subjects. A psychologist can now run experiments using subjects from across the USA or even multiple countries, by designing some experiment that subjects can participate in over the internet, while the subjects stay in the comfort of their homes.  

But while there may have been an increase in study group sizes used in psychology experiments, there has apparently been no such increase in the field of neuroscience. How could you honestly describe the state of experimental neuroscience? You might honestly describe it as an irredeemable cesspool consisting mostly of junk science studies that continue to have the same old fatal defects such as the use of way-too-small study group sizes. Well-designed studies in cognitive neuroscience seem to be in the minority, and are outnumbered by junk science studies guilty of very bad Questionable Research Practices. 

Scientific studies that use small sample sizes are typically unreliable, and often present false alarms, suggesting a causal relation when there is none. Such small sample sizes are particularly common in neuroscience studies, which often require expensive brain scans, not the type of thing that can be inexpensively done with many subjects. In 2013 the leading science journal Nature published a paper entitled "Power failure: why small sample size undermines the reliability of neuroscience." There is something called statistical power that is related to the chance of a study producing a false alarm. The Nature paper found that the statistical power of the average neuroscience study is between 8% and 31%. With such a low statistical power, false alarms and false causal suggestions will be very common. 

A scientific study with a statistical power of 50% is one that will have about a 50% chance of being successfully reproduced when someone attempts to reproduce it. Even when a statistical power of 50% is reached, the statistical power is not high enough for robust evidence to be claimed.  In order to be robust evidence for an effect, a study much reach a higher statistical power such as 80%.  When that power is reached, there is about an 80% chance that an attempt to reproduce the results will be successful. 

The Nature paper said, "It is possible that false positives heavily contaminate the neuroscience literature." 

An article on this important Nature paper states the following:

"The group discovered that neuroscience as a field is tremendously underpowered, meaning that most experiments are too small to be likely to find the subtle effects being looked for and the effects that are found are far more likely to be false positives than previously thought. It is likely that many theories that were previously thought to be robust might be far weaker than previously imagined."

Scientific American reported on the paper with a headline of "New Study: Neuroscience Gets an 'F' for Reliability."

So, for example, when some neuroscience paper suggests that some part of your brain controls or mediates some mental activity, there is a large chance that may simply be a false positive. As this paper makes clear, the more comparisons a study makes, the larger a chance for a false positive. The study has an example: if you test whether jelly beans cause acne, you'll probably get a negative result, but if your sample size is small, and you test 30 different colors of jelly bean, you'll probably be able to say something like "there's a possible link between green jelly beans and acne"  -- simply because the more types of comparisons, the larger the chance of a false positive.  So when a neuroscientist tries to look for some part of your brain that causes some mental activity, and makes 30 different comparisons using 30  different brain regions, with a small sample size, he'll probably come up with some link he can report as "such and such a region of the brain is related to this activity." But there will be a high chance this is simply a false positive.  

bad neuroscience lab

The 2013 "Power Failure" paper discussed above was widely discussed in the neuroscience field, but a 2017 paper indicated that little or nothing had been done to fix the problem. Referring to an issue of the Nature Neuroscience journal, the author states, "Here I reproduce the statements regarding sample size from all 15 papers published in the August 2016 issue, and find that all of them except one essentially confess they are probably statistically underpowered," which is what happens when too small a sample size is used. 

A 2017 study entitled "Effect size and statistical power in the rodent fear conditioning literature -- A systematic review" looked at what percentage of 410 experiments used the standard of 15 animals per study group (needed for a moderately compelling statistical power of 50 percent).  The study found that only 12 percent of the experiments met such a standard.  What this basically means is that 88 percent of the experiments had low statistical power, and are not compelling evidence for anything.


low statistical power in neuroscience


The 2017 scientific paper "Empirical assessment of published effect sizes and power in the recent cognitive neuroscience and psychology literature" contains some analysis and graphs suggesting that neuroscience is less reliable than psychology. Below is a quote from the paper:


"With specific respect to functional magnetic resonance imaging (fMRI), a recent analysis of 1,484 resting state fMRI data sets have shown empirically that the most popular statistical analysis methods for group analysis are inadequate and may generate up to 70% false positive results in null data. This result alone questions the published outcomes and interpretations of thousands of fMRI papers. Similar conclusions have been reached by the analysis of the outcome of an open international tractography challenge, which found that diffusion-weighted magnetic resonance imaging reconstructions of white matter pathways are dominated by false positive outcomes  Hence, provided that here we conclude that FRP [false report probability] is very high even when only considering low power and a general bias parameter (i.e., assuming that the statistical procedures used were computationally optimal and correct), FRP is actually likely to be even higher in cognitive neuroscience than our formal analyses suggest.

The paper draws a shocking conclusion that most published neuroscience results are false. The paper states the following: "In all, the combination of low power, selective reporting, and other biases and errors that have been well documented suggest that high FRP [false report probability] can be expected in cognitive neuroscience and psychology. For example, if we consider the recent estimate of 13:1 H0:H1 odds, then FRP [false report probability] exceeds 50% even in the absence of bias." The paper says of the neuroscience literature, "False report probability is likely to exceed 50% for the whole literature." 

In June of 2025 I searched on Google Scholar, trying to find some paper reporting on an improvement of sample sizes in neuroscience research. I could find no such paper. The sample sizes used in neuroscience research are very bad, and are not getting any better. Today's neuroscience research is a cesspool of dysfunction and misleading claims. There are no signs that it is improving its horribly dysfunctional ways. 

Why does this situation persist? There are two main reasons: economics and ideology. 

The economic explanation for bad science practices is explained rather well in the paper "The Natural Selection of Bad Science" by Paul E. Smaldino and Richard McElreath. In that paper we read this:

"Poor research design and data analysis encourage false-positive findings. Such poor methods persist despite perennial calls for improvement, suggesting that they result from something more than just misunderstanding. The persistence of poor methods results partly from incentives that favour them, leading to the natural selection of bad science. This dynamic requires no conscious strategizing—no deliberate cheating nor loafing—by scientists, only that publication is a principal factor for career advancement. Some normative methods of analysis have almost certainly been selected to further publication instead of discovery....We first present a 60-year meta-analysis of statistical power in the behavioural sciences and show that power has not improved despite repeated demonstrations of the necessity of increasing power. To demonstrate the logical consequences of structural incentives, we then present a dynamic model of scientific communities in which competing laboratories investigate novel or previously published hypotheses using culturally transmitted research methods. As in the real world, successful labs produce more ‘progeny,’ such that their methods are more often copied and their students are more likely to start labs of their own. Selection for high output leads to poorer methods and increasingly high false discovery rates."

The paper has a shocking confession by a scientist who has worked on search committees searching for scientists to be hired. The scientist states this:

"I’ve been on a number of search committees. I don’t remember anybody looking at anybody’s papers. Number and IF [impact factor] of pubs are what counts."

This is a description of an economic ecosystem in which what  determines a scientist's career advancement is not the quality and reliability of the papers he has published, but the mere quantity of such papers, and how many citations such papers are getting. 

The paper ("The natural selection of bad science") states this: "In fields such as psychology, neuroscience and medicine, practices that increase false discoveries remain not only common, but normative." In this context "normative" means "more the rule than the exception." The paper states, "Some of the most powerful incentives in contemporary science actively encourage, reward and propagate poor research methods and abuse of statistical procedures." Later the paper gives us some insight on the economics that help to increase the likelihood of scientists producing lots of low-quality research papers:

"If researchers are rewarded for publications and positive results are generally both easier to publish and more prestigious than negative results, then researchers who can obtain more positive results—whatever their truth value—will have an advantage. ...One way to better ensure that a positive result corresponds to a true effect is to make sure one’s hypotheses have firm theoretical grounding and that one’s experimental design is sufficiently well powered. However, this route takes effort and is likely to slow down the rate of production. An alternative way to obtain positive results is to employ techniques, purposefully or not, that drive up the rate of false positives. Such methods have the dual advantage of generating output at higher rates than more rigorous work, while simultaneously being more likely to generate publishable results. Although sometimes replication efforts can reveal poorly designed studies and irreproducible results, this is more the exception than the rule. For example, it has been estimated that less than 1% of all psychological research is ever replicated  and failed replications are often disputed. Moreover, even firmly discredited research is often cited by scholars unaware of the discreditation. Thus, once a false discovery is published, it can permanently contribute to the metrics used to assess the researchers who produced it....Campbell’s Law, stated in this paper’s epigraph, implies that if researchers are incentivized to increase the number of papers published, they will modify their methods to produce the largest possible number of publishable results rather than the most rigorous investigations."

What the paper is suggesting is that junk science is strongly incentivized in today's science research ecosystem.  A scientist is more likely to succeed in academia if he produces a high quantity of low-quality research papers than if he produces a lower quality of high-quality research. There are several online sources that keep track of the number of papers that a scientist wrote or co-wrote, and the number of citations such papers got.  There are no online sources that keep track of the quality and reliability of the papers that such a scientist produced.  In such an environment, a scientist will be more likely to get ahead if he produces many low-quality papers rather than a smaller number of papers that are more reliable and truthful in the results reported. 

junk science practices

The economic motivations of badly behaving neuroscientists and similar bad actors are sketched in my diagram below, and the post here explaining the diagram. At the top left corner is the starting point of "quick and dirty" experimental designs with way too few subjects. The diagram charts how various types of people in various industries benefit from such malpractice. 

academia cyberspace profit complex

Another huge explanatory factor that helps explain the massive persistence of junk neuroscience studies is ideology. What we should never forget is that neuroscientists are members of a belief community.  That belief community is dedicated to promoting various dubious belief dogmas such as the dogma that the brain is the source of the human mind, and the dogma that the brain is the storage place of human memories. So in many cases junk science studies that a peer reviewer or an editor would normally be ashamed to approve for publication will be approved for publication, because the study appears to support some dogma or narrative that is cherished by members of the neuroscientist belief community. 

church of academia

The main beliefs of the neuroscientist belief community are false beliefs.  Because of innumerable reasons discussed on this blog, there is no credibility in the claim that the brain is the source of the human mind, and there is no credibility in the claim that the brain is a storage place of human memories. When the beliefs of a belief community are true, the community does not need to rely on studies involving bad science practices or bad scholarly practices.  But when the beliefs of a belief community are false, that belief community may need to keep producing studies involving bad science practices or bad scholarly practices. That way the belief community can try to maintain an illusion that the evidence is favoring its cherished beliefs. 

Monday, May 5, 2025

The STAT Research Award Goes to Junk Neuroscience

Statnews.com is a site that tries to create an aura of a serious, respectable science news site. It bills itself as "your go-to source for the world of life sciences, medicine, and biopharma." My guess is that the site gets funding from pharmaceutical companies and medical device manufacturers, and that it exists largely to serve their interests. On the site's pages we don't get the usual swarm of ads that you see these days on so-called science news site; but there are some ads. You should always be suspicious of any science news site containing ads.  Every time you see an ad on the pages of Statnews.com, you should remember that online science news sites containing ads will tend to have clickbait headlines that drive people to click on headlines, so that they go see ads that make the site owners money. 

At a recent article at statnews.com we have an example of someone making a misleading statement. We have a biologist who boasts "I’ve published more than 340 papers that have garnered more than 100,000 citations."  But the biologist was merely a co-author of most of such papers, which typically had 5, 10 or as many as 39 authors each. It is not right to create the impression that you did by yourself some writing that actually required the work of hundreds of different authors in addition to yourself. 

Recently at the site we saw this headline: "Baylor crowned STAT Madness champion for second consecutive year, for insights into memory." This is followed by the utterly untrue claim that "researchers showed brain cells called astrocytes are involved in forming and recalling memories." 

We read of something called the 2025 STAT Madness competition, "a bracket style celebration of neuroscience research." We read of some voting process, although we get no details of how this worked.  We read the claim that the competition "stacked 64 entries against each other in a month long combination."  Sadly the winner of this competition was a very low-quality piece of junk science entitled "Learning-associated astrocyte ensembles regulate memory recall." You can read the abstract here

The paper is behind a paywall. Why has the Statnews.com site given an award for a study that the public cannot even access? But without  paying any money, I can use the abstract link to get all the information I need to determine that the study is junk science.  For the abstract link does allow me to look at some of the paper's figures  And those figures tell me enough to determine the junk science nature of the paper. 

Figure 1 is below, and it has two "freezing percentage" bar graphs at its bottom.


Figure 4 and Figure 5 look similar. Figure 4 (entitled "Reactivation of LAAs elicits memory recall") has three "freezing percentage" bar graphs.  Figure 5 (entitled "Ensemble-specific NFIA is necessary for context-specific memory") has two "freezing percentage" bar graphs. 

From these figures, I can tell that the paper is guilty of two of the methodological sins so very common in memory-related rodent research, either one of which is enough to show the paper is very low-quality research.

Fault 1: Judgments of Freezing Behavior Were Used to Try to Judge Animal Recall

 We can tell from the figures mentioned above that the researchers attempted to measure recall in a rodent by using a "freezing behavior" estimation. All rodent experiments that use such a method are examples of junk science. The technique involves putting a rodent in a cage in which there is some stimulus the animal was trained to fear, and then trying to judge fear or recall by trying to judge what fraction of a time interval the rodent was immobile, with the immobility being judged as the rodent "frozen in fear." Neither fear nor memory nor recall can reliably be measured by such a technique. How often a rodent is immobile in a cage is a random thing that does not reliably correlate with how much the rodent is afraid or how often the rodent recalls a fearful stimulus.  For a full discussion of the utter unreliability of "freezing behavior" judgments in neuroscience experiments, see my long post here, entitled "All Papers Relying on Rodent 'Freezing Behavior' Estimations Are Junk Science." 

freezing behavior in rodents

I lived in a large apartment building for more than a decade, and about once a month I would spot a mouse, which would usually cause me to shriek. It seemed that not one time did I ever see a mouse freezing in fear after getting this fearful stimulus, contrary to the assumptions of neuroscientists assuming that mice tend to freeze in immobility when they are afraid. It would seem to always be the opposite behavior: the mice would always flee. Tracking the immobility of a mouse in a cage (and assuming the more "freezing" the more fear) is not a reliable way of measuring fear or recall in a mouse. 

Fault 2:Way-Too-Small Study Group Sizes

The use of way-too-small study group sizes is more the rule than the exception in today's neuroscience rodent research. 

The paper "Prevalence of Mixed-methods Sampling Designs in Social Science Research" has a Table 2 giving recommendations for minimum study group sizes for different types of research. According to the paper, the minimum number of subjects for an experimental study are 21 subjects per study group. The same table lists 61 subjects per study group as a minimum for a "correlational" study. The table below is a shortened version of the Table 2 found in that study. Because of the small effect sizes and "levels of significance" typically reported in neuroscience experiments, we should regard the sample sizes required for neuroscience experiments as every bit as high as the numbers listed below. 

Research design/method

Minimum sample size suggestion

Correlational

64 participants for one-tailed hypotheses; 82 participants for two-tailed hypotheses (Onwuegbuzie et al ., 2004)


Causal-comparative

51 participants per group for one-tailed hypotheses;

64 participants for two-tailed hypotheses (Onwuegbuzie et al., 2004)


Experimental

21 participants per group for one-tailed hypotheses

(Onwuegbuzie et al., 2004)


Phenomenological

<= 5/10 interviews (Creswell, 1998); >= 6 (Morse, 1994


For correlational, causal-comparative and experimental research designs, the recommended sample sizes represent those needed to detect a medium (using Cohen’s [1988] criteria), one-tailed statistically significant relationship or difference with 0.80 power at the 5% level of significance.”


Source: “Prevalence of Mixed-methods Sampling Designs in Social Science Research” by Kathleen M.T. Collins, Anthony J. Onwuegbuzie and Qun G. Jiao. 


In her post “Why Most Published Neuroscience Findings Are False,” Kelly Zalocusky PhD calculates that the median effect size of neuroscience studies is about .51. She then states the following, talking about statistical power (something that needs to be substantially greater than .5 for any compelling result to be claimed): 

"To get a power of 0.2, with an effect size of 0.51, the sample size needs to be 12 per group. This fits well with my intuition of sample sizes in (behavioral) neuroscience, and might actually be a little generous. To bump our power up to 0.5, we would need an n of 31 per group. A power of 0.8 would require 60 per group."

Here Zalocusky is telling us that the sample size requirements (i.e. study group size requirements) are just as high as those suggested in the table above, and that a study group size of around 50 is needed to get a good statistical power of 0.8 (80%). 

What study group sizes were used for the paper "Learning-associated astrocyte ensembles regulate memory recall," the paper that won the 2025 STAT Madness competition?  You can tell by the number of dots that we see at the bottom of the figures such as the Figure 1 at the top of this post (and also Figure 4  and Figure 5 of the paper). Each dot is a data point for a particular rodent. We see only about 8 dots for each of the study groups used. So the study group sizes were a way, way too-small size of only about 8 subjects per study group. 

freezing behavior bar chart

The study is therefore worthless as evidence, having used study group sizes less than half (and probably less than a third) of the study group sizes needed for a reliable result. StatNews.com (or the voters it coordinated) made the very bad mistake of giving a prize to a very low-quality junk science paper. 

How to Do Work Like This Junk Science Result

I can give an example of how you could do work similar to the utterly shoddy work that was awarded the STAT prize. You might do a project trying to show that some people can predict which cards are at the top of a newly shuffled deck that has been cut. 

A very important of your work would be to not publish beforehand any exact plan for how the research would be done.  Then you could do some tests that would work line this:

(1) You shuffle a deck of cards three times, and then cut the deck. 

(2) You bring in someone, asking him to guess the first ten cards that you will draw from the top of the deck.

You could have a "Correct Guesses" sheet recording the drawn result and the guess result for each of the guesses. Now, it would help very much if you used a technique that did not reliably record the guesses the guesser made. It might work like this: you could ask the user to list a sequence of ten guesses. You would then deal the first ten cards out on the table. You would then record your best recollection of the sequence of ten cards that the user had called out. 

This would not be a reliable measurement technique, because you had failed to write down each guess just after the guesser made it. So you would now be relying on your memory of the sequence of 10 cards the user chose. You could easily bias things by failing to remember correctly one or more of the cards the guesser had stated. This would help you get a more favorable "correct guesses" total. 

Now doing such a test on a few subjects, you would have some data: the sequence of ten cards drawn, and what the ten guesses made were (or at least, your best guess about what those guesses were, based on your short-term memory). You would then have a variety of ways to analyze your data. You could use the entire sequences of ten guesses.  Or if you did not like the "percentage correct" result from that, you could use only the first five cards guesses. Or if you did not like the "percentage correct" result from that, you could use only the last five cards guessed. Or if you did not like the "percentage correct" result from that, you could use only the middle five cards guessed. 

Or, if you were not getting anywhere trying to show above-chance card guessing, you could only report correct guesses of the card number or card face. Or, if you still were not getting anywhere trying to show correct card guessing, you could only report correct guesses of the card suit (clubs, spades, hearts or diamonds). Or, if you still were not getting anywhere trying to show correct card guessing, you could only report correct guesses of the card color (black or red). Or if the entire set of data showed no effect above chance for the card guessing, you could report on only an above-chance result for a single guesser. 

What happened if you tried all these things for 5 subjects, and were still unable to find anything greater than chance? No problem. You could simply "file drawer" the results, filing them away in your files. You could then start a new experiment using a different set of 5 subjects. After a few such attempts, you would probably have something you could report as an "above chance" result for card guessing, given your small study group sizes, and given your freedom to analyze the data in any of dozens of different ways. False alarms tend to show up with small data sets. Most of those false alarms go away when you use a much larger study group size. 

For example, if you flip a coin 100 times you will get a number of heads very close to 50. But if you flip a coin only a small number of times, it is very easy to get a result much different from the result expected by chance. So, for example, it is not too hard to flip a coin six times and get four or more heads or four or more tails rather than the chance-expected result of three heads. The chance of such a result is greater than 33%. Do three trials of six flips each, and you will probably get one trial with four heads or four tails. 

The example given above is very much like the typical neuroscience rodent experiment. Just like the card-guessing experiment I described, a typical neuroscience experiment fails to publish in advance any detailed research plan, leaving the experimenters free to analyze data in any of innumerable ways. For example, there are no standards for how to judge "freezing behavior" in rodents. So an experimenter can attempt to judge how much a rodent was immobile after being exposed to some fear stimulus it was trained to fear, recording three minutes of data. If the experimenter does not like the "freezing percentage" recorded for the first three minutes, the experimenter can simply report on only the "freezing percentage" recorded for the first two minutes. If the experimenter does not like the "freezing percentage" recorded for the first two minutes,  the experimenter can simply report on only the "freezing percentage" recorded for the first 60 seconds. If the experimenter does not like the "freezing percentage" recorded for the first 60 seconds, the experimenter can simply report on only the "freezing percentage" recorded for the first 30 seconds. Since there is no convention of a standard time interval to use, and since researchers typically get away without even reporting what was the time interval corresponding to a "freezing percentage" graph, it is easy for a researcher to switch around the time interval in any way he wants, and not even use the same time interval for each freezing percentage" graph in his paper.

And just as the card-guessing experimenter I described used an unreliable technique for measuring a key element of the experiment (relying on the experimenter's memory of what 10 called guesses were), nowadays neuroscience rodent memory experiments rely on an utterly unreliable measurement technique: the technique of trying to judge an animal's recall by judging an animal's immobility during an arbitrary span of time, with that immobility called "freezing behavior." There are reliable ways of measuring whether a rodent recalled something. You can use heart-rate measurement to measure heart rate spikes, and that is a reliable way of telling whether a rodent is recalling a fearful stimulus (the heart rate of rodents spikes dramatically when they are afraid). Or you can use something like the Morris water maze, widely regarded as a reliable way of measuring how well a rat remembered something. Or you can use the Fear Stimulus Avoidance technique described below. 

reliable neuroscience recall measurement

But if you use the unreliable "freezing behavior" technique, your paper deserves nothing but scorn. And the more "freezing behavior" graphs that your paper has, the more it deserves scorn and contempt. 

Nowadays neuroscience memory research involving rodents tells us mainly one important thing about memory: that neuroscience rodent researchers seem to have memories so bad that they keep forgetting to follow good standards in doing research. 

bungling neuroscientist

bungling neuroscientist

bungling neuroscientist

bungling neuroscientist

Postscript: The same fatal defects of the paper discussed above are found in the recent junk science paper "EPSILON: a method for pulse-chase labeling to probe synaptic AMPAR exocytosis during memory formation." The paper uses way too-small study group sizes of only 3 mice, 4 mice and 6 mice. The paper also hinges upon a use of the totally unreliable method of trying to judge "freezing behavior" to try and judge whether fear recall occurred in mice. The longer the time interval over which so-called freezing behavior is judged, the less reliable is the claim to have measured fear or recall, because real "freezing behavior" would be a very short-lasting phenomenon, probably lasting less than 30 seconds. In this case the "freezing behavior" is judged over a length of three minutes, a particularly dubious length of time to be using when such a judgment about "freezing behavior" is done. 

We again have an example of the arbitrary decisions and lack of standards that go on when such "freezing behavior" judgments are made, the kind of deal in which any experimenter is free to use any criteria, in a "see whatever you want to see" manner. The authors tell us that any motion of less than .03 meters per second (about less than one inch per second) was counted as "freezing behavior," stating, "the time duration during which the speed was slower than 0.03 m s−1 was counted as freezing time."  That makes no sense. Moving .8 inch a second is not actually freezing behavior for a mouse. Freezing behavior is immobility.  

The authors failed to do a sample size calculation to determine the number of mice needed to get a result of decent statistical power. Rather than confess their failure to do this basic requirement of good experimental science (a sample size calculation), the authors strangely state, "Sample sizes were determined by the technical requirements of the experiments." That sounds like someone saying the equivalent of, "We couldn't use more mice because it would have been too much trouble," which is very lame. The authors do candidly confess, "Data collection was not conducted blinded to experimental conditions." No paper of this type should be taken seriously unless the authors followed a detailed blinding protocol, and the authors did no such thing. 

Proving that you can't trust anything merely because it appears in a Harvard publication, the Harvard Gazette has published a recent article boasting about this piece of very bad junk science, one with the bogus headline, "Tracking precisely how learning, memories are formed." Nothing of the sort was done, because of the very bad methods followed. 

Tuesday, July 2, 2024

The Mythology of "Memory Maintenance Molecules"

 In an article at the Nautilus web site, scientist Ken Richardson suggests that his fellow scientists have been guilty of some molecular mythology. He points out that scientists have repeatedly used “action verbs” in describing DNA, telling us that inside DNA are genes that “act,” “behave,” “direct,” “control,” “design,” are “responsible for,” and so forth. But then Richardson tells us “a counter-narrative is building” to correct such erroneous ideas, and then gives us reasons for thinking that genes are merely passive chemical units that do no such things.

Another example of molecule mythology involves a protein called PKMzeta. Some neuroscientists have suggested that PKMzeta has the ability to make memories last for decades in synapses, even though the proteins that make up synapses are very short-lived (having an average lifetime of two weeks or less). Quite a few of the papers or posts spreading this idea were written or co-written by the same person, Todd C. Sacktor. It is never explained clearly by any such theorists how a protein molecule could perform this great feat of magic. For anyone to explain such a thing clearly, he would first need to have a clear theory of how conceptual memories and episodic memories could be stored in synapses. No neuroscientist has ever presented a clear and explicit theory of any such thing. Neuroscientists merely vaguely tell us that somehow memory storage in a brain occurs through “synapse strengthening,” without presenting any clear theory of how that could be memory storage. 

Of course, if you do not have a clear theory of how memories could be stored (for even a few minutes) in synapses, you cannot possibly have a clear theory as to how some protein molecule such as PKMzeta could possibly cause memories stored in synapses to persist for decades, even though the proteins that make up such synapses are very short-lived, lasting an average of less than two weeks. Trying to defend against the charge that synapses are totally unsuitable for storing memories for decades, because of the short lifetimes of the proteins that make up synapses, a scientific paper states, “As long as PKMZ [PKMzeta] remains active and there is an absence of forces which terminate its activity (such as LTD), it will continue to sustain the biochemical changes at the synapse which serve as the neurobiological basis of memory, allowing the memory to persist for durations far exceeding the turnover of its component molecules.” But how could such a miracle of persistence occur, which would be like a message written in wet sand at the seashore persisting for decades, even though the wet sand was being replaced and written over whenever the tide came in? The science paper does not tell us.

Again, we have the case of an “action verb” inappropriately used to describe a molecule. We are told that PKMzeta has a “sustain” super-power allowing it to preserve fantastically complicated information supposedly stored externally in synapses made up of short-lived molecules that are constantly being replaced. There is nothing in the structure of PKMzeta that should cause us to believe it can do any such thing. No theorist has presented an explicit theory as to how anything like PKMzeta could preserve a memory. Such theorists may sometimes present chemical details to impress us, but such details do not constitute a theory unless a theorist gives explicit examples of precisely how specific memories (such as someone's memory of seeing Paris or someone's memory of details learned about World War I) could be permanently stored with the aid of PKMzeta.  No theorist has done any such thing. 

I suppose that if a PKMzeta molecule were able to cause memories to persist despite rapid protein turnover,  we might imagine it as some kind of "genius" molecule that has thoughts like this:

Oh, my goodness, I see that a memory is starting to degrade because of protein turnover! The memory now states, "Ottawa is the capitol of," which isn't even a full English sentence. Why, I'd better synthesize some new proteins to fill in for those proteins that died,  so there can be a nice complete English sentence. Now, what was that country that Ottawa is the capital of?  

Of course, anything the slightest bit like this is very hard to believe in. It would seem that the most minimal requirements that a molecule would have to fulfill in order to be a "memory maintenance molecule" would be the following:

(1) The molecule would have to somehow know whenever a particular protein molecule (that was part of a memory stored in a synapse) had died or disappeared because of the short lifespans of protein molecules.
(2) The molecule would have to somehow cause a replacement protein of the same type to appear in the same place as the vanished molecule, so that the memory did not degrade. 

The problem is that no one can envision a credible scenario under which a molecule could have either of these powers. To imagine how much of a miracle it would be for memories to persist despite constant protein turnover,  you can imagine a homeowner with ten picnic tables in his backyard, each of which is filled with leaves on which a word or two is written. Imagine these leaves spell out narratives, factual information, and ideas. But the problem is that about one day in three there are winds blowing the leaves off of the tables, and scattering them far away. Also, the leaves don't last longer than a year, because they tend to crumble. Now imagine the homeowner has to keep all this information preserved in the leaves, not just for a few nights but for 50 years. That would be a mountainous job.  An equally mountainous job would have to be done if memories were to be preserved in brains despite constant protein turnover causing proteins to persist an average of less than two weeks, and no one has explained how a molecule could possibly do such a feat.  Since synapses face not only rapid protein turnover inside them but also the problem that synapses don't last for longer than a year or two,  they have the same "double degradation" problem that such a homeowner would have with his information written on leaves. 




In the article here, a PKMzeta enthusiast is asked to explain how PKMzeta could cause memories to persist. The scientist gives a lengthy answer which fails to explain how PKMzeta could do such a thing. He merely says "a cluster of PKMzeta molecules can keep themselves turned on perpetually," and then claims that this supposed ability "is a plausible mechanism for memory persistence," without justifying that claim. This fragmentary theorizing is just hand waving. It has never been demonstrated that any cluster of PKMzeta molecules is capable of storing any information (such as a list of words) for a period as long as a month.  We can imagine hypothetical lab experiments that might try to show such a thing, but they have never been done. The paper here refers to "900 synaptic proteins." PKMzeta is only one of those 900 proteins in synapses, being no more common in synapses than an average synapse protein. You don't solve the "short lifetime of proteins" problem by trying to argue that one in 900 of those proteins might somehow have some stability.  As for the scientist's use of the word "plausible," it has been noted by others that "plausible" is the most abused word in theoretical science discourse, and that scientists often carelessly use the word "plausible" without ever doing anything at all to show a likelihood. 

But the PKMzeta enthusiasts have done a few studies which they claim lends credibility to their claims. I will describe a typical such study. A small number of mice are injected with something that suppresses the PKMzeta molecules in their body (or perhaps they are genetically engineered so that they don't have any PKMzeta). Memory experiments are then done. It is sometimes claimed that such mice perform not as well as normal mice. Such experiments have been hailed as support for the “memory maintenance” claims about PKMzeta.

There are several reasons why such studies do not at all show the claims about PKMzeta are correct. The first is that a result such as I described could never show that PKMzeta can save memories from destruction for years. Whenever memory is tested, it's hard to figure out what the cause is for a discrepancy between two test groups. A difference in a test result might be because (1) PKMzeta is involved in perceiving whatever observation is being tested; (2) or that PKMzeta is involved in memory storage; (3) or that PKMzeta is involved in memory retrieval; (4) or that PKMzeta has something to do with attention or focus or visual perception used in a memory test. A test discrepancy could never tell us which of these things was involved. And if some mice did worse in remembering things without PKMzeta, that might justify the small claim that PKMzeta has something to do with memory, but could never justify the vastly more extravagant claim that PKMzeta is capable of preserving memories for decades.

Another reason why such studies do not at all show the claims about PKMzeta are correct has to do with a general malaise in neuroscience. A general problem in modern neuroscience is the production of papers with marginal results that we cannot trust because of things such as too-small sample sizes and publication bias. Let us imagine that neuroscientists want to prove some idea that fits in with their ideological expectations. A great number of experiments might be done, almost all producing no support for the idea. But perhaps 1 in 20 might produce results marginally supporting the idea, probably because of chance variations in data. Now, today negative results are vastly less likely to get published than positive results. So if 19 researchers get a negative result, conflicting with what neuroscientists hope to get, it could be that 10 of them don't even bother to write up their results as a scientific paper, and that the other 9 do write up a paper but don't get it published (because of the journal bias against negative results). However the one researcher who (by chance) got a positive result will write up his result as a scientific paper. Since it will be a result neuroscientists were hoping to get, he will almost certainly get the result published.

This publication bias is a great problem affecting the reliability of scientific research. Because of it we should follow a precautionary neuroscience rule like this: don't believe something has been established unless the result turns up fairly consistently at a high level of significance, in studies with large sample sizes.

Has this happened in regard to memory experiments involving PKMzeta? Not at all. In 2011 a scientist reported three separate studies showing that inhibiting PKMzeta has no effect on memory if tested between 10 and 15 day after the memory forms.  In 2013 two groups of scientists published results conflicting with claims that PKMzeta might allow memories to persist a long time. One study by a team of scientists used genetically engineered mice that had no PKMzeta. It found that such mice “have no deficits in several hippocampal-dependent learning and memory tasks,” and concluded that PKMzeta is not required for memory or learning. Another study by a different team of scientists found that absence of PKMzeta “does not impair learning and memory in mice.” A 2015 study found that inhibiting PKMzeta has no effect on memory in tests performed 30 days after the memory forms. A 2016 paper also found that that inhibiting PKMzeta has no effect on memory in tests performed 30 days after the memory forms.

Such studies would seem to completely debunk claims that PKMzeta enables memories to persist for decades in synapses despite the short lifetimes in the proteins.

The scientists such as Sacktor who helped to spread the PKMzeta myth have tried to fight back with papers such as this 2016 paper. But in that very paper we see evidence that second-rate science is being used to try to prop up claims about PKMzeta. In Figure 7 the scientists tell us how many mice were used for their experiment involving the memory effects of PKMzeta deprivation. They used only 8 mice per study group. That's way too small a sample size to get a moderately convincing result. It is well known that at least 15 animals per study group should be used to get a moderately convincing result. If you use only 8 animals per study group, there's a very high chance you'll get a false alarm, in which the result is due merely to chance variations rather than a real effect in nature.  In fact, in her post "Why Most Published Neuroscience Studies Are False," neuroscientist Kelly Zalocusky suggests that neuroscientists really should be using 31 animals per study group to get a not-very-strong statistical power of .5, and 60 animals per study group to get a fairly strong statistical power of .8.  Compare these numbers to the 8 animals per study group mentioned in Figure 7 of the Sacktor paper. 

This is the same “too small sample size” problem (discussed here) that plagues very many or most neuroscience experiments involving animals. Neuroscientists have known about this problem for many years, but year after year they continue in their errant ways, foisting upon the public too-small-sample-size studies with low statistical power that don't prove anything because of a high chance of false alarms.

If you look up the PRKCZ gene behind the PKMZeta protein molecule, using this page and this page of the Human Protein Database, you will find no characteristics that seem unusual, and nothing suggesting any superstar status. The pages make no mention of the gene even being used in synapses, telling us that the gene is "mainly localized to the cytosol" and "in addition localized to the plasma membrane."   The very idea of some kind of "superstar protein" or "superstar gene" is contrary to the experience in recent decades of scientists, who have found in general that bodily functions almost always involve the coordinated ballet of very many different genes (typically hundreds of them to accomplish a particular task). 

The 2015 scientific paper here shows that PKMzeta rapidly degrades in synapses. The authors say that therefore a stable amount of PKMzeta "would be difficult to maintain at synapses and store memories over long time scales." The paper tells us “There is growing evidence against a role for PKMzeta in memory.” Figure 9 of the paper also shows that a kind of cousin molecule or "isoform" of PKMzeta (PKC lambda) also quickly degrades, experiencing a 50% loss or degradation every 10 hours. So it seems that there is no truth to the idea of PKMzeta (or PKC lambda) as some magic bullet that allows memories to persist for decades in synapses that are constantly having their proteins replaced.

Where does that leave neuroscientists? It leaves them without a leg to stand on in their claims that memories are stored in brains. Based on everything we know about synapses, there is no reason to believe that synapses are capable of storing a memory for even a month, let alone the 50 years that is how long older humans can remember things. As discussed here and here, equally grave problems prevent scientists from creating any credible account of how memories could be encoded into neural states or how seldom-retrieved facts learned many years ago could be instantaneously recalled from a brain that seems to lack any capability for fast look-ups from exact neural positions. We also know (as discussed here and here) that massive damage can occur to brains (such as surgical removal of half of a brain) while producing little effect on memory, which would seem to be impossible if memories are stored in brains. How long before we realize that human memory cannot be a neural thing, but must be a psychic or spiritual phenomenon?

Some people tell tall tales about the protein CAMKII similar to the tall tales told about PKMZeta. We are sometimes told that some alleged autophosphorlyation of CAMKII can help explain stable memories. Most of the reasons I have cited against PKMZeta also apply with equal strength to CAMKII. At this link we are told an experiment debunked the idea that  autophosphorlyation of CAMKII has a role in memory storage.  The lifetime of a CAMKII molecule is only 30 hours, according to this source. The book here makes this statement:

"In the mid-1980's there was much excitement about the idea that autophosphorlyated CaMKII might serve as a self-perpetuating signal that could subserve permanent memory storage. However, a variety of experimental results generated since then suggests that perpetual activation of CaMKII does not occur with LTP-inducing stimulation or memory storage."

This scientific paper says the following:

"Previous models have suggested that CaMKII functions as a bistable switch that could be the molecular correlate of long-term memory, but experiments have failed to validate these predictions....The CaMKII model system is never bistable at resting calcium concentrations, which suggests that CaMKII activity does not function as the biochemical switch underlying long-term memory."

This recent scientific paper says on page 9, "Overall, the studies reviewed here argue against, but do not completely rule out, a role for persistently self-sustaining CaMKII activity in maintaining" long term memory. Another paper says, "The autophosphorylation of CaMKII, once thought to help maintain long-term changes in synaptic strength, has since been revealed to be rather transient." 

Those who have studied the history of science are familiar with epicycles, a complicated speculation that was introduced into Ptolemy's theory of astronomy, to try to fix cases in which the theory did not match observations. We may say these CaMKII speculations and PKMZeta speculations are epicycles intended to fix the failing synaptic theory of memory storage.  But while the Ptolemaic epicycles were exact speculations, the CaMKII speculations and PKMZeta speculations are very vague, failing to specify any exact theory of memory storage. 

Last week a press release from New York University tried to breath life into the dead horse of the mythology of "memory maintenance molecules." The press release was a glaring example of what constantly occurs these days in university press releases: university PR offices boasting about grand accomplishments that were not actually accomplished. We have the untrue claim that a "new study in the journal Science Advances, conducted by a team of international researchers, has uncovered a biological explanation for long-term memories."  No, the study was just more Questionable Research Practices junk science, a study so poorly designed it is a wonder it got published.  We read this:

"It’s been long-established that neurons store information in memory as the pattern of strong synapses and weak synapses, which determines the connectivity and function of neural networks. However, the molecules in synapses are unstable, continually moving around in the neurons, and wearing out and being replaced in hours to days, thereby raising the question: How, then, can memories be stable for years to decades?"

No, it has not ever been established that " neurons store information in memory as the pattern of strong synapses and weak synapses" : no credible theory of how such a thing could work has ever been advanced, no memory information stored in neurons or synapses has ever been found through microscopic examination, and the instability of synapses (including the short lifetimes of their proteins) is the strongest reason for thinking that it cannot possibly be true that "neurons store information in memory as the pattern of strong synapses and weak synapses." 

The press release makes the groundless claim that the junk science paper it is publicizing shows "that KIBRA is the 'missing link' in long-term memories," referring to a molecule called KIBRA.  The press release makes the groundless claim that "their experiments in the Science Advances paper show that breaking the KIBRA-PKMzeta bond erases old memory."

The study is the study "KIBRA anchoring the action of PKMζ maintains the persistence of memory" which you can read here. We have a very bad example of Questionable Research Practices junk science. The study group sizes used are ridiculously small, with study groups as small as only 4 mice and 6 mice, and the average study group size being only about 7 mice. It is frequently pointed out to neuroscientists that experimental studies involving mice are generally worthless unless they use at least 15 subjects per study group; but neuroscientists keep senselessly continuing to use ridiculously low study group sizes.  Why do they do that? Because it allows them to "mine noise," and report false alarms that would vanish if a decent study group size was used. It's rather like someone trying to prove his prophetic powers by publishing a test in which he correctly predicted whether merely four consecutive coin flips were "heads" or "tails," conveniently failing to publish a larger test of his powers involving how well he predicted 15 consecutive coin flips.  You can get all kinds of false alarms when you use tiny sample sizes. 

low statistical power in neuroscience

The junk science paper above relies crucially on an attempt to measure fear, and the attempt to measure fear was the stupid, unreliable technique of attempting to judge "freezing behavior" in mice. All experimental studies relying on estimations of "freezing behavior" are junk science studies, for the reasons I explain in my post here.  There are good reliable ways of measuring fear in rodents, and whether a mouse still has a memory of something the mouse was trained to fear. One good method is to measure heart rate, which dramatically spikes when a rodent is afraid. Another good method is to detect movements in which a mouse avoids a stimulus an animal was trained to fear. The technique is shown in the diagram below:

good way to measure fear in mice

Attempting to measure whether a rodent remembered something fearful by doing estimates of "freezing behavior" is not a reliable way of measuring fear or memory, but instead an unreliable "see whatever you want to see" affair. All studies hinging on so unreliable a method are junk science studies, including the new study by Sacktor mentioned in the New York University press release. Contrary to the claims by Sacktor in his later paper and the claims in the New York University press release, no robust evidence has been produced to show that inhibiting either the KIPRA molecule or the  PKMζ molecule (PKMzeta) does anything to harm the memory of mice. In the paper and the press release Todd Sacktor tells the tall tales of memory maintenance mythology that he likes to tell, unbelievable tales that are not backed up by any robust experiments. 

Genuine "freezing behavior" in an animal would typically only be an instantaneous thing, lasting only a few seconds. The longer the length of time over which "freezing behavior" is judged, the more unreliable such a judgment is as a basis for judging whether the animal recalled a fearful stimulus. In Sacktor's latest paper discussed above, which you can read here,  fear recall is being judged by someone estimating how much an animal was non-moving over a length of four minutes.  That's a particularly unreliable use of the utterly unreliable technique of judging "freezing behavior" as a method of trying to determine whether recall occurred.  

A 2024 scientific paper makes this candid confession, using the phrase "still not completely understood" when it should be saying "are not at all understood":

"Despite over a hundred years of research, the cellular/molecular mechanisms underlying learning and memory are still not completely understood. Many hypotheses have been proposed, but there is no consensus for any of these."

The paper calls Sacktor's speculations about the  PKMζ molecule (PKMzeta) "very controversial." 

I can give you an example that helps show the difficulty of maintaining stable information from unstable components. Let's suppose a shaving creme company wants to publicize its product. It arranges for a line of people to appear outside of the main branch of the New York Public Library, a line in which each person will be displaying a letter made out of shaving creme. The total line of people will spell out the message: "Sale! 10% off on purchases of Barbisol shaving creme, the world's most comfortable shaving creme."  The letters will look like this, but each letter will rest on the outstretched palms of one person.


The problem is that the shaving creme letters will soon decay. How to keep the advertising message running all day?  There could be a system in which each person stands on a numerical position with a number between 1 and 97. If a person sees his shaving creme letter is disintegrating, he then sends a text message to some phone number, saying something like, "I'm leaving -- I'm at position number 12, and my letter is F."  Then some person getting all these text messages can keep sending one of his helpers to each position mentioned in a message, filling in the letter mentioned in the text message.  Through such a system the message might last all day, even though each letter only lasts for less than an hour. Note well some of the requirements here, which include:

(1) A system for representing words by use of symbolic tokens (the English alphabet). 
(2) Some skill for creating these tokens in shaving creme letters (maybe someone who has practiced this skill). 
(3) An addressing system by which each person in the line knows his ordinal position in the line.
(4) A  message system by which components that are about to fail send out a message to some receiving system notifying it to replace their failing token, telling that receiving system of which token to replace, and what the address was of the token to replace. 

No similar system could ever be possible in the human brain. The human brain has no addresses or position numbers or position notation system or coordinate system. Neurons and synapses have no knowledge of the English alphabet, and no capability of writing synapse states or neuron states corresponding to letters of the English alphabet. A protein about to decay in a synapse would never know it was about to decay, and would never be capable of sending some external receiver a "replace me" message. And such a message could never have the address or position coordinates of a synapse protein to be replaced, because tiny components in the brain have neither  addresses nor position coordinates. When you also  consider that synapses are all entangled in 3D space, making them geometrically more difficult to work with than a simple one-dimensional line, you may start to realize how mythical is the notion that stable memories lasting decades could be made from constantly-replaced components with lifetimes of only a few weeks. 

Postscript: A 2018 paper tells us a little about the appalling state of research practices in research involving rodents:

"There is a crisis in pre-clinical biomedical research
involving laboratory animals. Too many papers publish
results which turn out to be irreproducible. One estimate puts the cost at $28 billion being wasted per annum in the United States alone. The causes of this irreproducibility crisis have not
been fully identified. But it has been known for many
years that experiments are often poorly designed, inadequately analysed, and misreported. A survey of
271 papers chosen at random involving rats, mice
and non-human primates showed that 87% did not
report random allocation of experimental subjects to
the treatments and 86% did not report 'blinding' 
when measuring the results. None of the papers gave
any justification for their choice of sample size, and a
substantial number of papers failed even to state the
sex, age or weight of the animal."

A 2021 paper ("Increasing the statistical power of animal experiments with historical control data" by V. Bonapersona et. al. ) gives us the damning graph below:

poor practices in neuroscience

The graph shows an analysis of 1900+ neuroscience papers. A statistical power of 80% is considered a good goal to reach (when such power is reached there will be about an 80% chance that a reported effect is real). The first graph shows that the average neuroscience paper has a miserably weak statistical power of only about 15%.  The graph on the right shows the number of animals used in these papers, with a median of only about 10 per study group. It is largely because of such low study group sizes that the papers are getting such poor statistical power.  The study group sizes in the Sacktor paper discussed above are sub-standard, even within the dismally poor practices being followed by today's neuroscientists, where the median is a way-too-low number of about 10 animals per study group.  We read this:

"For a common effect size of Hedge’s g= 0.5 (Welch’s independent samples t-test, α=0.05), ten animals per group would correspond to a statistical power of 18%, 30 animals per group to 48% power and 65 animals per group to 81% power...Through a systematic search (Supplementary Notes 1 and 2), we identified a large sample of animal studies in the areas of ‘neuroscience’ and ‘metabolism’ (n...=1,935) that were previously included in meta-analyses (n...=69). These animal studies had an overall median statistical power of 18% (Fig. 1a), which was roughly equal in the two fields (neuroscience, 15%; metabolism, 22%)....We estimated that, at best, 12.5% of a large sample of rodent studies were sufficiently powered (that is, prospective power was larger than 80%). This estimate is a best-case scenario, as it is not yet adjusted for any subsequent multiple testing, experimental bias, P hacking and/or fishing, selective reporting, etc.." 

Postscript: Scientific American has an article covering Sacktor's latest paper, giving us yet another example of its very poor journalism regarding neuroscience research. We have an appalling failure to inform the reader of the most basic facts about Sacktor's latest study. At no point is the reader told that the research merely involved rodents. At no point are we told about the appallingly small study group sizes used, such as only 7 rodents.  Sacktor is allowed to get away with a groundless boast that he "nailed it," and no mention is made of how the study hinged upon an unreliable technique for measuring memory (the faulty judgement of "freezing behavior" discussed above).  The author falls for Sacktor's boasts "hook, line and sinker." 

I may also mention that the study fails to discuss how an effective blinding protocol was followed.  We get a mere sketchy vague  statement that an experimenter judging how much "freezing behavior" occurred was "blind to the conditions," but that does not constitute a description of an effective blinding protocol. When very small study group sizes (such as only seven mice) are used it is easy for a supposedly blind judge to know (by visual recognition) whether or not some animals being analyzed were part of some group that had been chemically treated, and which were desired to be described as acting in a particular way, such as "freezing" more. Effective blinding in an experiment (necessary whenever subjective judgments are made) can only be achieved by following a careful, detailed blinding protocol taking at least a long paragraph to describe; and that apparently did not occur in this study. No reliable measurement or reliable analysis of memory performance has occurred. 

Post-postscript: A perennial promoter of junk neuroscience research, Quanta Magazine had a 2025 article promoting Sacktor's groundless boasts, one with the bogus title "The Molecular Bond That Helps Secure Your Memories." Sacktor's paper with the ridiculously low sample sizes such as only four mice or six mice is passed off as a discovery, with no mention at all of the absurdly low study group sizes used. Another neuroscientist guilty of the same type of very poor research practices (such as the use of way-too-small study group sizes and utterly unreliable "freezing behavior" judgments to try to measure animal recall) is quoted as praising Sacktor's research. Again and again the article uses the word "discovery" or "discovered" for something that was not actually discovered, and again and again the article uses the word "showed" for something that was not actually showed. 

Post-post-postscript: The 2025 paper here states that  "the synaptic turnover rate is as high as 1% per day in the visual cortex." That is a rate of about 100% replacement every four months. If that is the typical rate at which synapses are replaced, synapses cannot be the storage place of memories lasting decades. Quoting an even higher rate of synaptic turnover, the paper here states, "A recent imaging study revealed that the synaptic turnover rate in hippocampal CA1 cells is very high, with an estimated lifetime of 1–2 weeks (Attardo et al., 2015)."