Showing posts with label claims of neural correlates of mental activity. Show all posts
Showing posts with label claims of neural correlates of mental activity. Show all posts

Sunday, January 4, 2026

A Paper Gives a Severe Blow to "Neural Correlates of Cognition" Claimants

 Those trying to make claims about neural correlates of cognition have mainly relied on brain scans. The approach of such claimants has been to use fMRI scans, which do not directly measure neural activity, but instead measure differences in oxygen or blood flow in different regions of the brain. The "neural correlates of cognition" theorists try to claim that when some particular cognitive activity is done, some tiny region of the brain shows more activity, as shown in fMRI scans. 

Such claims have always been very dubious, for several reasons:

(1) Typically the reported difference in activity during some cognitive activity as shown in the fMRI scan is some unimpressive tiny difference such as 1 part in 200. There is no reason for thinking that so small a change is any good evidence of any part of the brain working harder when some type of mind activity occurs. We would expect so small a fluctuation to occur merely by chance. 

Those making claims about neural correlates of cognition have tended to be guilty of a very bad type of visual deception I call "lying with colors." When the deception occurs, some difference of only about 1 part in 200 will be depicted in bright red, against a white background. That creates the impression that the difference is some big difference much greater than the actual difference. What goes on is illustrated in the diagram below. 

deceptive fMRI visuals

(2) Typically there is no reliable statistical basis for any claim of a  reported difference in activity during some cognitive activity. To make such a claim with any reliability, you would need to do a study scanning hundreds or thousands of people. But typically a scientific study making such "neural correlates of cognition" claim will involve some way-too-small study group size such as only 15 or 20 subjects. 

(3) There has always been another reason for doubting the reliability of studies of this type: the fact that neither brain electrical activity nor brain chemical activity is being  directly measured by fMRI machines when brain scans are done. Instead, what is being measured is blood flow or oxygen levels in the brain. Those making "neural correlates of cognition" claims will typically try to persuade us that greater blood flow or oxygen means greater brain activity related to cognition. 

Recently a new scientific study cast great doubt on such claims. The study is entitled "BOLD signal changes can oppose oxygen metabolism across the human cortex." A press release announcing the study has the headline "40% of MRI signals do not correspond to actual brain activity, study suggests."  We read this:

"For almost three decades, functional magnetic resonance imaging (fMRI) has been one of the main tools in brain research. Yet a new study published in Nature Neuroscience fundamentally challenges the way fMRI data have so far been interpreted with regard to neuronal activity. According to the findings, there is no generally valid coupling between the oxygen content measured by MRI and neuronal activity.

First author Dr. Samira Epp emphasizes, 'This contradicts the long-standing assumption that increased brain activity is always accompanied by an increased blood flow to meet higher oxygen demand. Since tens of thousands of fMRI studies worldwide are based on this assumption, our results could lead to opposite interpretations in many of them.' "


This result is a giant torpedo in the hull of those who have tried to use brain imaging studies to back up claims of "superior activation" during certain types of brain activity. It gives all the more reason for thinking that the vast majority of such studies are worthless junk, suitable only for activities such as wrapping freshly caught fish and lining bird cages. 

boasts of brain scientists
A sinking ship

Postscript: A year 2026 paper by 9 scientists says this: "Despite many years of research, the quest to identify neural correlates of perceptual consciousness (NCC) remains unresolved."

Thursday, December 18, 2025

No, Brains Have Nothing Like a Large Language Model (LLM)

Quanta Magazine is a widely-read online magazine with slick graphics. On topics of science the magazine again and again is guilty of the most glaring failures. Quanta Magazine often assigns its online articles about great biology mysteries (involving riddles a thousand miles over the heads of PhDs) to writers who lack even a bachelor's degree in biology, and who may also lack any history of writing very much about biology. Often it will assign such articles to be written by people identified as "writing interns."  The articles at Quanta Magazine often contain misleading prose, groundless boasts or glaring falsehoods. I discuss some examples of such poor journalism in my posts here and here and here and here.

The writers at Quanta Magazine often sound like the most credulous pushovers for scientists making dubious boasts. They often seem like the science journalists depicted below:

pushover science journalists

A recent  example of a puff-piece article in Quanta Magazine is its fawning article entitled "The Polyglot Neuroscientist Resolving How the Brain Parses Language." The article title is a very misleading one. It is not brains that parse language. It is people who parse language. Parsing language (interpreting what someone meant upon hearing or reading something) is an example of a very high-level mental faculty.  Neuroscientists have no credible tale to tell of how a brain could produce such a faculty. 

The puff piece article is written by someone identified as a "writer and filmmaker," who has written many articles at the Quanta Magazine dealing with topics related to AI and technology, but almost no articles at the site relating to psychology or neuroscience.  What we get is an article devoted to glorifying neuroscientist Ev Fedorenko. You can tell what a "going gaga" scientist glorification affair is occurring by the fact that the article has five huge photos of Fedorenko, each of which fills up my computer monitor. 

We have this extremely misleading statement by the neuroscientist, not matching anything ever found in a brain:

" 'You can think of the language network as a set of pointers,'  Fedorenko said. 'It’s like a map, and it tells you where in the brain you can find different kinds of meaning. It’s basically a glorified parser that helps us put the pieces together — and then all the thinking and interesting stuff happens outside of [its] boundaries.' ”

The brain has no maps, and it has no pointers. I know pointers well, having used them extensively when I was a C programmer early in my programming career. A pointer is a variable that stores the address where some data is stored. Brains have no addresses, and nothing corresponding to a programming variable. So in a brain there can be nothing like a pointer. 

The article has a reference to a scientific paper by Fedorenko, one entitled "The language network as a natural kind within the broader landscape of the human brain." The paper says, "In this Review, we discuss brain areas that are specific to language —what we refer to as the language network — and position them in relation to perception, motor planning and cognition (Fig. 1)." There are no regions in the brain "specific to language" in the sense of being able to parse language. 

The paper announces that it will back up such claims largely by appealing to brain scan studies, saying, "We primarily draw on fMRI data from studies that have relied on the individual-subject functional localization approach14,19 (Box 1), which was essential in clarifying the distinctions discussed.." We then have in Figure 1 an extremely dubious visual showing five brains. We have some "function localization" claims:

  • One brain visual has a small fraction of it colored blue, and this is labeled as the "perception" area of the brain. 
  • Another brain visual has a very small fraction of it colored red, and we are told that is the "motor planning" area of the brain. 
  • Another brain visual has five parts of it colored purple, and we are told that this is the "language" areas of the brain. 
  • Another brain visual has quite a few parts of it colored green, and we are told that this is the "knowledge and reasoning" areas of the brain. 
  • Another brain visual has quite a few parts of it colored green, and we are told that this is the "intended meaning" areas of the brain. 

People have been making these kind of functional localization claims about the brain for almost 200 years, and almost always the claims have been very dubious. For example, below is an illustration from the beginning of the 1834 book A System of Phrenology by George Combe. Notice how the brain areas have little numbers next to them, with the bottom legend explaining the claims made about localization of function in the brain. 

phrenology

What basis does Fedorenko have for the brain function area maps in her paper, Figure 2, one rather reminding you of the phrenology image above? She says her basis is brain scan data.  The Supplementary Data part of the paper promises that if you click on a link you will get the data that was used to generate the brain function area maps. Following the links takes you to a little page that is almost worthless for inspecting data, as it involves a .zip file consisting of files in an .nii format that the average person will be unable to load or read. Consequently it is all but impossible for anyone to check the evidence basis for Fedorenko's brain function area maps, which should not be regarded as reliable. Neuroscientists dramatically disagree with each other when they produce such brain function area maps, which typically fail to have a sound evidence basis. 

A figure on page 294 of the paper  gives us the bar chart below: 

The purple bars at left are probably incorrect

The scale on the left shows us that the bars refer to fMRI percent signal changes in different regions of the brain during different activities. You should note well that the largest bars show a variation of only 2%. For almost all of the activities listed, the percent signal change is less than 1%. On average the change in brain signal strength is merely about 1 part in 200. A change that small is no real sign of brains working harder during some cognitive activity. We might expect that mere chance variations would produce differences that small. 

Except for the two purple bars on the left-most edge, the diagram is consistent with what I have often stated on this blog. For example, in my post "Brain Imaging Shows No Appreciable Neural Correlates of Memory Activity," I quoted quite a few studies showing a variation of only  1 part in 200 when people had their brains scanned by fMRI scanners while they were doing various memory-related activities. I noted that so small a change fails to provide any good evidence of brains causing mental activities, as we might expect a 1 in 200 fluctuation in brain activity to occur by chance, even if brains don't make minds. 

But what about the purple bars on the left of the chart? What is the basis for the claim being made in the graph that during sentence comprehension there is up to a 2% change in brain signal strength? We fail to get a justification for the data shown. The references in the paper include some papers referring to sentence comprehension. But none of them seem to back up the claim above. Specifically:

  • We have a reference to a paper "Cognitive control and parsing: Reexamining the role of Broca’s area in sentence comprehension." It makes no claim about percent signal changes during sentence comprehension. 
  • There's a reference to a paper "Retrieval and Unification of Syntactic Structure in Sentence Comprehension: an fMRI Study Using Word-Category Ambiguity." It's behind a paywall, and its abstract makes no claim about percent signal changes. 
  • There's a reference to a paper "The cortical language circuit: from auditory perception to sentence comprehension."  It's behind a paywall, and its abstract makes no claim about percent signal changes. 
  • There's a reference to a paper "fMRI reveals language-specific predictive coding during naturalistic sentence comprehension." The paper does not make any claim about a percent signal change during sentence comprehension. 
  • There's a reference to the paper "Sentence complexity and input modality effects in sentence comprehension: an fMRI study." It's a study involving only 20 subjects, and does not claim in the main body of its text to have detected any percent signal change of 1% of higher. But there is a footnote in which the authors say that after doing some dubious-sounding fiddling with the data they got some kind of 1% difference of some type. We cannot have much confidence in such claim, as it stated only in a footnote. 
  • There's a reference to the paper "Form and Content: Dissociating Syntax and Semantics in Sentence Comprehension." It does not make a claim about a percent signal change. 
  • There's a reference to a paper "Location of lesions in stroke patients with deficits in syntactic processing in sentence comprehension." We can ignore it, because the issue is how much normal brain signals change during sentence comprehension.
  • There's a reference to a paper "Time course of semantic processes during sentence comprehension: an fMRI study." It does not make a claim about a percent signal change. 
Rather than relying on the references Fedorenko has given, we can follow the alternate approach of doing a Google image search using the phrase " 'percent signal change' + ' sentence comprehension' ." This gives us papers such as these:
  • The paper here involving sentence comprehension indicates a percent signal change of less than 1 part in 200 (less than half of a percent). 
  • The paper here ("Neural correlates of syntactic movement: converging evidence from two fMRI experiments") reports a relatively large percent signal change of about 1 part in 100 for people listening to sentences. But the study group sizes are so small (involving only 11 subjects for one experiments, and 10 subjects for another experiment) that the paper cannot be counted as good evidence for anything. 
  • The very low-quality paper here ("Neural correlate of the construction of sentence meaning") co-authored by Fedorenko must be disregarded because it reported a "percent signal change" based on EEG readings rather than fMRI readings, because it used a sample size of only 6 epilepsy patients with implanted electrodes, and also because of the unreliability of looking for percent signal changes in the brain waves of seizure-prone patients, whose brains produce all kinds of unpredictable EEG spikes. 
  • The paper here ("Top-down and bottom-up contributions to understanding sentences describing objects in motion")  used an almost equally poor study group size of 12 subjects, and found a percent signal change of .05 percent, 1 part in 2000. 
  • The paper here ("Language processing in the occipital cortex of congenitally blind adults") used one too-small study group size of 9 blind adults and another possibly halfway-adequate study group size of 22 control subjects with normal vision. Its Experiment 1 involving language processing found a percent signal change of less than .1 (less than 1 part in 1000) for the 22 control subjects. Another experiment involve language processing found a percent signal change of less than .5 (less than 1 part in 200) for the 22 control subjects.
  • The paper here ("Brain activity associated with selective attention, divided attention and distraction") finds about a 1% percent signal change in the auditory cortex during language processing. But the study group size is an unimpressive 15 subjects. This auditory cortex region of the brain is not part of the "language processing" area claimed by Fedorenko. In her paper she refers to some regions of the brain and says, "These areas are distinct from the language network as well as from general-purpose sensory and motor areas, such as the primary auditory or primary motor cortex," apparently indicating that what she thinks is a "language network" in the brain is something outside of the auditory cortex. 
Typically there is no test/retest reliability when doing these little studies involving only 15 or 20 or 25 subjects. Robust evidence would only come from a much larger study group size. The New Scientist story below tells us that thousands of participants are needed for studies like these:


The purple bars on the graph shown above are very probably incorrect. Scientific papers documenting brain scanning during language comprehension do not show any robust evidence of a 1% or 2% change in brain activity during sentence comprehension. And there is no sound evidential basis for Fedorenko's diagram claiming to show the location of five language areas of the brain. 

As discussed here, a recent study indicated that 40% of MRI signals do not even correspond to brain activity, which further undermines the knowledge boasts that  Fedorenko has made, and provides another reason for disbelieving in the accuracy of her brain activity maps.  We read this: "Researchers at the Technical University of Munich (TUM) and the Friedrich-Alexander-University Erlangen-Nuremberg (FAU) have found that an increased fMRI signal is associated with reduced brain activity in around 40% of cases. At the same time, they observed decreased fMRI signals in regions with elevated activity." This is the kind of finding that should make us disbelieve Fedorenko's fMRI-based claims of a brain "language network." 


The puff piece article in Quanta has a "softball questions only" interview with Fedorenko in which she states this:

"There’s a core set of areas in adult brains that acts as an interconnected system for computing linguistic structure. They store the mappings between words and meanings, and rules for how to put words together. When you learn a language, that’s what you learn: You learn these mappings and the rules. And that allows us to use this 'code' in incredibly flexible ways."

There is no robust scientific basis for these imaginative claims. To the contrary, the most powerful microscopes have examined very much brain tissue from very many still-living people and very many very recently deceased people, and no one has found the slightest trace of any sentence, word or letter stored in a brain. 

Later on Fedorenko states, " Then the language network parses that, finding familiar chunks in the utterance and using them as pointers to stored representations of meaning." Excluding the merely genetic chemical representations in DNA, there is zero evidence for any such "stored representations of meaning" anywhere in the brain.  Later she kind of gives away the speculative nature of her claims, by saying "there may well be cells that respond to particular aspects of language."  Note the use of "there may well be" rather than "there are."   

Fedorenko offers as evidence for this claim the preprint "Modality-Specific and Amodal Language Processing by Single Neurons."  It is a "Jesus in my toast" exercise in noise-mining. Recordings of neuron firings were taken after 1400 invasive microwire electrodes had been implanted in the brains of 21 very sick patients with the worst types of epilepsy. The authors tried to find some neurons that fired more often when certain words were spoken. Because neurons fire between 1 and 200 times per second, any such exercise will always be able to find some neurons that fired more often when certain words were spoken. It is misleading to call such a thing "selectivity" as the paper did. Similarly, if I set up a computer program to try to correlate the passing of cars outside my house and the speaking of words on my TV set, and let the program run long enough, I will probably be able to find that certain words were spoken more commonly on my TV set when a car passed in front of my house.  That would be mere silly noise-mining, and that is all that is going on in this preprint, which fails to provide any decent evidence of any neuron responses to language different from what we would expect from chance variations, giving 1400 microwires implanted in brains. 

Noise-mining studies like this raise grave moral concerns about the reckless research endangerment of very sick epilepsy patients.  The sickest of epilepsy patients often have electrodes implanted for evaluation of where surgery should be done to reduce their symptoms. But such electrodes are usually not the microwire electrodes used for experiments like this. The deep implanting of such microwire electrodes in brains (70 microwires per patient for this study) has serious risks that are not medically justified, such as a risk of brain bleeding.  We should doubt any claim to have got any adequate degree of "informed consent" from the seizure-racked people who have these kind of deep microwire electrode implants. 

A scientific paper tells us, "Sixty-five years after single units were first recorded in the human brain, there remain no established clinical indications for microelectrode recordings in the presurgical evaluation of patients with epilepsy (Cash and Hochberg, 2015)." In other words, there is no medical justification for implanting microelectrodes or microwires in the brains of epilepsy patients. Complications from the insertions of such electrodes may include death, with two deaths reported here. I consider the exploitation and endangerment of sickest epilepsy patients in these type of poorly designed noise-mining experiments to be a serious moral scandal. 

research abuse of epilepsy patients
Does it work like this?

For much more on the topic of why most subjects in neuroscience experiments do not understand the risks they are taking, read my post here. Epilepsy patients requiring surgery are often those who cannot read well (because of learning delays caused by their very frequent seizures), and who will not well-understand any jargon-filled form they are asked to sign. Typically neuroscientists doing electrode implant experiments with epilepsy patients fail to provide in their papers the informed consent form used, and also fail to describe whether special measures were taken to make sure those getting the microwire electrode implants really understood the risks of participating in an experiment of this type. 

We should laugh at Fedorenko's answer when she is asked whether there is an LLM (Large-Language Model) inside the brain, and she replies "pretty much." Nothing like that exists in the brain. An LLM or Large Language Model is a gigantic structure of data, computer programming software and data processing software built after some big server farm (consisting of very many individual computers) crunches data gobbled up by crawling the Internet, scooping up the contents of millions of web pages. That does not correspond to anything in a human brain. 

The idea of a language-parsing ability in a brain seems implausible when you consider that human brain anatomy has not changed in more than 20,000 years, but the main human languages (such as English) are less than 3000 years old. 

It is commonly claimed that the left half of the brain is needed for language. But if you read my series of posts labeled "loss of left half of brain," using the link here, you will find many cases that defy such a dogma.  For example, you will read of a case described like this: "He exhibited no lack of intelligence, yet after his death it was discovered that his left brain was practically destroyed and replaced by a watery substance." And you will read of how Beth Usher could still tell all of her "knock-knock" jokes just after the left  half of her brain had been removed to treat intractable seizures. You will read that on page 109 of a paper we read of three other cases of the removal of the left half of the brain, and we read that "speech and verbal comprehension were present immediately after left hemispherectomy in all three cases." The same post describes good verbal comprehension in patient E.C. after the left half of his brain was surgically removed. 

In a scientific paper ("Why Would You Remove Half a Brain? The Outcome of 58 Children After Hemispherectomy −−The Johns Hopkins Experience: 1968 to 1996") we read about how surgeons at Johns Hopkins Medical School performed fifty-eight hemispherectomy operations on children over a thirty-year period. Eleven of these children had the left hemisphere of their brains removed; most of the rest had the right hemisphere of their brains removed.  The paper states this:

"Despite removal of one hemisphere  [i.e. one half of the brain], the intellect of all but one of the children seems either unchanged or improved....Although there have been major concerns about loss of language after left hemispherectomy, all eleven of these children have regained virtually normal language....It is tempting to speculate, that the continuous electrical activity of these severely dysfunctional hemispheres interferes with the function of the other, more normal hemisphere. This might explain why motor function improves after hemispherectomy and why language recovers after removal of the dysfunctional left hemisphere, but does not seem to fully transfer before surgery. Perhaps it also partially explains intellectual improvement in these children after removal of half of the cortex. We are awed by the apparent retention of memory after removal of half of the brain, either half, and by the retention of the child’s personality and sense of humor."

Monday, November 25, 2024

The Brains of 60 Subjects Seemed to Look the Same During Eyes Closed Mental Rest, Recall and Math Activity

The EEG is a device that can detect electrical activity from parts of the brain. When an EEG device is used, electrodes are placed next to different parts of the skull. The device will pick up a dozen or more different lines that show electrical activity in different parts of the brain. 

Brains have a great deal of signal noise, and the abundance of such noise is one of several major reasons for disbelieving that the brain is the source of human thinking and recall which can occur with incredible accuracy, such as when people perfectly recall very large bodies of text and perfectly perform extremely difficult math calculations without using tools such as computers, pencils or paper. The analysis of brain waves obtained by EEG devices is an area of science where bad methods, pareidolia and junk analysis is very abundant.  There is an abundance of people trying to use fancy statistical methods to try to extract identifiable "signals" or "signs" from data that is very noisy and polluted. Muscle movements abundantly contaminate EEG readings. 

A widely used publicly available dataset of EEG data is available on a site called Physionet. On a page entitled "EEG During Mental Arithmetic Tasks" it is possible to download EEG data for 36 subjects. The data includes EEG readings taken during "rest activity" and EEG readings taken when the subjects were told to perform mathematical operations.  The paper here ("Electroencephalograms during Mental Arithmetic Task Performance") describes how the data was gathered.  The data set is sometimes called the "EEG During Mental Arithmetic Tasks" or it may be called something like the "Physionet EEG mental arithmetic task dataset."

I don't recommend trying to download this data, because it uses some file format that your spreadsheet or text editor will not be able to understand.  But at the page here, we have some comments by a person who downloaded this data, and also downloaded a utility program that allows him to see the data represented as particular wavy lines. 

After showing us a picture showing one subject whose brain wave lines looked different when he was doing the math tasks, the writer states, "Other participants didn’t see much change at all while doing their tasks." By this he means that when he looks at the brain waves of such participants, they don't look different when the subjects were doing the math tasks (compared to when they were resting). The writer also states, "In fact, some data looked like the brain had more activity while doing nothing at all."  We see one visual with brain wave lines showing "baseline" activity for Subject 15, and another visual showing brain wave lines during that subject's performance of math tasks. The first visual shows wavy lines that are a lot wavier that the second visual, contrary to the idea that mental activity would involve more active brain waves. 

You can read some scientific papers written by scientists that create algorithms or models that analyze data sets such as this, algorithms or models trying to detect whether a particular set of EEG readings was or was not taken when a patient was engaging in heavy thinking. A typical paper of this type will discuss several different algorithms or models the scientists tested. We may be told that the most successful algorithm had something like a 75% success rate in predicting whether a set of EEG readings were produced rest activity or thinking.  

Such a thing is unimpressive when you consider that the data set being used for testing is usually small. In many cases half of the patient data will be used to "train" the model, and the other half will be used to test the model. So maybe the data for only 8 or 10 patients will be used to test the model.  The odds of accidental success on guessing whether the person's mind was active or not (even if the model is worthless) are something like this (I used the StatTrek binomial probability calculator to calculate some of the odds):

Eight patients:

Chance of 8 guesses all correct = 2 to 8th power = 1/256.

Chance of 7 out of 8 guesses correct = .035

Chance of 6 out of 8 guesses correct = .014

Ten patients:

Chance of 10 guesses all correct = 2 to 10th power = 1/1024

Chance of 9 out of 10 guesses correct = .01

Chance of 8 out of 10 guesses correct = .05

Chance of 7 out of 10 guesses correct = .17

Now, with odds like these it means very little if some scientific paper says that it tried several different predictive models, and found that one of the models had a 70% predictive accuracy. You might rather easily get that level of success by pure chance, even if the model is worthless or if the "mental activity" scans have no identifying characteristics.  We must also remember here factors such as what is called publication bias and what is called the file drawer effect. Publication bias is that scientific journals tend to reject negative results, and accept for publication only papers reporting positive results. The file drawer effect is that scientists are free to try different things without publishing their failures, and without submitting failed attempts for publication. So a scientist who produces a slightly successful predictive model analyzing EEG data may have in his file drawers 40 failed attempts involving unsuccessful predictive models. Getting maybe a "70% successful" predictive model on the 20th or 30th try does not mean that the EEG data actually shows a difference when people are thinking versus when their minds are resting. 

Then there is the fact that the gathering of EEG data must be done very carefully for any data set that compares intensive mental activity with rest activity. Visual activity, muscle activity and stress can produce traces in EEG data.  So, for example, it might be easy to detect the difference between rest activity and mental activity if the subject is motionless and closes his eyes during rest activity, and the subject uses a keyboard to type answers during the mental activity.  In that case the difference would come from the fact that during the rest activity there is no use of the eyes and muscles, and during the mental task there is use of the eyes and muscles. 

The paper here ("Electroencephalograms during Mental Arithmetic Task Performance") describes some poor methods of gathering rest data and mental arithmetic data used to create the "EEG During Mental Arithmetic Tasks" data set that has been the basis of quite a few scientific papers.  We are told this:

"Mental arithmetic performance is considered as a standardized stress-inducing experimental protocol. Serial subtraction during 15 min is considered to be a psychosocial stress. In this way, our study design required intensive cognitive activity from the subjects. Intensive mental load is accompanied by a change in the emotional background when the subject makes additional effort to resolve tasks, so one can talk about evoked emotions in this case.. During EEG recording, the participants sat in a dark soundproof chamber, comfortably reclined in an armchair. Prior to the experiment, participants were instructed to try to relax during the rest state and were informed about the arithmetic task—participants were asked to count mentally without speaking or using finger movements, accurately and quickly, in the rhythm they had determined. After 3 min of adaptation to experimental conditions, EEG registration of the rest state with closed eyes was made (over the next 3 min). Then the participants performed a mental arithmetic task—serial subtraction—for 4 min."

We are also told that the scientists kept only a subset of the original data gathered, throwing out about half of the data:

"Based on EEG visual inspection by a qualified electroneurophysiologist, 30 of the 66 initial participants were excluded from the database due to poor EEG quality (excessive number of oculographic and myographic artifacts), so the final sample size is 36 subjects."

It is easy to see how that could have gone wrong. The desire to get a set of EEGs with mental activity brain scan data looking during different from rest state brain scan data might have come into play, creating a bias in so subjective a selection of which subjects to keep. 

There's much gone wrong here. We have no description of a rest state which is a clear description of a lack of mental activity. Were the subjects hearing something told them during the rest state? That isn't a rest state. Did any of the subjects move during the rest state? That isn't a rest state. Were the subjects counting during the rest state? We can't even tell from the wording above. Did the subjects have their eyes closed when they were doing the mental subtractions? We don't know. Were the subjects disqualified if they violated the instructions by softly speaking as they counted backwards? Apparently not. The subjects were told to follow a rhythm during counting, an instruction which might have tended to produce sounds or motions such as tapping. The subjects were not told to be motionless, but merely told not to use their fingers (an instruction that would not exclude arm movements or foot tapping movements or a rocking motion in their reclining armchair). Also, the subjects were asked about what was the final number after their mental subtractions. That might have created a possible element of anxiety, in which people would be worried about whether the final number (after their mental subtractions) would be a correct one. Such anxiety might have shown up in the EEG readings, which might have shown signs of anxiety that were not signs of mental effort.  Also, based on subjective whims of a human judge only about half of the data collected has been put in the public data set. The mental activity requested (serial subtraction) is a mental activity that almost seems designed to create distress and frustration in subjects, which may show up as EEG blips that are not signs of thinking. 

Data like this has no value unless there is a crystal-clear description of the exact procedure used during the rest state and the mental activity state. That description should include a precise detailing of whether the subjects had their eyes opened, an exact quotation of what they were told, an exact description of whether the subjects moved or spoke, a description of what (if any) methods were used to prevent the subjects from moving, and so forth.  Comparing mental rest states and mental activity states (from EEG data) cannot be done effectively unless the mental activity states occur under the exact sensory conditions and movement conditions of the rest state, and it would seem the only good method would be for patients to have eyes closed (without any sounds) both in the rest state and the mental activity state, without any possible source of mental anxiety in either state.  All papers based on the data set described (the "EEG During Mental Arithmetic Tasks" data set) would seem to have little value because of the failure (in the paper describing how the data was gathered) to describe an effective, well-documented protocol for distinguishing between real rest activity and sightless, soundless, motionless mental activity without any element of potential anxiety. 

I can tell you how a valid data set of EEG data might be created for the comparison of rest data and mental activity data.  People would be blindfolded in a dark silent room. They would be told that when they first hear a first electronic beep, they should remain motionless for two minutes and think of absolutely nothing other than the blackness of outer space. They would be told that when they hear the second beep, they should remain motionless and start some arithmetic activity such as adding the number 7 continually, continuing for two minutes until they hear the third beep, at which point the EEG readings will stop. The people would also be told to remain motionless and without any expression throughout the whole four minutes of testing.  They would also be told that no one will ask them what the final number was in their minds, so that there is no reason for any anxiety. They would be told, "Don't worry at all if you think one of your numbers is wrong -- just keep adding 7 to whatever was your last number was." A variety of sensitive motion detectors could be used to exclude any subjects who moved significantly. The number of subjects in the final data set would be at least 60, requiring an original pool of test subjects much greater.  Exclusion of subjects would be based on an objective criteria such as motion detector activation, rather than some arbitrary exclusion based on subjective human exclusions. Heart rate data would be gathered, and any subjects showing signs of increased heart rate during the mental activity phase (a sign of stress) would be excluded from the data set. Sensitive sound detectors would also listen for people who softly counted the numbers, excluding such subjects. Ideally, the subjects would wear mouth devices preventing any soft counting. 

Papers based on an analysis of data gathered in such a way (with a sufficient study group size) would fail to show any analysis method correctly predicting whether the rest state or the mental activity state occurred, tending to confirm the idea that thinking is not actually produced by the brain.  The accuracy of any such method over multiple tests would never be some high percentage such as 80%. 

In neuroscience papers attempting to do EEG analysis to find neural correlates of mental activity,  we tend to see some of the same problems found in papers attempting to do fMRI analysis to find neural correlates of mental activity.  The biggest problem is insufficient study group sizes.  Claims are made such that if you analyze some EEG data in such-and-such a way, you will be able to tell (with such-and-such an accuracy) whether or not mental activity occurred.  The claims are made on the basis of tiny data sets such as 8 or 10 or 12 patients. Such claims should never convince unless they are done on large data sets involving more than 50 subjects, and unless the data sets are fully documented by a discussion of a sound procedure used to gather the data sets.  Almost always what is being picked up is not signs of mental activity but signs of muscle activity, speech, vision or emotional states. 

Here are some examples of papers that we should not be taking seriously because of defects I will mention. All of these are examples of "how not to do an EEG study looking for brain wave correlates of mental activity." 

  • "What does delta band tell us about cognitive processes: A mental calculation study" (link). The study got data on only 18 subjects. The mental calculation activity required muscle movement, and the rest activity did not. So the EEG data was not gathered so that pure mental activity was compared to pure mind resting, and "neural correlates of thinking" claims are invalid. 
  • "Real-Time Mental Arithmetic Task Recognition From EEG Signals"  (link).  Data was not gathered in a way to exclude physical differences between rest states and activity and  not gathered in a way to exclude emotional differences between rest states and mental activity.  We are told, "In the relax task, subjects were asked to open their eyes and try to be relaxed. There was no mental arithmetic task to fulfill in this session. Subjects were required to breathe deeply and focus on their breath."  Then we are told in the mental activity state "subjects were required to complete arithmetic calculations as quick as possible." Any differences detected may have been due purely to differences in stress, differences in muscle activity and differences in breathing.  
  • "EEG activation patterns during the performance of tasks involving  different components of mental calculation" (link). We have no description of a data gathering method that excluded muscle activity or caused identical levels of muscle activity during the rest period and the mental calculation period.  Any differences detected may have been due purely to differences in muscle activity. 
  • "EEG microstate features according to performance on a mental arithmetic task" (link). This paper has little value because it used the "EEG During Mental Arithmetic Tasks" data set which is defective for reasons I have explained above. 
  • "Automated Classification of Mental Arithmetic Tasks Using Recurrent Neural Network and Entropy Features Obtained from Multi-Channel EEG Signals" (link). This paper has little  value because it used the "EEG During Mental Arithmetic Tasks" data set which is defective for reasons I have explained above. 
  • "Impact of mental arithmetic task on the electrical activity of the human brain" (link). This paper has little value because it used the "EEG During Mental Arithmetic Tasks" data set which is  defective for reasons I have explained above. 
  • "Mental arithmetic task detection using geometric features extraction of EEG signal based on machine learning" (link). This paper has little value because it used the "EEG During Mental Arithmetic Tasks" data set which is  defective for reasons I have explained above. 

  • "Do specific EEG frequencies indicate different processes during mental calculation? (link). The EEG data was gathered from only ten subjects, and the "rest" state involved no real rest, but looking at a visual and saying, "Nothing." The math calculation involved hard problems such as "a complex arithmetic task, e.g. (24 + 39)/9 = , to which the subject had to give the solution verbally immediately after a warning response signal was presented,"  We have no description of a data gathering method that excluded muscle activity or caused identical levels of muscle activity during the rest period and the mental calculation period.  Any differences detected may have been due purely to differences in muscle activity or differences in stress between the easy task of saying nothing and the stressful task of having to answer the hard math problem "immediately." 
  • "Mental Arithmetic Task Recognition Using Effective Connectivity and Hierarchical Feature Selection From EEG Signals" (link). EEG data was gathered from 29 subjects who we are told alternated between a short period of "mental arithmetic" and "rest." We have no indication of whether this "mental arithmetic" was silent or involved speech or muscular activity.  So we can't tell whether muscular activity was the same during the rest period and the mental activity period. 
  • "Mental arithmetic task classification with convolutional neural network based on spectral-temporal features from EEG" (link). This study used a too-small dataset made from only 12 subjects.
  • "Electroencephalographic Study of Real-Time Arithmetic Task Recognition" (link). There were only eight subjects, and a professional EEG equipment was not even used, but only a cheap consumer device.  There was also no rest state for comparison. 
  • "EEG Based Mental Arithmetic Task Classification Using a Stacked Long Short Term Memory Network for Brain-Computer Interfacing" (link). This paper has little value because it used the "EEG During Mental Arithmetic Tasks" data set which is defective for reasons I have explained above. 
  • "A Modified Multivariable Complexity Measure Algorithm and Its Application for Identifying Mental Arithmetic Task" (link). This paper has little value because it used the "EEG During Mental Arithmetic Tasks" data set which is defective for reasons I have explained above. 

The paper "Investigating neural efficiency of elite karate athletes during a mental arithmetic task using EEG" discusses a relatively good protocol for gathering data during rest and mental activity. We are told that during the rest stage subjects were told to keep their eyes closed and do nothing, and during the activity stage subjects kept their eyes closed and silently counted backward from 600, subtracting 3 each time (e.g. 597, 594, 591, and so forth).  But there were only ten subjects, and the paper does not report any great success in distinguishing rest states and activity states, with the investigators concentrating on other things.  

A scientific paper ("A test-retest resting, and cognitive state EEG dataset during multiple subject-driven states" by Yulin Wang and others) laments, "Given the various advantages of EEG including non-invasive, high temporal resolution, easy-to-operate, and cheap as a neuroimaging technique, it is surprising that there exist relatively fewer high-quality, open-access, big EEG datasets when compared to magnetic resonance imaging (MRI) datasets to enable the investigation of the brain function." Correct. In general, neuroscientists involved in EEG analysis have not done their job correctly, and have failed to create large publicly available brain wave EEG data sets using very careful methods like those I describe above, which would minimize the confounding factors of signal artifacts created by muscle movement and emotional states. 

The paper tries to help this situation by creating an EEG public dataset. The effort has some good elements,  but some shortcomings.  Data was gathered for 60 subjects during an eyes open rest state, an eyes closed rest state, and some mental activity states. We read this:

"During resting-state EEG recording, participants were instructed to view a fixation point for five minutes (Eyes Open) and then close eyes for another five minutes (Eyes Closed). They needed to keep still, quiet, and relaxed as much as they can, and try to avoid blinking for Eyes Open (EO) session and stay awake for Eyes Closed (EC) session. EEG cognitive state:  The present experiment consisted of three subject-driven cognitive states: retrieval of recent episodic memories, serial subtractions, and (silent) singing of music lyrics."

Alas, we are not told whether there was any method to exclude subjects who did not follow the instructions to "keep still, quiet and relaxed as much as they can" (methods such as motion detectors), and we do not know whether subjects failing to follow such instructions were excluded. Also, we are not told that the same instructions to "keep still, quiet and relaxed as much as they can" were given to the subjects while they were performing the cognitive tasks. So we don't know whether the levels of motion were the same when the subjects rested and when they did the cognitive tasks. But on the plus side, the number of subjects used (60) is pretty good, and there is also a good "test/retest" feature in which each subject was tested on multiple days. 

Figure 6 of the paper gives us this very interesting visual showing something called the "averaged power spectrum" for all of the 60 subjects. We have five colored lines, two of which (light blue and yellow) represent the rest states, and the other representing the mental activity states. It is interesting that all of the lines are the same, except that for the "eyes open" rest state, part of the line looks a little different. Referring to an "eyes-closed" that was a state of mental inactivity, the paper tells us "the spectrum of the four states of eyes-closed, subtraction, music, and memory are particularly similar." 

EEG rest versus activity

This is what we would expect under a "your brain does not make your mind" assumption. There is no significant brain signal difference between someone resting his mind with his eyes closed, and someone doing mental activities.  We see something similar in Figure 7 of the paper, which shows us something called the "power distribution of alpha rhythm." The Eyes Closed rest state (EC, in which people's minds were supposed to be inactive) looks the same as when the people were doing mental activity and mental recall. The last four columns on this chart all look the same, and the second column is the Eyes Closed rest state (the last two columns being memory activity and math activity). 

EEG rest versus mental activity

I find the two visuals above to be quite consistent with the claim that your brain is not the source of your mind and not the storage place of your memories. 

Saturday, June 24, 2023

Why Most Correlation-Fishing Experimental Neuroscience Is Worthless

At the Myths of Vision Science blog, written by a vision scientist, there is a post quoting quite a few juicy tidbits in which neuroscientists speak candidly about how they are stumbling about in the dark and following poor methods. We hear of neuroscientists calling their work BS that doesn't replicate. The author states, " As Mehler and Kording (2018) have discussed, despite the intrinsic inability of post hoc sample correlations to generalize due to massive confounding, neuroscience practitioners continue to employ language in their publications improperly implying causality."  In another post, the scientist author explains the situation more clearly, stating this: 

“Neuroscience, as it is practiced today, is a pseudoscience, largely because it relies on post hoc correlation-fishing....As previously detailed, practitioners simply record some neural activity within a particular time frame; describe some events going on in the lab during the same time frame; then fish around for correlations between the events and the 'data' collected. Correlations, of course, will always be found. Even if, instead of neural recordings and 'stimuli' or 'tasks' we simply used two sets of random numbers, we would find correlations, simply due to chance. What’s more, the bigger the dataset, the more chance correlations we’ll turn out (Calude & Longo (2016)). So this type of exercise will always yield 'results;' and since all we’re called on to do is count and correlate, there’s no way we can fail. Maybe some of our correlations are 'true,' i.e. represent reliable associations; but we have no way of knowing; and in the case of complex systems, it’s extremely unlikely. It’s akin to flipping a coin a number of times, recording the results, and making fancy algorithms linking e.g. the third throw with the sixth, and hundredth, or describing some involved pattern between odd and even throws, etc. The possible constructs, or 'models' we could concoct are endless. But if you repeat the flips, your results will certainly be different, and your algorithms invalid...As Konrad Kording has admitted, practitioners get around the non-replication problem simply by avoiding doing replications.”

Here is an example. Suppose neuroscientist Joe wants to show that some particular region of the brain is activated more strongly than normal when people recall an old memory. Joe has 10 people undergo brain scans, and at a particular point in time Joe asks them to recall some memory from their childhood. Later Joe scrutinizes the brain scans of the people. He is looking for some tiny hundredth or thousandth of the brain that shows a little more activity when the recall occurred. Since brain areas have random variations in activity from moment to moment, differing in activity by a hundredth or a two hundredth from one minute to the next, it will be rather easy for Joe to find some hundredth or thousandth of the brain that he can declare as showing greater activity when the subjects recalled their old memories. 

Joe is engaging in correlation fishing, what is sometimes called data mining. If Joe merely reports a difference in some brain area of a hundredth or a two-hundredth, he has provided no real evidence that this brain area is more active when memory recall occurs.  Purely by chance we would expect one such hundredth or thousandth of the brain to show a hundredth or a two-hundredth more activity at any particular moment, purely because of chance variations. 

There are several conventions or tendencies in the world of experimental neuroscience that aid and abet Joe in this particular piece of junk science he is producing. 

(1) A lack of pre-registration. Pre-registration is when a scientist commits himself to testing one exact hypothesis, and also spells out exactly how data will be gathered and analyzed, before any data is gathered. It is generally recognized that pre-registration greatly reduces the amount of junk science. If, for example, Joe had to pick some particular hundredth of the brain for analysis, and test the hypothesis that this particular hundredth of the brain is more active during memory recall, it would be much less likely that Joe would produce a false alarm result coming from mere correlation fishing.  But, sadly, pre-registration is rarely practiced in experimental neuroscience. Once some neuroscientist has gathered data, he is free to check for correlations in 1001 different places, using a hundred and one different analysis pipelines, each a different way of analyzing the data. It will therefore be easy for some spurious correlation to be found, one that is not a real sign of a causal relation. 

(2) Small study group sizes, and failing tests put in file drawers (no guaranteed publication). If you limit yourself to small study group sizes, it's always easier to find spurious correlations that do not involve causal relations. For example, let's suppose you are looking for a correlation between birth year and death year, or birth month and death month, or birth day of week and death day or week, or any of those correlations occurring between a father and a son or a mother and a son. You won't find any such thing examining data on the lives of 1000 children and their parents, but it wouldn't be too hard to find such a correlation if you merely need to show such a correlation within a group of ten people. You could try using data from 10 or 20 people, and if you don't find anything, you could just keep trying, using a different set of 10 or 20 people. Before long you would be able to report finding such a correlation, even though it involves no causal relation. 

Sadly, this state of affairs matches what goes on in experimental neuroscience. It is very, very common for neuroscientists to publish papers based on small study group sizes such as those involving only 10 or 15 subjects. Also, very few studies occur as "registered reports" in which there is guaranteed publication. Neuroscientists are aware of what is called publication bias, meaning the tendency of journals to not publish results that are negative. So suppose someone like neuroscientist Joe can find no correlation between brain activity and recall activity, after testing with 10 subjects. He can just "file drawer" his study, and start over using another 10 subjects.  He can keep doing this several times, until he has some marginal correlation to report, reporting on only the subjects in his most recent iteration of his experiment.

(3) Weak variations in brain activity regarded as adequate evidence.  Neuroscientist Joe couldn't get away with his shady correlation fishing if there were standards such as a standard of requiring a 1% difference in brain activity as evidence of correlation.  But there are no such standards or conventions.  Many a neuroscientist has published papers reporting correlations involving differences in brain activity of no more than 1 part in 200 or 1 part in 500. Such reports are almost all false alarms. 

(4) Easy-to-obtain ".05" statistical significance as a standard for publication.  Somehow there arose in experimental science a convention that if you could show something has a "statistical significance" of .05, then that's good enough for publication in a science paper.  This is a very loose and weak standard that is all too easy to reach. Roughly speaking, with such a standard, anything is regarded as "statistically significant" as long as you would expect it to show up by chance only 1 time in 20 or less.  But when neuroscientists have not committed themselves (by pre-registration) to testing one particular hypothesis in one particular way, they are free to try 101 different ways to analyze their data looking for correlations. It's all too easy to find something that can be reported as "statistically significant." Even if you fail after fifty or a hundred attempts, you can just "file drawer" your study, start over, and you'll probably be able to report some "statistically significant" correlation on Version 2 or Version 3 of your experiment. In his book "The Cult of Statistical Significance," economist Stephen Thomas Zilliac laments this habit of judging science papers as being acceptable if they reach .05 statistical significance.  He says that science took a giant wrong turn by adopting such a convention.  He says, "Statistical significance is neither necessary nor sufficient for a scientific result." 

torture the data until it confesses
This can be done when there's no pre-registration

Underpowered studies are those with a statistical power of under 50%. A scientific paper says this about the practice of accepting p=0.05 as a "good enough" mark for publication of science papers:

"If you use p=0.05 to suggest that you have made a discovery, you will be wrong at least 30% of the time. If, as is often the case, experiments are underpowered, you will be wrong most of the time....Button et al. [9] said, 'We optimistically estimate the median statistical power of studies in the neuroscience field to be between about 8% and about 31%' ".

(5) Phony visuals are allowed in correlation-fishing papers. It should be a standard in experimental neuroscience that any paper will be rejected if it misleadingly has a visual giving people the impression that a difference in brain activity was greater than it was. Unfortunately, no such standard exists. Correlation-fishing studies routinely include "lying with colors" visuals that dishonestly depict differences of only 1 part in 200 or 1 part in 500, making them look like much greater differences such as 1 part in 10. See here for how this type of visual deception occurs. 

At the end of the long quote above was the sentence, "As Konrad Kording has admitted, practitioners get around the non-replication problem simply by avoiding doing replications." That's right. As scientist Randall J. Ellis stated in 2022, "There are no completed large-scale replication projects focused on neuroscience."

Involving only very slight reported differences in brain activity, correlation-fishing experimental neuroscience studies do nothing to show that the brain is the source of the human mind or the storage place of human memories. Search for the phrase "percent signal change" and you will find that almost all of such studies are reporting changes of only about 1 part in 200 or smaller. In the rare case when a signal change of 1% is reported, it is usually because of head movements, which are a large source of false alarms in such studies. A person being brain scanned and not perfectly following instructions to keep his head motionless can cause a brain scan blip that is reported as a region activation.  A paper states, "The signal derived from functional MRI (fMRI) can also be greatly perturbed by motion; two detailed reports by Power and colleagues describe the complex and variable manner by which different types of motion can impact fMRI acquisitions and increase the proportion of spurious correlations across the brain." Another paper says this:

"A general relationship between head motion and changes in BOLD signal across the brain can be seen in every subject examined in this paper (N=119 in four cohorts)....Any and all movement tends to increase the amplitude of rs-fcMRI signal changes."

Scientists use a variety of types of "motion scrubbing" to try and get rid of the effects of head movements during brain scans, and there is no standard for such data massaging; it's "roll your own." The paper notes that "motion scrubbing tends to decrease many short-range correlations, and to increase many medium- to long-range correlations. " We can imagine how it is for some scientist fishing for correlations between mental activity and brain activity in brain scans. If he is dissatisfied with the size of the correlation reported, he can just tweak his "motion scrubbing" technique to easily get some more correlation that can be reported.