Showing posts with label bad computer programming in neuroscience studies. Show all posts
Showing posts with label bad computer programming in neuroscience studies. Show all posts

Friday, March 21, 2025

Claimed Evidence for "Concept Cells" Is Just Noise-Mining Nonsense

 Quanta Magazine is a widely-read online magazine with slick graphics. On topics of science the magazine again and again is guilty of the most glaring failures.  The articles at Quanta Magazine often contain misleading prose, groundless boasts or the most glaring falsehoods. I discuss some examples of such poor journalism in my posts here and  here and here and here and here.

The latest piece of nonsense in Quanta Magazine is an article trying to persuade us that scientists have discovered "concept cells" in the brain. No such thing has occurred. What is mainly going on is noise mining,  the morally dubious exploitation of very sick epilepsy patients, and scientists trying to get citations and attention by wrongly applying unjustified nicknames to cells. 

The research discussed works like this:

(1) Electrodes are implanted in the brains of very sick epilepsy patients requiring surgery for epilepsy, supposedly for the sake of surgical evaluation on where to perform the surgery (although we should suspect that additional electrodes are being implanted so that this type of noise-mining research can be done). 

(2) The patients are shown some visual stimuli, and EEG readings of their brain waves are taken while they see this stimuli. 

(3) The data is then analyzed, with searches made for particular neurons that fired more often when some particular image was seen.  Scientists then make a triumphal declaration that a "concept cell" was found, on the basis of the claim that some neuron was firing more frequently than we would expect when some visual stimulus was seen. 

This is noise-mining, like someone searching 1000 photos of clouds looking for one that looks like the ghost of an animal.  Neurons in the brain fire continuously, at a rate between 1 time per second and 200 times per second. Anyone tracking the firing of 300 neurons while someone is seeing different things will be able to find a few neurons that seemed to fire more often when something was seen. Similarly, if someone records the ups and downs of 300 stocks on the New York Stock exchange, and tries to correlate them with the occurrence of images coming from his television set showing an old movie, he will be able to find a few stocks that went up or down more often when some particular image was displayed. But that would be mere noise-mining. 

It is never justified to speak of single-neuron "responses" to a concept or an experience.  A single neuron does not respond to something a  person sees or recalls or thinks about. A neuron fires continuously at a rate of about 1 time per second or more, with random variations. It is always misleading to try and suggest a stimulus and response relation between a neuron firing and something someone saw or thought of or recalled.  This is like tracking many flu-infected people who each cough hundreds of times a day, and boasting about having found some "concept coughs," claiming that some of the coughs are a "response" to an image a person sees on a TV.  

The main paper discussed is the late 2024 paper "Concept and location neurons in the human brain provide the 'what' and 'where' in memory formation." The paper does nothing to find any evidence of memories stored in brains. All that is going is noise-mining of 3681 neuron firings in 13 epilepsy patients who had electrodes implanted in their brains. 

We should have no trust in the statistical analysis done in the paper, which is largely dependent upon a large body of programming code that looks like it is poorly written, and is performing all kinds of obscure or arbitrary convolutions and manipulations of data. You can see the programming code here. The EEG data is being processed in many a strange way, and is being passed through all kinds of contortion processes including many  doubly-nested loops doing God-only-knows-what, as in this example:

for k=1:numXbins

    for j=1:numYbins

        n(k,j)= sum ( wavs(:,k) <= ybins(j)+ybinSize/2 & wavs(:,k) > ybins(j)-ybinSize/2);

    end

end

No peer reviewer could ever untangle the programmatic "witches' brew" that is going on in the programming code of this paper. 

torturing data until it confesses

 A look at the peer reviewer comments on the paper gives us some hints about the paper's defects. One peer-reviewer asks this:

"These recordings come from epilepsy patients. How were possible epilepsy-related confounds mitigated?"

Here's what this comment refers to: epilepsy patients scheduled for surgery have all kinds of weird brain wave anomalies cropping up in their brain waves as recorded by EEG devices.  The possibilities for getting false alarms from EEG readings from epilepsy patients is endless. The paper authors respond to this question in an unconvincing way, by claiming that they were using EEG analysis software that did something to reduce such a problem. It is not a convincing response. 

When we read calculations of p-values in papers like this, we typically get no detailed discussion of how such a p-value was computed. We should have little trust in the accuracy of the calculated p-value. In the case of this paper, one of the peer reviewers says that he did his own calculation, and got a p-value number drastically different from  one of the p-values published in the paper.  

brain wave noise mining
The game of "keep torturing the data until it confesses"

This experiment was based on nonsensical assumptions.  The brain has billions of neurons and trillions of synapses. There never was any reason to believe that studying the firing rate of any single neuron  during the observation of sights by subjects would produce any evidence that such a neuron encodes or recognizes or represents or is sensitive to any concept. The idea that you would get meaningful evidence of such a thing from analyzing the firing of only 3681  randomly selected neurons (from a brain of billions of neurons) never made any sense, particularly given the extremely large random variations in the firings of noisy neurons. If an individual neuron encoded a concept, we would expect that an experiment such as this would have less than 1 chance in a million of success, since 3681 is less than a millionth of the total number of neurons in a brain. 

In any study in which subjects are moving muscles, you can never assume that some neuron has some connection to a concept because greater firing activity occurred when that concept was displayed or chosen.  It is well-known that muscle movements abundantly contaminate EEG readings that indicate how much neurons are firing. So in any study in which subjects are not perfectly immobile, you have no way of knowing whether some increased neuron firing is merely due to some increase or difference in muscle movement. In this study subjects were not perfectly immobile when the EEG readings occurred. They were instead performing activities with their hands. For example, we read, "The participant was asked to confirm every image location by tapping it within the presentation time window (1.5–3.5 s)."

The idea of a concept cell is nonsense. No one has any coherent idea of how a single cell could represent a concept or why a cell would have a tendency to respond more frequently when a viewer is exposed to just one concept. Since the  paper is so dependent upon black-box spaghetti code that is programmatically fooling around with their gathered data in many strange and tangled ways, no one should have any confidence that evidence of a "concept cell" was found. 

Below is a visual (Figure 2) from a scientific paper that did single-cell recordings of the firings of individual neurons in monkeys. Each one of the vertical bars represents a spike in the firing rate of a neuron. There is no clear relation between the firing rate and the stimulus presented to the monkey,  The "clumpy-bursty" example was one cherry-picked to show the strongest evidence of response to the stimulus.  A visual like this makes clear that spikes or blips in the firing rate of a neuron occur randomly many times a day. It is deceptive to pick out a case where one of these blips occurs when the stimulus occurred, and to call such a blip a "response" to the stimulus. But such a deception is what is occurring in papers claiming to have found "concept cells." 

blips in neuron firing rates

The paper discussed above ("Concept and location neurons in the human brain provide the 'what' and 'where' in memory formation")  gives us no assertion that the microelectrodes implanted in the brains of these very sick epilepsy patients were inserted only for medical reasons, to evaluate where they should have surgery. Since we have got no such assertion, we should suspect that one or more of the microelectrodes were unnecessarily implanted in the brains of very sick patients, for the sake of this poor-quality study based on nonsensical assumptions. Implantation of microelectrodes in the brain comes with very serious risks.  People requiring epilepsy surgery are very sick people who should not be put at higher risk for the sake of low-quality research such as this. 

The paper discussed above implanted many microelectrodes in epilepsy patients. The type of electrode normally used for surgical evaluation of epilepsy patients is a much larger type of electrode called a macroelectrode. A scientific paper tells us, "Sixty-five years after single units were first recorded in the human brain, there remain no established clinical indications for microelectrode recordings in the presurgical evaluation of patients with epilepsy (Cash and Hochberg, 2015)." In other words, there is no medical justification for implanting microelectrodes in the brains of epilepsy patients. The paper tells us that the microelectrodes were "inserted through the hollow clinical macro electrodes, and protruding from the tips by ~4 mm." This was medically unnecessary and potentially hazardous insertion of many wires into the brains of very sick patients.  We seem to have here some very sick patients being put at needless risk merely so that junk science can be produced. 

reckless neuroscientist

Here is a quote from a scientific paper:

"The effects of penetrating microelectrode implantation on brain tissues according to the literature data...  are as follows:

  1. Disruption of the blood–brain barrier (BBB);
  2. Tissue deformation;
  3. Scarring of the brain tissue around the implant, i.e., gliosis 
  4. Chronic inflammation after microelectrode implantation;
  5. Neuronal cells loss."
I strongly advise any people who participated in any brain scanning experiment or any neuroscience experiment involving electrode implants to permanently keep very careful records of their participation, to find out and write down the name of the scientific paper corresponding to the study, to write down and keep the names of any scientists or helpers they were involved with, to permanently keep a copy of any forms they signed, and to keep a very careful log of any health problems they ever have. Such information may be useful should such a person decide to file a lawsuit or a claim seeking monetary damages. 

Postscript: The term "gnostic cells" has sometimes been used to mean the same thing as "concept cells."

More noise-mining nonsense is found in the recent paper "Lack of context modulation in human single neuron responses in the medial temporal lobe," which you can read here. This time the study group is even smaller, consisting of only 9 very sick patients.  The authors have failed to make their code publicly available through any easy access method, as if they were embarrassed by their programming. But using the link here you can download their code as a .zip file, scan the .zip file for viruses, and then extract it, which is a laborious way to have to inspect code. I did that, and found the usual very-poorly-documented spaghetti code programming horror that you typically find in projects like this. We have undocumented loops like the one below, doing some unfathomable rigmarole contortions of the brain wave data:

for iresp=1:size(resp,2)
    if ~isempty(resp(iresp).trecall_phasic_ms)
        resp_rec = resp_rec + 1;
        active_cluster = resp(iresp).spike_times_Rec;
        phasic_recall = resp(iresp).trecall_phasic_ms;
        ntrials(1) = size(active_cluster{resp(iresp).responsive_Storiesindex(1)},1);
        ntrials(2) = size(active_cluster{resp(iresp).responsive_Storiesindex(2)},1);
        strength_pair = NaN*ones(max(ntrials),2);
    
        equiv_strength(resp_rec).chan = resp(iresp).channel_number;
        equiv_strength(resp_rec).class = resp(iresp).cluster;
        equiv_strength(resp_rec).pair = resp(iresp).responsive_Storiesindex;
            
        for istim=1:2 
            phas_vec = phasic_recall{resp(iresp).responsive_Storiesindex(istim)};
            spikes1 = active_cluster{resp(iresp).responsive_Storiesindex(istim)};
            spikes1 = arrayfun(@(k) spikes1{k}-phas_vec(k),[1:length(phas_vec)]','UniformOutput',false);
            strength_pair(1:ntrials(istim),istim) = cell2mat(cellfun(@(x) sum((x< twin_off) & (x> twin_on)),spikes1,'UniformOutput',0));
        end
            
    
        nsamp = min(ntrials);
        equiv_strength(resp_rec).samples = ntrials;
        equiv_strength(resp_rec).meandiff_Hz = (diff(nanmean(strength_pair)))/((twin_off-twin_on)/1000);
        equiv_strength(resp_rec).delta = sqrt(2/nsamp*(norminv(alpha)+norminv(alpha/2))^2);
        [equiv_strength(resp_rec).test_resu, equiv_strength(resp_rec).pval, equiv_strength(resp_rec).pooledSD, equiv_strength(resp_rec).meandiff] = TOST_2023(strength_pair(~isnan(strength_pair(:,1)),1), strength_pair(~isnan(strength_pair(:,2)),2), 'welch',equiv_strength(resp_rec).delta,alpha);
    end
end

And we have the equally ugly unfathomable bit of monkey business below:

for ii=1:length(spikes)
        if iscell(spikes{ii})
            all_spks=cell2mat(spikes{ii}');
        else
            all_spks=spikes{ii}(:);
        end

        spikes_tot=all_spks(all_spks < tmax_epoch+half_ancho_gauss & all_spks > tmin_epoch-half_ancho_gauss);
        spike_timeline = hist(spikes_tot,(tmin_epoch-half_ancho_gauss:sample_period:tmax_epoch+half_ancho_gauss))/ntrials(ii);
        n_spike_timeline = length(spike_timeline); %should be the same length as ejex
        integ_timeline_stim = conv(spike_timeline, int_window);
        integ_timeline_stim_cut = integ_timeline_stim(round(half_ancho_gauss/sample_period)+1:n_spike_timeline+round(half_ancho_gauss/sample_period));
        aver_fr{ii} =  smooth(integ_timeline_stim_cut(which_times),smooth_bin);

        % subplot(length(spikes),1,ii)
        % plot(ejex(which_times(1:downs:end)),aver_fr{ii}(1:downs:end))
        % maxi(ii) = max(aver_fr{ii});
    end  

I can merely say the more time a programmer spends looking at the programming code used by this paper, the less confidence he will have that the paper did anything to establish the existence of "concept cells." This is black box "witches' brew" monkeying around with brain wave data that is best described with the phrase "they kept torturing the data until it confessed in the weakest whisper."

A visual from the paper gives you the type of data that was being tortured to try to get something. Random cases of more noise blips from noisy and extremely variable neurons (during a few seconds) are being passed off as "concept responses." It's like someone tracking each and every noise blip from a nearby street construction crew using a jackhammer,  and trying to correlate particular noise spikes with images appearing on his TV set. That would be a very silly case of noise mining, and what is going on in this paper is just as silly. 


One of the authors of this poor-quality paper is Rodrigo Quian Quiroga, who has long been quoted as making a misleading claim about a "Jennifer Aniston" concept cell.  For example:
  • In a 2017 article "Concept cells: the building blocks of declarative memory functions" by  Quian Quiroga, he incorrectly stated, "One of the first such neurons found in the hippocampus fired to seven different pictures of the actress Jennifer Aniston and not to 80 other pictures of known and unknown people, animals and places." This was a claim that did not match the data in his 2005 paper where we have in Figure 1A a visual depiction of more than 50 firings of that neuron when the subject was shown a picture other than Jennifer Aniston.
  • In Figure 1 of the year 2020 "Searching for the neural correlates of human intelligence" article by Quian Quiroga, he misleadingly states that the Jennifer Aniston neuron "did not respond to about 80 pictures of other persons," a claim which does not match the data in his 2005 paper where we have in Figure 1A a visual depiction of more than 50 firings of that neuron when the subject was shown a picture other than Jennifer Aniston.  
  • In a 2025 Salon article with many misstatements,  Quian Quiroga made this incorrect statement: "'Twenty years ago … I was doing experiments with a patient, and then I showed many pictures of Jennifer Aniston, and I found a neuron that responded only to her and to nothing else...It was very clear that in an area called the hippocampus that is known to be critical for memory, we have neurons that represent, in this case, specific people, or in general, specific concepts. "  The claim does not match what was reported in Quian Quiroga's 2005 paper, where  we have in Figure 1A a visual depiction of more than 50 firings of that neuron when the subject was shown a picture other than Jennifer Aniston.  
 In fact, Figure 1A of the 2005 paper shows that the famed "Jennifer Aniston neuron" did fire more than seven times when a picture of a basketball player was shown, that it did fire more than seven times when a picture of another basketball player was shown, that it did fire more than five times when a picture of a snake was shown, and that it did fire more than six times when a picture of the Leaning Tower of Pisa was shown. The paper says this: "To hold their attention, patients had to perform a simple task during all sessions (indicating with a key press whether a human face was present in the image)."  It is known that the muscle movements can abundantly affect EEG readings and whether a neuron fires. A difference between key presses when the Jennifer Aniston picture was shown and things other than faces were shown can easily account for why one neuron may have fired differently when pictures of Jennifer were shown, without any need at all to evoke an idea of a "concept cell" for Jennifer Aniston. Maybe Quian Quiroga's misstatements on his "Jennifer Aniston neuron" had something to do with the fact that he later managed to get a book deal, a deal for a book with a title mentioning his claimed "Jennifer Aniston" neuron. 

When he discusses this "Jennifer Aniston neuron" outside of the original paper, Quian Quiroga seems to never mention that the subject with that neuron was asked to press a button whenever he saw a face, and that muscle movements are known to increase brain wave spikes picked up by electrodes or EEG devices.  That seems like the "secret sauce" behind his "Jennifer Aniston" neuron. 

There is nothing very impressive about this case of the neuron's firing more often while someone saw a few pictures of Jennifer Aniston. While unlikely to occur on any one day, it is the kind of result you would expect an eager noise miner to produce after spending very long periods of time looking for some result better than chance in a dataset of random data. Similarly, if someone is very eager to find a cloud shape looking like the ghost of an animal, and he spends day after day scanning photos of clouds looking for such a thing, he will probably find one or two clouds that look rather like animal ghosts. Activity like this is correctly described as misguided noise-mining "Jesus in my toast" pareidolia, aided by misstatements in which the meager result is described in a misleading way. 

Thursday, February 27, 2025

Programming Gone Astray: Iteration Inanity of the Neuroscientists' Distortion Loops

 Quanta Magazine is a widely-read online magazine with slick graphics. On topics of science the magazine again and again is guilty of the most glaring failures. Quanta Magazine often assigns its online articles about great biology mysteries (involving riddles a thousand miles over the heads of PhDs) to writers who lack even a bachelor's degree in biology. Often it will assign such articles to be written by people identified as "writing interns."  The articles at Quanta Magazine often contain misleading prose, groundless boasts or the most glaring falsehoods. I discuss some examples of such poor journalism in my posts here and here and here and here

The writers at Quanta Magazine are very often guilty of bootlicking, a word meaning excessive deference to an authority or a superior.  The latest example of bootlicking at the magazine is an article entitled "How 'Event Scripts’ Structure Our Personal Memories." The subtitle makes this very untrue claim: "By screening films in a brain scanner, neuroscientists discovered a rich library of neural scripts — from a trip through an airport to a marriage proposal — that form scaffolds for memories of our experience."  The claim has no basis in fact. The article follows it with this equally untrue claim: " 'Event scripts' are distinct neural fingerprints that encode repeated sequences of events, such those that unfold during a trip through the airport."  No such things have been found. 

The article begins telling us tall tales about neuroscientist Christopher Baldassano, incorrectly stating this: "Then, in 2018, Baldassano found it: neural fingerprints of narrative experience, derived from brain scans, that replay sequentially during standard life events. " No such thing happened. The article is referring to a very low-quality paper co-authored by Baldassano, one entitled "Representation of Real-World Event Schemas during Narrative Perception." 

The study had the following flaws:

(1) The study group sizes in this task-based fMRI study were skimpy, consisting of only 15 or 16 subjects per study group. Referring to study group sizes twice as large, an article on neursosciencenews.com states this: "A new analysis reveals that task-based fMRI experiments involving typical sample sizes of about 30 participants are only modestly replicable. This means that independent efforts to repeat the experiments are as likely to challenge as to confirm the original results."

(2) No blinding protocol was used. 

(3) The paper was not preregistered, and did not test any hypothesis formulated before gathering data, using a method specified before gathering data. 

(4) The paper is a bad example of "keep torturing the data until it confesses" methodology.  The paper has graphs that are not based on simple brain scans, but are instead based on brain scan data after it has been manipulated through the most convoluted pathway of arbitrary contortions. 

Below from the paper is a discussion of only a small fraction of the "keep torturing the data until it confesses" nonsense that was occurring:

"For each story, four regressors were created to model the response to the four schematic events, along with an additional nuisance regressor to model the initial countdown video. These were created by taking the blocks of time corresponding to these five segments and then convolving with the HRF from AFNI (Cox, 1996). A separate linear regression was performed to fit the average response of each group (in the 100-dimensional SRM space) using the regressors, resulting in a 100-dimensional pattern of coefficients for each event of each story in each group. For every pair of stories, the pattern vectors for each of their corresponding events were correlated across groups (event 1 from Group 1 with event 1 from Group 2, event 2 from Group 1 with event 2 from Group 2, etc., as shown in Fig. 2a) and the four resulting correlations were averaged. This yielded a 16 X 16 matrix of across-group story event similarity. To ensure robustness, the whole process was repeated for 10 random splits of the 31 subjects, and the resulting similarity matrices were averaged across splits...To explore the dimensionality of the schematic patterns, we reran the analysis after preprocessing the data with a range of different SRM dimensions, from 2 to 100. The resulting curve of z values versus dimensionality for each region was then smoothed with the LOWESS (Locally Weighted Scatterplot Smoothing) algorithm implemented in the statsmodels python package (using the default parameters). To generate the searchlight map, a z value was computed for each vertex as the average of the z values from all searchlights that included that vertex. The map of z values was then converted into map of q values using the same false discovery rate correction that is used in AFNI (Cox, 1996)....The resampled data (time courses on the left and right hemispheres, and in the subcortical volume) were then read by a custom python script, which implemented the following preprocessing steps: removal of nuisance regressors (the 6 degrees of freedom motion correction estimates, and low-order Legendre drift polynomials up to order [1  duration/150] as in Analysis of Functional NeuroImages [AFNI]) (Cox, 1996), z scoring each run to have zero mean and SD of 1, and dividing the runs into the portions corresponding to each stimulus. All subsequent analyses, described below, were performed using custom python scripts and the Brain Imaging Analysis Kit (http://brainiak. org/)."

This is only a small fraction of the contortion inanity that was going on. The paper has many other paragraphs sounding like the one just quoted.  To see the ugliness of the manipulation muddle that was occurring, you must look at the programming code. The authors have made their code public, and you can see it using the link here.  Looking at their programming scripts, we see an appalling example of arbitrary, unjustifiable  algorithms, the most convoluted spaghetti code.  The brain scan data is being passed through many types of poorly documented programming loops that are doing God-only-knows-what kind of mystifying manipulation. Below is only a tiny part of the bizarre manipulations that were going on.


You might call this "iteration inanity." The output is some kind of utterly artificial "witches' brew" that cannot be called the original data gathered or anything like the original data gathered. We should not have any confidence in any of the main graphs in the paper, because they are all produced by passing brain scan data through spaghetti code convolution contortions similar to the one shown above.  This is a severe example of "keep torturing the data until it confesses," what we might call a Spanish Inquisition level of torturing.  We have some utterly artificial transmogrification mess that is the result of obscure arbitrary  programming manipulations, some gobbledygook rigmarole. The authors have not found any "event scripts" or patterns in the brain.  The only thing they have found is something they have created themselves by spaghetti-code programming that distorts and manipulates the original brain scan data. 

spaghetti code neuroscience

keep torturing data until it confesses
Was this how the mess arose?

When good programmers are writing straightforward programming code and they know what they are doing, they tend to use intelligible variable names such as ThisYearsAccruedInterest or TotalAccruedInterest.  Bad programmers use unintelligible variable names such as "d" or "ev" or "cc" or "np," like in the example above, without any comments documenting the variable names, often because they don't even understand what the variables correspond to, and cannot give an intelligible name corresponding to the variable.  As a general rule, we should tend to distrust any scientific programming that uses undocumented variable names of one or two letters such as "d" or "ev" or "cc" or "np," because the use of such cryptic variable names is a strong reason for suspecting that incompetent programmers are at work. 

The Quanta Magazine article then has a link to another paper by Baldassano and others, one entitled "Top-down attention shifts behavioral and neural event boundaries in narratives with overlapping event script."  The paper relies on the same kind of iterative inanity as the previously mentioned paper. We see the same type of loony-looking loops that make all kind of weird, arbitrary transfigurations and contortions and manipulations of the original data, with only the scarcest comments in the source code to explain what is being done. It's another big heap of spaghetti code nonsense doing God-only-knows-what to the original data.  You can see the manipulation mess by looking at the Python files here

Nothing real about the brain is being revealed here. If any "scripts" or patterns were discovered, the authors were merely discovering the outputs of their own data-manipulating programming loops. To claim the output of such distortion loops as being something in the brain is as misleading as picking up 100 stones from the seashore,  forming them into a sculpture of a cat, and then claiming that the waves produced a sculpture of a cat. 

It is rather obvious that our Quanta Magazine writer has not learned how to distinguish good neuroscience research from very bad neuroscience research. That writer states this:

"In 2004, the neuroscientist Uri Hasson and his colleagues at the Weizmann Institute of Science in Israel started carving a path through the thicket of voxels. In one of their studies, five people, while lying in a brain scanner, watched 30 minutes of The Good, the Bad and the Ugly (1966), a spaghetti western starring Clint Eastwood. Comparing the data from the five participants, the researchers noted when and where brain activity surged or waned in unison."

Why would anyone even bother to mention a research study using so obviously too-small a study group size of only five subjects? The writer then gives us one more bum steer. We are given false claims about a study by a neuroscientist named Chen:

"In 2012, Chen joined Hasson’s lab, then at Princeton, and extended the approach to memory. She had people watch the first episode of the television show Sherlock (2010), featuring Benedict Cumberbatch as a modern take on the legendary detective. Then the study participants talked through their memory of it, while still lying in the scanner. The experiment worked. Chen and her colleagues were able to match brain activity recorded during participants’ recollections to specific scenes around 60 seconds long — for example, when Sherlock meets Watson."

The claim is false, because the study was some very low-quality work. The link given in the Quanta Magazine article is to the paper "Shared memories reveal shared structure in neural activity across individuals" which you can read here. The study had the following defects:

  • The study used too-small study groups such as one with only 8 subjects and another with only 9 subjects. The authors confess, "No statistical methods were used to pre-determine sample sizes but our sample sizes are similar to those reported in previous publications." It is well-known that neuroscience experiments typically use way too few subjects for results with good statistical power, so you do not have a good excuse for failing to do a sample size calculation (to determine a good study group size) by appealing to other experimenters using study group sizes like yours. 
  • The study failed to use a blinding protocol.  The authors confess, "Data collection and analysis were not performed blind to the conditions of the experiments." 
  • Instead of simply using the original brain scan data, the authors performed very many obscure and arbitrary convolutions, contortions and distortions of the original data. 
A very long part of the paper describes all the weird data manipulations and convoluted contortions that were occurring. Here is only a very small fraction of that part:

"We performed a resampling analysis wherein the individual participant correlation values for recall-recall and movie-recall were randomly swapped between conditions to produce two surrogate groups of 17 members each, i.e., each surrogate group contained one value from each of the 17 original participants, but the values were randomly selected to be from the recall-recall comparison or from the between participant movie-recall comparison. These two surrogate groups were compared using a t-test, and the procedure was repeated 100,000 times to produce a null distribution of t values. The veridical t-value was compared to the null distribution to produce a p-value for every voxel. The test was performed for every voxel that showed either significant recall-recall similarity (Fig. 3B) or significant between-participant movie-recall similarity (Fig. 3B), corrected for multiple comparisons across the entire brain using an FDR threshold of q = 0.05 (see Methods: Pattern similarity analyses); voxels p < 0.05 (one-tailed) are plotted on the brain (Fig. 4A)"

The authors have not provided a link to their source code. Based on the descriptions of their methods, we may assume that they were using the same kind of unjustifiable distortion loops that go on in the papers of Baldassano.  People with the worst programming are the least likely to make the code public. The Chen paper "Shared memories reveal shared structure in neural activity across individuals" is low-quality work that fails to follow good standards of research. No robust evidence has been provided of "shared structure in neural activity" when the same memories are experienced.  The authors seem to have merely discovered something they created themselves through their strange contortions and manipulations of data. 

The Quanta Magazine article is a very bad example of bootlicking. We have all kinds of claims that scientists accomplished something, when most of these things were not actually done, because the methods used were so poor.  Instead of such fanboy swooning, the author should have put the methods of the discussed neuroscientists under stringent critical scrutiny, which would have mainly revealed the defective methods being used. 

Part of the problem with studies like this is that we do not get any chronological account of the different attempts at fooling around with the brain scan data that was produced. We only get a result of some final algorithmic result that the authors had, after a long process of "keep torturing the data until it confesses in the weakest whisper." We may presume that what often goes on is something rather like this:

Programmer: Well, that ends my 18th programming attempt to squeeze some "patterns" out of this brain scan data, and I still have nothing. I'm getting nowhere. It's like trying to squeeze blood from a stone. 
Scientist: Keep trying! Be more creative!  Add, pad; slice, dice;  merge, purge; mix, fix; ruffle, shuffle; sift, shift; combine, align; inflate, conflate; shrink, link and sync; crop, drop, and swap; bend, blend, rend and mend; ditch, hitch, stitch and switch. Try every kind of programming loop you can think of, to try to gin up something from this data that we can call a pattern, or something we can claim as a possible representation. 
Programmer: Do I have to save all the earlier versions of my code that failed?
Scientist: Hell no. We only describe the FINAL version of the code in our paper. 

Monday, September 9, 2024

Neuroscientists Senselessly Think They Can Perform Innumerable Contortions of Brain Data, and Then Claim a Discovery

Claims by neuroscientists that they have found "representations" in the brain (other than genetic representations) are examples of what very abundantly exists in biology: groundless achievement legends. There is no robust evidence for any such representations. 

Excluding the genetic information stored in DNA and its genes, there are simply no physical signs of learned information stored in a brain in any kind of organized format that resembles some kind of system of representation. If learned information were stored in a brain, it would tend to have an easily detected hallmark: the hallmark of token repetition.  There would be some system of tokens, each of which would represent something, perhaps a sound or a color pixel or a letter. There would be very many repetitions of different types of symbolic tokens.   Some examples of tokens are given below. Other examples of tokens include nucleotide base pairs (which in particular combinations of 3 base pairs represent particular amino acids), and also coins and bills (some particular combination of coins and bills can represent some particular amount of wealth). 

symbolic tokens

Other than the nucleotide base pair triple combinations that represent mere low-level chemical information such as amino acids, something found in neurons and many other types of cells outside of the brain, there is no sign at all of any repetition of symbolic tokens in the brain. Except for genetic information which is merely low-level chemical information, we can find none of the hallmarks of symbolic information (the repetition of symbolic tokens) inside the brain. No one has ever found anything that looks like traces or remnants of learned information by studying brain tissue. If you cut off some piece of brain tissue when someone dies, and place it under the most powerful electron microscope, you will never find any evidence that such tissue stored information learned during a lifetime, and you will never be able to figure out what a person learned from studying such tissue.  This is one reason why scientists and law enforcement officials never bother to preserve the brains of dead people in hopes of learning something about what such people experienced during their lives, or what they thought or believed, or what deeds they committed.    

But despite their complete failure to find any robust evidence of non-genetic representations in the brain, neuroscientists often make groundless boasts of having discovered representations. What is going on is pareidolia, people reporting seeing something that is not there, after wishfully analyzing large amounts of ambiguous and hazy data. It's like someone eagerly analyzing his toast every day for years, looking for something that looks like the face of Jesus, and eventually reporting he saw something that looked to him like the face of Jesus.  It's also like someone walking in many different forests, eagerly looking for face shapes on trees, and occasionally reporting a success, or like someone scanning the sky, looking for clouds that look like animal shapes.

pareidolia

The latest example of nonsensical neuroscientist pareidolia is to be found in a press release from Columbia University, and the paper that press release describes in very misleading terms.  The press release has the phony headline "
Scientists Capture Clearest Glimpse of How Brain Cells Embody Thought." When you read a headline like that, you should remember a sad truth that has been glaringly obvious for many years now: the fact that university press releases on topics of science are no more trustworthy than corporate PR press releases.  There is the most gigantic amount of lying, hype and misrepresentation in university press releases these days, and such baloney occurs in equal amounts in the press releases of every major university. So don't think for a moment than you can  trust a press release because it came from Harvard or Columbia or Yale or Oxford University. I wish I had a dollar for every bogus press release that has been issued by such institutions. The subtitle of the press release is the 100% untrue claim "Recordings from thousands of neurons reveal how a person’s brain abstractly represents acts of reasoning."

We have an utterly groundless claim by a neuroscientist that he and his colleagues found a “uniquely revealing dataset that is letting us for the first time monitor how the brain’s cells represent a learning process critical for inferential reasoning." We then have an equally groundless claim by another neuroscientist that "this work elucidates a neural basis for conceptual knowledge, which is essential for reasoning, making inferences, planning and even regulating emotions.”

Do the authors claim to have seen some structure in the brain corresponding to these claims? Certainly not. They did not do any brain imaging such as using MRI scans. All they had to work with is EEG readings, readings of brain waves. Would anyone have seen any sign of such claimed representations by visually examining the wavy lines of these EEG readings? Certainly not. 

The press release reveals that what is going on is an affair that can be described as "keep torturing the data until it confesses in the weakest voice." We read this:

"The researchers recast the volunteers’ brain activity into geometric representations – into shapes, that is – albeit ones occupying thousands of dimensions instead of the familiar three dimensions that we routinely visualize. 'These are high-dimensional geometrical shapes that we cannot imagine or visualize on a computer monitor,' said Dr. Fusi. 'But we can use mathematical techniques to visualize much simplified renditions of them in 3D.' ” 

What a joke. They didn't see any such representations of knowledge or thought by a simple examination of the brain wave data they acquired. So they kept fiddling with their data and manipulating the data and contorting the data with some kind of absurdly convoluted and byzantine analysis pathway, and then claimed to see representations or shapes in such super-manipulated data. It's like someone taking 1000 pictures of the clouds in the sky, and then playing around all day with image manipulation filters until he got something that looked like an animal shape in one of the clouds. 

The paper being described (which you can read here) describes the huge chain of arbitrary manipulations and arbitrary naming that went on. It's a laughably arbitrary and byzantine analysis pathway in which many dozens of arbitrary analysis decisions are being made. 

spaghetti code neuroscience analysis

Here from the paper is a description of just a small fraction of the "keep torturing the data until it confessed" spaghetti code craziness that went on:

" A cross-session-group PS was then computed by applying the same alignment to a pair of held-out conditions, one on either side of the current dichotomy boundary. Alignment and cross-group comparisons were performed in a space derived using dimensionality reduction (six dimensions). For a given dichotomy, two groups of sessions with N and M neurons were aligned by applying singular value decomposition to the firing-rate normalized condition averages of all but two of the eight task conditions, one on either side of the dichotomy boundary. The top six singular vectors corresponding to the non-zero singular values from each session group were then used as projection matrices to embed the condition averages from each session group in a six-dimensional space. Alignment between the two groups of sessions, in the six-dimensional space, was then performed by computing the average coding vector crossing the dichotomy boundary for each session group, with the vector difference between these two coding vectors defining the ‘transformation’ between the two embedding spaces. To compare whether coding directions generalize between the two groups of sessions, we then used the data from the two remaining held-out conditions (in both session groups). We first projected these data points into the same six-dimensional embedding spaces and computed the coding vectors between the two in each embedding space. We then applied the transformation vector to the coding vector in the first embedding space, thereby transforming it into the coordinate system of the second session groups. Within the second session group embedding space, we then computed the cosine similarity between the transformed coding vector from the first session group and the coding vector from the second session group to examine whether the two were parallel (if so, the coding vectors generalize). We repeated this procedure for each of the other three pairs of conditions being the held-out pair, thereby estimating the vector transformation of each pair of conditions independently. The average cosine similarity was then computed over the held-out pairs. All possible configurations of conditions aligned on either side of the dichotomy boundary are considered (24 in this case), and the maximum cosine similarity over configurations is returned as the PS for that dichotomy (plotted as ‘cross-half’ in Extended Data Fig. 3z)."

This is only a very small fraction of the "manipulate the data like crazy" nonsense that was going on. The paper lists a total amount of statistical rigmarole that seems ten times more complicated than what is described in the quote above. Never be impressed when you read about such operations. The more complicated such paragraphs are, the more it shows that the original data did not have the final result claimed, and that the authors had to play "keep torturing the data until it confesses" games, using a long series of arbitrary data manipulations  and data contortions to try to "gin up some success."  Strangely, we read almost nothing in the way of justifying these bizarre data manipulations and data contortions. It is as if the authors thought they had the right to dream up the most enormously  convoluted data manipulation scheme, without justifying the bizarre data distortions they were doing. 

A look at some of the programming code used shows that all the data was being passed through doubly-nested loops that were doing God-only-knows-what:

%for every area in the dataset
for i = 1:length(cell_groups)
    
    %run analysis for the inference absent sessions
    idx_current = intersect(cell_groups{i},find(ismember(sessions,inference_absent)));

    for i_rs = 1:n_resample
        [avg_array,~] = construct_regressors(neu,n_samples(i),idx_current);
    
        [t_1,t_2]        = sd(avg_array,n_perm_inner,n_samples(i));
        sd_{i,1}         = cat(2,sd_{i,1},t_1);
        sd_boot{i,1}     = cat(2,sd_boot{i,1},t_2);
        [ccgp_{i,1}]     = cat(2,ccgp_{i,1},ccgp(avg_array,...
                           n_perm_inner,false,n_samples(i)));
        [ccgp_boot{i,1}] = cat(2,ccgp_boot{i,1},ccgp(avg_array,...
                           n_perm_inner,true,n_samples(i)));
        ps_{i,1}         = cat(2,ps_{i,1},ps(avg_array,...
                           n_perm_inner,false));
        ps_boot{i,1}     = cat(2,ps_boot{i,1},ps(avg_array,...
                           n_perm_inner,true));
    end
    
    %run again for the inference present sessions
    idx_current = intersect(cell_groups{i},find(ismember(sessions,inference_present)));

    for i_rs = 1:n_resample
        [avg_array,~] = construct_regressors(neu,n_samples(i),idx_current);
    
        [t_1,t_2]        = sd(avg_array,n_perm_inner,n_samples(i));
        sd_{i,2}         = cat(2,sd_{i,2},t_1);
        sd_boot{i,2}     = cat(2,sd_boot{i,2},t_2);
        [ccgp_{i,2}]     = cat(2,ccgp_{i,2},ccgp(avg_array,...
                           n_perm_inner,false,n_samples(i)));
        [ccgp_boot{i,2}] = cat(2,ccgp_boot{i,2},ccgp(avg_array,...
                           n_perm_inner,true,n_samples(i)));
        ps_{i,2}         = cat(2,ps_{i,2},ps(avg_array,...
                           n_perm_inner,false));
        ps_boot{i,2}     = cat(2,ps_boot{i,2},ps(avg_array,...
                           n_perm_inner,true));
    end

end

Every single piece of data is being passed into a function called construct_regressors(), but what is that function doing? We cannot tell, because the code for that function has not been supplied.  We should be suspicious that this construct_regressors() function was doing something so arbitrary and convoluted that the authors were embarrassed to publish the code for that function. 

What we have here is something like the situation described in the visual below:

What we have here is something like the situation described in the visual below:

keep torturing the data till it confesses

To pass off the results of so vast an amount of data monkeying as a discovery is a case of BS and baloney. No representations in the brain or brain waves have been discovered here. All we have is scientists manipulating and contorting data like crazy, and then displaying some pareidolia by passing off their super-manipulated data as an example of "representations."  

Can you imagine what a scandal would arise if climate scientists tried to get away with even one tenth of this amount of manipulating and contorting and distorting their data? Skeptics of their work would start "screaming bloody murder," and howl about how scientists were failing to use their original data, and using instead manipulated, contorted, twisted, distorted data.  But it seems that neuroscientists senselessly think that it is okay for them to play around endlessly with data from brains,  and that they have the right to contort and twist and distort such data in dozens of different ways, and then pass off the result (an utterly artificial construction) as something they can then call "what the brain does."  

What can you call data like this which has undergone so many contortions and manipulations and distortions and transfigurations that it is something almost totally different from the raw data originally gathered? You might be tempted to call it "fake data," but that isn't quite right, because the authors have described all the transformations they did of the data. So rather than calling it "fake data," a description that is not quite right, we can merely say that it is data that has been so enormously contorted and manipulated and transfigured that it is data that cannot be claimed as evidence telling us about brain states or brain activity. 

Experimental neuroscience is in a state of great sickness and dysfunction. The use of Questionable Research Practices seems like more the rule in experimental neuroscience than the exception. When scientists think that it is okay for them to perform endless manipulations and contortions and distortions and transfigurations of their data, and to then pass off the resulting artificial mess as "what we got from the brain," then it is a sign that neuroscience has sunk to an extremely low nadir of dysfunction. 

Brain waves don't represent anything that someone has learned. Brain waves are no more representations of something learned than clouds are representations of something learned. Brain waves are streams of random data, as random as the stream of clouds passing above a house or  a city.  Not one iota of evidence of brain representations has been presented in this paper. The paper is entitled "Abstract representations emerge in human hippocampal neurons during inference." An honest title for the paper would have been "We got something we called 'abstract representations'  after we manipulated and contorted brain wave readings in dozens of weird ways."

For additional examples of neuroscientists using computer programs to play "keep torturing the data until it confesses," read my post here, entitled "Programming Gone Astray: Iteration Inanity of the Neuroscientists' Distortion Loops."

bad programming by neuroscientists
Contortion craziness