Questionable Research Practices and shoddy methods are extremely abundant in today's neuroscience research. How is it that you can detect such examples of poor research? I will give here a method. The method mainly involves looking for certain types of defects in scientific papers, but also involves looking for defects in press announcements about such papers.
Step #1: open up a "defect list" file
You will be using this file to record any defects you find in either the original scientific paper or any of the press announcements that occur about that paper. You can create such a file by getting a blank piece of paper, opening up a new file using a tool such as Notepad, opening up a file using Google Docs, clicking on the Notes utility on your I-Pad, and so forth.
Step #2: Look for a claim in the headline of an article or press release announcing the research that is not justified by any claims in the text of the article or the text of the academic press release.
An article that you read announcing the research may or may not be the academic press announcing the research. If the press article is not the original academic press release, it may have a link to the academic press release. The academic press release will typically have a link to a newly published scientific paper, and such a link may also be found in some article based on the press release.
An extremely common defect of press articles about scientific papers is that they very often make boastful headline claims that are not justified by anything claimed or established in the articles underneath such headlines. This often occurs for economic reasons, to create the effect known as clickbait. Clickbait is when online articles have sensational-sounding headlines or interesting-sounding headlines that lure you into going to some web page that has ads. The people running or funding such pages thereby get advertising revenue when such pages are viewed.
Clickbait is enormously abundant in the world of science journalism. University press releases very often contain headlines never justified by anything mentioned in the story underneath such headlines. Press articles based on such press releases very often contain headlines never justified by anything mentioned in the story underneath such headlines, or never justified by anything stated in the body of the press release such articles were based on.
A simple starting point in detecting unjustified neuroscience research announcements is to simply compare headlines to the text underlying such headlines, and note cases in which the headline is unjustified. Record any such cases in your "Defects list" file, nothing the URL of the corresponding press article or press release.
I will give a very simple example of such a thing. A headline may announce "Scientists Unlock the Secret of Human Memory Retrieval" But the underlying story may refer to research that only dealt with mice. In such a case the unjustified hype is obvious -- the research told us nothing about human memory.

Step #3: find a copy of the scientific paper that is the basis of the press release or press article.
Generally the press article or press release will have a link to a scientific paper that is the basis of the research announcement. In the easiest case, you will simply be able to click on that link to get the full text of the paper. But in many cases when you click on the link, you will go to a page that merely has the abstract of the paper. There may be some "Full Text" link that asks you to pay money. Such a barrier to you reading the paper is called a paywall. I strongly advise against ever paying money merely to research the quality of a neuroscience paper. Most neuroscience research papers these days are poor quality papers guilty of multiple examples of Questionable Research Practices.
But if you find yourself blocked by a paywall, there are some things you can do to try to get the paper:
(1) Go to the Google Scholar site (https://scholar.google.com/), and copy the name of the paper into the search bar. See whether the paper shows in the search results, with a link to the full text of the paper.
(2) Go to the biology preprint server (https://www.biorxiv.org/) and copy the name of the paper into the search bar. See whether the paper shows in the search results. If it does, you will typically be able to get the full text of the paper by clicking on the Full Text tab on that site.
Step #4: examine the title of the paper and read its abstract, looking for a claim in the title that is not matched by any claim in the abstract
It is surprisingly common these days for the titles of neuroscience research papers to make claims that are not justified by any statements made in either the abstract of the paper or the full text of the paper. If you find any discrepancy between the title of the paper and the results announced in the abstract, record such a discrepancy in your "defects list" file.
Step #5: if the paper is an experimental research file, look for evidence of insufficient study group size
Since the use of way-too-small study group sizes is amazingly predominant these days in experimental neuroscience research, the abstract of every paper should tell how many subjects were used in each of the study groups. But like people who are trying to hide their shortcomings, the abstracts of experimental neuroscience research papers rarely list the study group sizes used. So you will usually need to search the text of the paper for an indication of the study group sizes used.
Any experimental neuroscience paper using fewer than 15 subjects in any of its study groups should be regarded as a paper that has used a way-too-small study group size. There are actually reasons for thinking that any use of fewer than 25 subjects in any of the study groups is a reason for doubting the quality of the study, particularly if the work involves brain scans of humans.
Do not stop looking for study group sizes if you see a statement indicating a fairly large study group size such as 50. What very often happens in neuroscience papers is that the paper will announce a fairly large number of subjects (such as stating "50 mice were analyzed"), but will then divide this group up into smaller study groups, so that the smallest study groups used is much smaller than such a fairly large number. Look for any cases of any study group sizes smaller than 15.
How do you find what study group sizes were used? The easiest way is to search in the text for the phrases "n=" or "n =". It is a custom in neuroscience research papers to list study group sizes using phrases such as "n =8." For example, the text may vaguely refer to "subjects" or "mice" without specifying how many. Then the text of the paper may state the exact number of subjects by using a phrase such as "n = 10." Another way to search for study group sizes is to search for the phrases "mice," "rats," "subjects" or "humans" and look for a number preceding such phrases.
Whenever any such searches reveal a study group size of less than 15 or 20, you have discovered prima facie evidence of a too-small study group size, and you should record such a defect in your "defects list" file. Rarely you will find a paper that makes no mention of how many experimental subjects were used. The failure to record so vital a fact is itself a defect that you should note in your "defects list" file.
When an experimental neuroscientist is doing his job right, he will use some statistical method to do what is called a power analysis or a sample size calculation or a power size calculation. This involves some mathematical calculation of what sample size was needed to achieve some particular degree of statistical power. Many science journals require that a paper state whether or not such a calculation was done. Search for the phrase "sample size calculation" or "power size calculation" or "power calculation" to see whether such a calculation was done. You will often read a confession that no such calculation was done. If you find such a confession, write that down in your "defects list" file. If you fail to find any mention of such a calculation, that is also a defect that should be noted in your defects list file.
Step #6: look for a failure to use controls
Almost any experimental neuroscience experiment should be using controls. In experimental science a control can be a subject that does not have some characteristic or variable or intervention being tested, or a control can be a neutral state that does not match some experience or condition being tested. For example, if you are testing some medicine, you can give 15 subjects the medicine, and give 15 other subjects some placebo that is not some medicine. Or, if you are testing whether some cognitive activity such as memory recall causes increased activation of some brain region, you might take 15 brain scans while a subject was engaging in memory recall, and 15 brains "control" scans on some other day in which the subject was asked to think of nothing.
It is easy to check whether a scientific study made use of controls. Just do a text search in the paper for the word "control" looking for usage that indicates controls were used. Typically phrases such as "control subjects" or "control state" will be used. If you fail to find any evidence controls were used, record that failure in your "defects list" file.
Step #7: look for a failure to follow a detailed blinding protocol
In experimental neuroscience a blinding protocol is usually needed for a robust result. A blinding protocol is a procedure that helps to minimize the chance of a biased analysis. I can give some examples to illustrate the concept. Imagine you brain scan 15 subjects who were asked to recall memories while their brains were scanned, and you also brain scan 15 other subjects who were asked to think of nothing while their brains were being scanned. Then suppose you give the brain scans to some analyst. If the analyst knows that the first 15 brain scans were from people who were recalling things, and the second 15 subjects were from people thinking of nothing, and also that the purpose of the test is to look for brain differences in memory recall, such an analyst will be all too likely to "see what his bosses are hoping he sees," and produce a biased result. The risk of such bias could be avoided by a careful blinding protocol. Each set of brain scans could be assigned a random number, with someone recording which number matched a person recalling something, and which number corresponded to a person not recalling something. If there was a stack of 30 folders, each containing one subject's brain scans, the folders could be shuffled so that the analyst could not tell which one came from some one recalling something. It could then be stated in the paper that the person analyzing the brains was "blind to which subjects had engaged in memory recall."
There is another type of blinding that could occur. Instead of the brain scan analyst being told that half of the subjects were engaging in memory recall, and half were not, the analyst could be told nothing at all about what the people were doing during the scans. This would make it all the more unlikely that the analyst would see some effect that wasn't really there. Normally in every study there are multiple ways in which blinding should occur.
A failure to follow blinding protocols is one of the most egregious defects of today's neuroscience research. Most experimental neuroscience studies fail to follow any blinding protocol. The failure can easily be found if you have the full text of the paper. Simply do a text search for the word "blind." If you fail to find meaningful uses of the word "blind" in the text of the paper, thereby indicating a failure to follow a blinding protocol, record that failure in your "defects list" file.
One or two uses of the word "blind" in the text of a paper does not show that an effective blinding protocol was used. It is all too easy for some experimenter to have an ineffective blinding protocol that fails to achieve much of any real blinding. The smaller the study group size, the easier it is for any attempt at blinding to be ineffective. I will give an example. Suppose a study used only seven rats for an experimental group, and seven rats for a control group. The seven rats in the experimental group might be given some modification not given to the seven rats in the control group. After being given foot tags to identify them with random numbers, the fourteen rats might then be given to an analyst asked to test for some difference. But if the analyst was involved in applying the modification, it might be all too easy for him to recognize which rat had the modification, and which did not. For example, only the rats given the modification might have some surgical mark showing they had the modification. What this example shows is that to be effective, a blinding protocol must be a detailed, carefully thought-out plan that prevents "sham blinding" in which someone supposedly blind to which subjects were in the control group is not really blind to such a thing.
If you either fail to see the word "blind" being used in the text of an experimental neuroscience paper, or if you find the word "blind" or "blinding" only being used once or twice "in passing," you should note the result in your "defects list" file. Also note it in your "defects list" file if you fail to find a detailed discussion of a blinding protocol. If a study merely claims that some analysis was done by analysts "blind" to whether the subjects were control subjects, and fails to discuss how a careful detailed plan was followed to prevent "sham" blinding, that also should be recorded in your "defects list" file.
Step #8: look for convoluted analysis pathways that may have "conjured up" some illusory result
In my post "Convoluted 'Spaghetti Code' Analysis Pathways Help Neuroscientists Conjure Phantasms That Don't Exist," I gave two examples of scientific papers that used ridiculously convoluted analysis pathways. Don't be impressed when you see such methods, which may seem like gobbledygook or rigmarole. Such over-complicated methods are typically not signs of good experimental methods, but instead a failure to follow a straightforward technique for analyzing data. What is going on often can be describe as "keep torturing the data until it confesses."
Besides doing a quick scan looking for byzantine analysis pathways that sound like statistical "monkey business," there are some things you can look for:
(1) Look for the word "iterations" which typically indicates that some data was passed through a programming loop in which the data may have been distorted or contorted or convoluted.
(2) Look for the phrases "processing" (particularly "data processing," "preprocessing" or "pre-processing" or "post-processing" or "postprocessing"), phrases which indicate that data has been passed through some computer program. Once data has been passed through a computer program, there are any number of ways in which the data can be contorted, distorted or corrupted. Often computer programs processing neuroscience data are written by scientists who are not professional computer programmers, and who can often produce unreliable results. If the programming job is given to a professional programmer, the person may be someone who does not understand the data, leading to unreliable results. Programming code used in scientific research is very often poorly written and poorly commented, producing effects that may be known only to the original programming. Very often the result is a "black box" situation in which even the original programmer does not understand what is happening to the data.
If you find evidence of such dubious-sounding analysis pathways or dubious pre-processing or post-processing of brain scan data or neuroscience data, add some lines to your "defects list" file recording such a finding.
Step #9: look for fake data in the science paper, which may be described using the word "simulated" or "simulation."
Neuroscientists often look for things they cannot find but are hoping to find. When they cannot find such things, they often resort to generating simulated data to try to fill in the gap. Whenever you hear the word "simulated" in a neuroscience paper, you should presume that this actually means "fake." Search for the words "simulated" or "simulation" in a neuroscience paper. Note the resorting to such simulated data in your "defects list" file. Trying to pass off simulated data rather as if it was real-world data is a sleazy trick you should "throw a flag on" when critically analyzing neuroscience papers.
Step #10: look for p-hacking
Modern experimental science has the silly rule that a result is treated of worthy of publication if someone found it to have a "statistical significance" of .05 or less. This rule ends up being a very silly one in many cases in which it is very easy to get such a result when it is merely a false alarm. Very roughly you can think of a statistical significance of .05 as a result you might get by chance once in 20 tries. Given a lack of pre-registration in scientific studies, it is rather easy to get such a result. You can just keep trying something multiple times, calling these trials Experiment 1, Experiment 2, and so forth. When you get a result that you would get by chance once in 20 times, you can then write up that result, and describe only it.
How do you find evidence of unimpressive results such as this in a scientific paper? Statistical significance is reported using a phrase such as "p < .05" or "p < .01" or "p < .001." You can search for such phrases. Whenever you find the phrase "p < .05" it is a sign that an unimpressive result was obtained. Write down any such occurrence in your "defects list" file. The more examples you find of the phrase "p < .05" the stronger the case you can make that "p-hacking" went on.
Don't be impressed if you find a stronger statistical significance reported in addition to a marginal statistical significance of "p < .05." What often happens is that the most relevant result will be some marginal borderline result of "p < .05" but other results not very relevant will be reported with a stronger statistical significance such as "p < .01" or "p < .001." This is often done as a kind of window dressing to give you the impression that the paper has more impressive results than it has. The results with higher statistical significance may be irrelevant to the main claim made by the paper.
Step #11: look for lack of pre-registration
It has been pointed out many times that when an experimenter fails to state before gathering data a hypothesis to be tested and a detailed research plan for how to gather and analyze data, there will be a much higher chance of some false alarm being reported. A researcher who is free to "make up his method as he goes along" will be free to keep playing around with data analysis pathways until he seems to find something he was hoping to find. Pre-registration (also called the use of "registered reports") is when a researcher publishes a hypothesis to be tested and a detailed research plan before any data is gathered. It is widely recognized that following such a method greatly reduces the number of false alarms that are reported.
It is easy to search in a scientific paper for whether pre-registration occurred. Simply search in the text of the paper for the terms "registered report," "pre-registered" or "pre-registration." If no such terms are found, note this in your "defects list" file.
Step 12: look for unreliable measurements of memory, such as attempts to judge "freezing behavior."
There are reliable ways to judge whether an animal remembered something, and also unreliable ways. A reliable way to judge whether a a rodent trained to fear some stimulus (such as a shock plate) is to measure heart spikes when the animal is placed near the pain-inducing stimulus. Heart rates very dramatically spike in rodents when they are afraid. Another reliable way to judge whether a a rodent trained to fear some stimulus such as a shock plate is to use some setup such as the one shown below, in which an animal remembering the fear stimulus will take the harder path towards a reward rather the easier path.
A very unreliable way of measuring rodent recall of a fearful stimulus is to try to judge "freezing behavior," defined simply as immobility. My post here explains why such a technique is very unreliable as a way to judge whether recall or fear occurred. Attempts to judge so-called "freezing behavior" are massively used in experimental neuroscience involving rodents and memory. This is an example of a dysfunctional research community tradition. If you find that the neuroscience paper you are examining made any use of judgments of "freezing behavior," note that in your "defects list" file. All neuroscience papers relying on "freezing behavior" judgments are junk science.
Once you have taken such steps, you have the basis for a critical review of a neuroscience paper. Very often your "defects list" file will contain multiple examples of Questionable Research Practices, often more than five. Mentioning such defects, you can explain to a reader why some triumphal announcement of neuroscience progress is unjustified.






No comments:
Post a Comment