Technology

An “Omics” Answer to the Replication Crisis

The role of scientists is to discover things, and the more valuable those things are, the better. But science is an ongoing process, and very few answers are ever definitive. We are continually gaining more knowledge about the world around us and revising our theories and hypotheses. So the outcome of scientific inquiry is not a fixed “fact,” but rather a knowledge claim that is presented in a scientific article, and that the reader is invited to accept. About ten years ago it became clear — first in psychology, then in other fields — that many of these knowledge claims, even those we had treated as canonical beliefs, turned out to be wrong. The findings could not be replicated, and the “replication crisis” was born.

Since then, many of us have been worrying away at the issue, trying to understand what has been going wrong in science. I first became involved in this area as a trainee neurologist, interested in running clinical trials of new treatments for stroke. Looking at the data emerging from the lab, there was no lack of promising drugs we might test. But when we examined things closely, every drug that had been tested in clinical trials was ineffective, apart from one exception (a clot-busting treatment more like a plumbing fix than genuine neuroprotection). It turns out the same has been true across neurological diseases, including Alzheimer’s disease, Huntington’s disease, Parkinson’s disease, Amyotrophic Lateral Sclerosis, and brain tumors. There was a basic disconnect between the research being done in the lab and what we were seeing in the clinic, and it had a real human cost. So how do we bring the two into better alignment?

Crucially, we need to be far more deliberate in the way we use published research findings (from the lab) to shape what we do next (e.g. clinical trials). We need to bring together and integrate — systematically — the full detail of research claims and where they came from, so that our understanding has more richness and depth. This kind of large-scale information integration is well established in the “-omics,” where vast amounts of data about genes (genomics) or proteins (proteomics), for instance, are merged into a single knowledge system. I propose that we do the same with the oceans of data embodied in published research claims, in a new approach I’m calling “publomics.”

But first, how did we get here?

At least part of the explanation for the replication crisis lay in weaknesses in the provenance of the original knowledge claims. When we examine how those original studies were carried out, we find a number of recurring shortcomings. These include:

No blinding. Experiments need both a test arm and a control arm. “Blinding” means the researcher does not know whether the subject being analyzed (whether proteins, cells, mice, or human subjects) belongs to the test or control group. Results are more trustworthy when the researcher is blinded, and this is standard practice in clinical trials, but it is often absent or unreported in preclinical experiments, which leads to biased results. We also want experimental subjects to be assigned to the test and control arms at random — but again, this is rarely done and rarely reported in the lab.

Small sample sizes. We also need to know that the study was large enough to support the claims being made. For example, it is pretty obvious that men are on average taller than women. But to have a good chance of detecting this difference in an experiment, you would need at least 20 people to find a statistically significant result. Unfortunately, effects this large are uncommon and unlikely in biology, and larger study populations are needed to make well-supported claims. Results based on small sample sizes can contaminate the literature with knowledge claims built on artifacts.

Publication bias. Perhaps unsurprisingly, the scientific literature is filled with positive results. Null or negative results have traditionally been difficult to publish and are often left on abandoned hard drives, even though they represent valuable knowledge claims. Research users looking for evidence get a distorted picture, which — like Johnny Cabot in the 1944 film Here Come the Waves — accentuates the positive, eliminates the negative, and latches on to the affirmative. While that may work in the movies, it is not a sound basis for drug development or genuine scientific progress.

These and other common design flaws combine to mislead the clinical trialist, the research funder, and the investor about what counts as sound, rigorous science. More importantly, this lets down patients — the people, and their families, who endure the consequences of these diseases. It is heartbreaking to watch people go through cycle after cycle of excitement and promise and hope, only to see yet another promising therapy fail in clinical trials. We owe it to them to do better.

There has to be a better way

So what would “better” look like? A more realistic assessment of the prospects for scientific success might not generate as much buzz, but it could be much more efficient. That would mean examining all the available evidence, not just cherry picking the best bits. It would mean doing due diligence on the research claims made — were the studies randomized and blinded, with preregistered designs, sufficiently large sample sizes, and power calculations, or are the authors silent on these important details? Just because these things were not done does not mean the findings are not true — but it makes it more likely that they are not true.

Fortunately, there is a well-established strategy for finding this out, through approaches called systematic review and meta-analysis. In systematic review, you use a defined literature search strategy to identify relevant research. It does not guarantee that you will find everything, but it is substantially better than starting at the first page of PubMed and choosing the material that looks interesting. A systematic review is often followed by a meta-analysis, which lets you combine, quantitatively, the results from different publications. These analyses can show whether the findings from different papers on the same subject are broadly consistent, or all over the place. For example, does the treatment work in both sexes, at all ages, in different species? If so, this would strongly support the hypothesis. If not, you may be dealing with an artifact, edge, or corner case. If the literature does not give you that complete picture, it may make more sense to fill in those gaps with further pre-clinical experimental work before starting a clinical trial.

To fully reap the benefits of our collective advances in science and technology for health gain, a careful and systematic evaluation of the available information, with due diligence on the strengths, and relevance of the research claims made, is essential. If this were done at critical stages — such as the initiation of human clinical trials — it would reduce the risk that the premise information on which those trials was based was not well founded. That would benefit the trial participants (who wants to be in a trial where the chances of success are low?), the clinical trialists (who could then focus on trials of agents with a better prospect of success), and those who pay for the trials (de-risking their investment). This is not only true for the move to clinical trial, but for any transition that requires a major commitment of time, people, or money.

There is of course a catch. If it is that simple, why has it not become standard practice?

Firstly, nobody wants to see their beautiful ideas set aside — the last thing we want is for someone to prove that we are wrong. So although many scientists welcome, in the abstract, the principle that our theories should be tested to destruction — we would very much prefer that that did not happen.

But even if you were completely committed to doing a systematic review, there are challenges. This research-on-research, or “meta-research,” is not well funded (generally, tools designed to test the theories of scientists to destruction are not warmly received by funding panels made up of those scientists). That means we have not built the research capacity — we do not have enough people with the right skills — to do all the systematic reviews and meta-analyses that need done. And the process itself can be burdensome. Looking for research into Alzheimer’s disease, for instance, identifies over 400,000 publications. Figuring out what is relevant and what is not, and judging the quality of the research, is a Herculean task.

This is where technology comes in

If you search PubMed for research on Alzheimer’s disease — one of our most pressing clinical needs — it lists over 170,934 results. At half an hour per article for 40 hours a week that is over 41 years of reading. It is simply not possible for one person to integrate these claims into an organized knowledge system — our understanding is instead based on abstraction and simplification, and lacks granularity. What if we could take these claims — and their context — and describe phenomena which are always observed, which are never observed, and which are observed in some contexts but not in others? What if we could map these claims from animal research with knowledge claims from human studies, from cell culture? What if we could do it all at the touch of a button?

Fortunately, this is where the emerging tools of big data can help. Advances in text mining and natural language processing let us automate many of the steps in systematic review, so that specific research questions can be addressed at much lower cost and effort.

This prospect is what I mean by “Publomics” — taking the automation approach to its logical extreme, to sit with genomics and proteomics and metabolomics and all the other “-omics” that have transformed our understanding of biology (and provided a richer understanding of human health as well). Every research artifact — publication, or pre-print, or dataset — would be richly and automatically indexed and curated in a central repository. In addition to the details of the experiments (the cells or animals or people studied, the diseases modeled, the interventions tested, the outcomes measured) and the outcome data (the numbers, not the interpretation), this would also include details of how the researchers sought to reduce risks of bias in their work, for instance through randomisation or blinding or more robust sampling.

Then, a systematic search, with all the tricky and tedious work involved today, would no longer be required. All that would be needed for a meta-analysis is the selection of the outcomes of interest, and the pressing of a button.

With apologies to my colleagues in the field, we would make systematic review redundant.

The benefits of this “Publomics” approach extend beyond the replication crisis, by also enabling higher-order analyses, across literatures, of interesting questions like: the concordance between different outcome measures, and the construction of evidence-based networks, describing the different components of mechanistic pathways and how the functioning of those pathways differs for instance across variables like species, sex, and age.

To illustrate how this could operate, let’s use multiple sclerosis as an example. It is a devastating disease that needs better treatments. In animal studies of new drugs, researchers record and report a range of measures, including the effects of candidate drugs on inflammation, on injury to nerve fibres (axon loss), on injury to the surrounding myelin (demyelination), and on animal behavior. By bringing together data from nearly 300 animal studies that assessed drug effects on at least two of these outcomes, our meta-analysis found that early treatment aimed at inflammation was highly effective, but over time, drugs that did not improve axon loss had no effect on behavior. This distinction is vital, because most patients arrive with established disease, implying that targeting inflammation will not help many patients — and that treatments protecting against axon loss are needed, an important insight for drug development.

If that is what you can uncover by pooling data from 300 experiments, what value lies in a body of 30,000 publications, and how can we make use of it? Text mining and machine learning methods can take us part of the way, but without high-quality source information (risks of bias in the primary studies, detailed annotation of the experimental design), they can provide only the broadest outline — like snapping an ultra high-resolution image with a blurred camera. With Publomics, we could address that issue by ensuring that our focus, our level of detail, is perfectly sharp.

There is already a strong example of this strategy in the clinical field. The Trialstreamer platform (https://trialstreamer.robotreviewer.net/) uses a high-resolution screen to find reports of human clinical trials, then relies on automation tools to pull out information on the study population, the drug being tested, the outcomes measured, and risks of bias in the research design.

But laboratory research is more complex, because papers often present several experiments at the same time, and disentangling them can be difficult (for humans as well as for machines). Nevertheless, it is becoming feasible. In our prototype Systematic Online Living Evidence Summary of animal research testing interventions for Alzheimer’s disease (https://camarades.shinyapps.io/LivingEvidence_AD/), for example, we have coded the animal model, intervention, outcomes measured, and risks of bias in almost 20,000 publications using transgenic animal models of Alzheimer’s disease. The next step is to find a way to extract the outcome data, so they can feed into meta-analytical or machine learning platforms. These are not yet ready for prime time, but they are not far off.

There will, of course, be resistance. The effort needed to get us there may be viewed by some as too great, and the accuracy of our tools less than perfect. There may also be concern that making such understanding available at the click of a button could weaken the status and cost of expert opinion. But we do not have to pick between experts and machines: expert opinion that draws from and builds on a Publomics approach is likely to be far more valuable than either approach alone.

Having tools like this will change the way we can use, apply, and deploy research findings from thousands of scientists around the world, and will not only provide one answer to the replication crisis and give us greater confidence in our understanding of the literature. It may also help us discover cures sooner and more effectively, changing the lives of many people whose existence is blighted by disease.

About the author

Malcolm MacLeod is a professor at the University of Edinburgh and co-founder of the Collaborative Approach to Meta-Analysis and Review of Animal Data from Experimental Studies (CAMARADES).