Thursday, March 17, 2022

Where I'm going, if I get there at all

About a decade ago, the field of psychology was rocked by what is now known as the replication crisis. (I say "rocked," but if you're not in the habit of tuning in to social-science news, you probably experienced this crisis like the seagulls in Finding Nemo, who floated so far removed from a deep sea explosion that it looked to them like a fart bubble.) 

The crux of the crisis -- which is still ongoing and has expanded to other branches of the social and natural sciences -- was the simple discovery that a whole bunch of studies failed to hold up when other research teams attempted to replicate them.

Take, for example, the social-psychological theory of ego depletion, a.k.a. decision fatigue, a.k.a. the depletion model of self-control. It's basically idea that everyone has a finite reservoir of willpower to be spent over a given period of time. (Imagine you spend all afternoon resisting the doughnuts and cake in the break room. You go home, and your spouse offers you a cookie. You're more likely to give in and take the cookie than if you hadn't already spent all day exercising your will power.) The fact that this theory has its origins in Freud might've been a warning sign -- still, ego depletion sounds plausible, feels intuitively true. Researchers hypothesizing the existence of ego depletion published hundreds of validating studies throughout the '90s and '00s. 

In 2016, however, a large-scale effort to analyze the actual validity of this phenomenon found... well, pretty much nothing. Since then, ego depletion has become one of the poster children of the replication crisis, a leading example of well-meaning social scientists finding significance where none has really been shown. This revelation prompted a leading researcher on ego depletion, Michael Inzlicht, to write a sorrowful, soul-searching blog post reflecting on his career. "As someone who has been doing research for nearly twenty years, I now can’t help but wonder if the topics I chose to study are in fact real and robust. Have I been chasing puffs of smoke for all these years?" he lamented. "I’m in a dark place. I feel like the ground is moving from underneath me and I no longer know what is real and what is not." Other titans of popular social-psychological theory (power poses! priming!) are in similar turmoil.

How does this kind of thing happen? It wouldn't be a full-blown crisis if it wasn't complicated. Conscious and unconscious bias, p-hacking, and seemingly innocuous practices in data analysis have all been identified as contributing to the problem.

But the ringleader of the crisis -- the four-star general under whom all the biases and bad practices are marshaled -- is probably simple publishing bias. Researchers are under tremendous pressure to publish, and the incentive structure for publishing heavily favors positive/significant results. There are plenty of studies (called null studies) that don't prove a positive result: they fail to show that the experiment had any significant effect whatsoever. These studies are, well, kind of boring and difficult to publish; researchers are disappointed when they happen. But they're also extremely important. Consider how our perception of the prevalence of skin cancer would be skewed if doctors only ever reported the results of biopsies that came back malignant and we had no idea how many came back benign.

On the plus side, this means there is at least one clear path to reform: incentivize null results! Make them a central part of teaching and textbooks. Devote journals to replication studies and null findings and recruit prominent members of the field to ensure they become prestigious. Help researchers, students, and institutions internalize the message that null isn't dull.

---

In medicine, null findings have something of a cousin in the concept of expectant management (a.k.a. watchful waiting, a.k.a. active surveillance). It's the puffed-up medical term for "doing nothing."

In the wild, you've probably encountered expectant management in its lowest-stakes form: going to the clinic for, say, an ear infection, only to hear the doctor say, "It'll probably clear up on its own, but if things haven't improved a week from now, give me a call." Other common contexts include kidney stones ("Let's see if it passes naturally"), cancer treatment ("The surgery is very risky, so let's wait to see if it gets worse before we act"), pregnancy ("You're at full term: we could induce, but let's wait a couple weeks to see if your water breaks on its own"), and infertility ("We're not sure why you're not getting pregnant; we could begin treatment, or you could keep trying").

As you might imagine, patients aren't always thrilled to have expectant management be their treatment plan. If you're seeking medical care, it's probably because you feel like an intervention of some kind is warranted. You want to do something.

Expectant management exists for good reason, though: intervention almost always carries drawbacks or risks, which we often downplay in our eagerness to act. Even a paltry antibiotics prescription, if prescribed unnecessarily, can contribute to the global scourge of antimicrobial resistance. And intervention doesn't even always yield a better or faster result than simply doing nothing.

Plus, most physicians would likely be quick to correct me: expectant management isn't the same as doing nothing, much in the same way that a null result isn't the same as no results. Watchful waiting entails monitoring and actively staying on top of the situation. Finding no significant effect in an experiment is itself a valuable result. 

But here's the thing: both expectant management and null studies ask us to confront the void, to stand by even though intervention is possible, to embrace non-significance rather than going back for a positive result. We shouldn't pretend that inaction is easy, or that researchers have nothing to lose by abandoning the quest for a significant p-value. 

Expectant management is a useful parallel for null studies, I think, because its stakes and frustrations are visceral, where the replication crisis is dry and academic. What they have in common is a mandate to accept absence, to make peace with the nonevent. What expectant management teaches us is that this acceptance is going to be an uphill battle.