Comments

Showing posts with label replication crisis. Show all posts
Showing posts with label replication crisis. Show all posts

Monday, August 27, 2018

Revolutions in science; a comment on Gelman

In what follows I am going to wander way beyond my level of expertise (perhaps even rudimentary competence). I am going to discuss statistics and its place in the contemporary “replication crisis” debates. So, reader be warned that you should take what I write with a very large grain of salt. 

Andrew Gelman has a long post (here, AG) where he ruminates about a comparatively small revolution in statistics that he has been a central part of (I know, it is a bit unseemly to toot your own horn, but heh, false modesty is nothing to be proud of either). It is small (or “far more trivial”) when compared to more substantial revolutions in Biology (Darwin) or Physics (Relativity and Quantum mechanics), but AG argues that the “Replication revolution” is an important step in enhancing our “understanding of how we learn about the world.” It may be right. But…

But, I am not sure that it has the narrative quite right. As AG portrays matters, the revolution need not have happened. The same ground could have been covered with “incremental corrections and adjustments.” Why weren’t they? The reactionaries forced a revolutionary change because of their reactions to reasonable criticisms by the likes of Meehl, Mayo, Ioannidis, Gelman, Simonsohn, Dreber, and “various other well-known skeptics.” Their reaction to these reasonable critiques was to charge the critics with bullying or insist that the indicated problems are all part of normal science and will eventually be removed by better training, higher standards etc. This, AG argues, was the wrong reaction and required a revolution, albeit a minor one relatively speaking, to overturn.

Now, I am very sympathetic to a large part of this position. I have long appreciated the work of the critics and have covered their work in FoL. I think that the critics have done a public service in pointing out that stats has served to confuse as often (maybe more often) than it has served to illuminate. And some have made the more important point (AG prominently among them) that this is not some mistake, but serves a need in the disciplines where it is most prominent (see here). What’s the need? Here is AG:[1]

Not understanding statistics is part of it, but another part is that people—applied researchers and also many professional statisticians—want statistics to do things it just can’t do. “Statistical significance” satisfies a real demand for certainty in the face of noise. It’s hard to teach people to accept uncertainty. I agree that we should try, but it’s tough, as so many of the incentives of publication and publicity go in the other direction.

And observe that the need is Janus faced. It faces inwards to relieve the anxiety of uncertainty and it faces outwards in relieving professional publish-or-perish anxiety. Much to AG’s credit he notices that these are different things, though they are mutually supporting. I suspect that the incentive structure is important, but secondary to the desire to “get results” and “find the truth” that animates most academics. Yes, lucre, fame, fortune, status are nice (well, very nice) but I agree that the main motivation for academics is the less tangible one, wanting to get results just for the sake of getting them. Being productive is a huge goal for any academic, and a big part of the lure of stats, IMO, is that it promises to get one there if one just works hard and keeps plugging away. 

So, what AG says about the curative nature of the mini-revolution rings true, but only in part. I think that the post fails to identify the three main causal spurs to stats overreach when combined with the desire to be a good productive scientist.

The first it mentions, but makes less off than perhaps others have. It is that stats are hard and interpreting them and applying them correctly takes a lot of subtlety. So much indeed that even experts often fail (see here). There is clearly something wrong with a tool that seems to insure large scale misuse. AG in fact notes this (here), but it does not play much of a role in the post cited above, though IMO it should have. What is it about stats techniques that make them so hard to get right? That I think is the real question. After all, as AG notes, it is not as if all domain find it hard to get things right. As he notes, psychometricians seem to get their stats right most of the time (as do those looking for the Higgs boson). So what is it about those domains where stats regularly fails to get things right that makes it the case that they so generally fail to get things right?  And this leads me to my second point.

Stats techniques play an outsized role in just those domains where theory is weakest. This is an old hobby horse of mine (see here for one example). Stats, especially fancy stats, induces the illusion that deep significant scientific insights are for the having if one just gets enough data points and learns to massage them correctly (and responsibly, no forking paths for me thank you very much). This conception is uncomfortable with the idea that there is no quick fix for ignorance. No amount of hard work, good ethics, or careful application suffices when we really have no idea what is going on. Why do I mention this? Because, in many of the domains where the replication crisis has been ripest are domains that are very very hard and where we really don’t have much of an understanding of what is happening. Or maybe to put this more gracefully, either the hypotheses of interest are too shallow and vague to be taken seriously (lots of social psych) or the effects of interest are the results of myriad interactions that are too hard to disentangle. In either case, stats will often provide an illusion of rigor while leading one down a forking garden path. Note, if this is right, then we have no problem seeing why psychometricians were in no need of the replication revolution. We really do have some good theory in the domains like sensory perception, and here stats have proven to be reliable and effective tools. The problem is not with stats, but with stats applied where they cannot be guided (and misapplications tamed) by significant theory.

Let me add two more codicils to this point.

First, here I part ways with AG. The post suggests that one source of the replication problem is with people having too great “an attachment to particular scientific theories or hypotheses.” But if I am right this is not the problem, at least not the problem behind the replication crisis. Being theoretically stubborn may make you wrong, but it is not clear why it makes your work shoddy. You get results you do not like and ignore them. That may or may not be bad. But with a modicum of honesty, the most stiff necked theoretician can appreciate that her/his favorite account, the one true theory, appears inconsistent with some data. I know whereof I speak, btw. The problem here, if there is one, is not generating misleading tests and non-replicable results, but if ignoring the (apparent) counter data. And this, though possibly a problem for an individual, may not be a problem for a field of inquiry as a whole. 

Second, there is a second temptation that today needs to be seriously resisted but that severely leads to replication problems: because of the ubiquity and availability of cheap “data” nowadays, the temptation to think that this time it’s different is very alluring. Big Data types often seem to think that get a large enough set of numbers, apply the right stats techniques (rinse and repeat) and out will plop The Truth. But this is wrong. Lars Syll puts it well here in a post entitled correctly “Why data is NOT enough to answer scientific questions”:

The central problem with present ‘machine learning’ and ‘big data’ hype is that so many –falsely- think that they can ge away with analyzing real-world phenomena without any (commitment to) theory. But –data never speaks for itself. Without a prior statistical set-up, there actually are no data at all to process. And – using a machine learning algorithm will only produce what you are looking for.

Clever data mining tricks are never enough to answer important scientific questions. Theory matters.

So, when one combines the fact that in many domains we have, at best, very weak theory, and that nowadays we are flooded with cheap available data the temptation to go hyper statistical can be overwhelming.

Let me put this another way. As AG notes, successful inquiry needs strong theory and careful measurement. Not the ‘and.’ Many read the ‘and’ as an ‘or’ and allow that strong theory can substitute for paucity of data or that tons of statistically curated data can substitute for virtual absence of significant theory. But this is a mistake. But a very tempting one if the alternative is having nothing much of interest or relevance to say at all. And this is what AG underplays: a central problem with stats is that it often tries to sell itself as allowing one to bypass the theory half of the conjunction. Further, because it “looks” technical and impressive (i.e. has a mathematical sheen) it leads to cargo cult science, scientific practice that looks "science" rather than being scientific. 

Note, this is not bad faith or corrupt practice (though there can be this as well). This stems from the desire to be, what AG dubs, a scientific “hero,” a disinterested searcher for the truth. The problem is not with the ambition, but the added supposition that any problem will yield to scientific inquiry if pursued conscientiously. Nope. Sorry. There are times when there is no obvious way to proceed because we have no idea how to proceed. And in these domains no matter how careful we are we are likely to find ourselves getting nowhere.

I think that there is a third source of the problem that resides in the complexity of the problems being studied. In particular, the fact that many phenomena we are interested in arise from the interaction of many causal sub-systems. When this happens there is bound to be a lot of sensitivity to the particular conditions of the experimental set up and so lots of opportunities for forking paths (i.e. p-hacking) stats (unintentional) abuse. 

Now, every domain of inquiry has this problem and needs to manage it. In the physical sciences this is done by (as Diogo once put it to me) “controlling the shit out of the experimental set up.” Physicists control for interaction effects by removing many (most) of the interfering factors. A good experiment requires creating a non-natural artificial environment in which problematic factors are managed via elimination. Diogo convinced me that one of the nice features of linguistic inquiry is that it is possible to “control the shit” out of the stimuli thereby vastly reducing noise generated by an experimental subject. At any rate, one way of getting around interaction effects problem is to manage the noise by simplifying the experimental set up and isolating the relevant causal sub-systems.

But often this cannot be done, among other reasons because we have no idea what the interacting subsystems are or how they function (think, for example, pragmatics).  Then we cannot simplify the set up and we will find that our experiments are often task dependent and very noisy. Stats offers a possible way out. In place of controlling the design of the set up the aim is to statistically manage (partial out) the noise. What seems to have been discovered (IMO, not surprisingly) is that this is very hard to do in the absence of relevant theory. You cannot control for the noise if you have no idea where it comes from or what is causing it. There is no such thing as a theory free lunch (or at least not a nutritious one). The revolution AG discusses, I believe, has rediscovered this bit of wisdom.

Let me end with an observation special to linguistics. There are parts of linguistics (syntax, large parts of phonology and morphology) where we are lucky in that the signal from the underlying mechanisms are remarkably strong in that they withstand all manner of secondary effects. Such data are, relatively speaking, very robust. So, for example, ECP or island or binding violations show few context effects. This does not mean to say that there are no effects at all of context wrt acceptability (Sprouse and Co. have shown that these do exist). But the main effect is usually easy to discern. We are lucky. Other domains of linguistic inquiry are far noisier (I mentioned pragmatics, but even large parts of semantics strike me as similar (maybe because it is hard to know where semantics ends and pragmatics begins)). I suspect that a good part of the success of linguistics can be traced to the fact that FL is largely insulated from the effects of the other cognitive subsystems it interacts with. As Jerry Fodor once observed (in his discussion of modularity), the degree to which a psych system is modular to that degree it is comprehensible. Some linguists have lucked out. But as we more and more study the interaction effects wrt language we will run into the same problems. If we are lucky, linguistic theory will help us avoid many of the pitfalls AG has noted and categorized. But there are no guarantees, sadly.



[1]I apologize for not being able to link to the original. It seems that in the post where I discussed it, I failed to link to the original and now cannot find it. It should have appeared in roughly June 2017, but I have not managed to track it down. Sorry.

Thursday, November 16, 2017

Is science broken/breaking?

Is science broken/breaking and if it is what broke/is breaking it? This question has been asked and answered a lot lately. Here is another recent contribution by Roy and Edwards (R&E). Their answer is that it is nearly broken (or at least severely injured) and that we should fix it by removing the perverse incentives that currently drive it. Though I am sympathetic to ameliorating many of the adverse forces R&E identify, I am more skeptical than they are that there is much of a crisis out there. Indeed, to date, I have seen no evidence showing that what we see today is appreciably worse than what we had before, either in the distant past (you know in the days when science was the more or less exclusive pursuit of whitish men of means) or even the more recent past (you know when a PhD was still something that women could only dream about). I have seen no evidence that show that published results were once more reliable than they are now or that progress overall was swifter.  Furthermore, from where I sit, things seem (at least from the outside) to be going no worse than before in the older “hard” sciences and in the “softer” sciences the problem is less the perverse incentives and the slipshod data management that R&E point to so much as the dearth of good ideas that can allow such inquiries to attain some explanatory depth. So, though I agree that there are many perverse incentives out there and that there are pressures that can (and often do) lead to bad behavior, I am unsure whether given the scale of the modern scientific enterprise things are really appreciably worse today than they were in some prior golden age (not that I would object to more money being thrown at the research problems I find interesting!). Let me ramble a bit on these themes.

What broke/is breaking science? R&E point to hypercompetition among academic researchers. Whence the hypercompetition? Largely from the fact that universities are operating “more like businesses” (2). What in particular? (i) The squeezed labor market for academics (fewer tenure track jobs and less pleasant work environment), (ii) The reliance on quantitative performance metrics (numbers of papers, research dollars, citations) and (iii) the fall in science research funding (from 2% of GDP in 1960 to .78% in 2014)[1] (p. 7) work together to incentivize scientists to cut corners in various ways. As R&E puts it:

The steady growth of perverse incentives, and their instrumental role in faculty research, hiring and promotion practices, amounts to a systematic dysfunction endangering scientific integrity. There is growing evidence that today’s research publications too frequently suffer from lack of replicability, rely on biased data-sets, apply low or sub-standard statistical methods, fail to guard against researcher biases, and overhype their findings. (p. 8)

So, perverse incentives due to heightened competition for shrinking research dollars and academic positions leads scientists interested in advancing their research and careers to conduct research “more vulnerable to falsehoods.” More vulnerable than what/when? Well, by implication than some earlier golden age when such incentives did not dominate and scientists pursued knowledge in a more relaxed fashion and were not incentivized to cut corners as they are today. [2]

I like some of this story. I believe as R&E argues that scientific life is less pleasant than it used to be when I was younger (at least for those that make it into an academic position). I also agree that the pressures of professional advancement make it costly (especially to young investigators) to undertake hard questions, ones that might resist solution (R&E quotes Nobelist Roger Kornberg as claiming: “If the work you propose to do isn’t virtually certain of success, then it won’t be funded” (8)).[3] Not surprisingly then, there is a tendency to concentrate on those questions for whose solution techniques already available apply and that hard work and concentrated effort can crack. I further agree that counting publications is, at best, a crude way of evaluating scientific merit, even if buttressed by citation counts (but see below). All of this seems right to me, and yet…I am skeptical that science as a whole is really doing so badly or that there was a golden age that our own is a degenerate version of. In fact, I suspect that people who think this don’t know much about earlier periods (there really was always lots of junk published and disseminated) or have an inflated view of the cooperative predilections of our predecessors (though how anyone who read the Double Helix might think this is beyond me).

But I have a somewhat larger problem with this story: if the perverse incentives are as R&E describes them then we should witness its ill effects all across the sciences and not concentrated in a subset of them. In particular, it should affect the hardcore areas (e.g. physics, chemistry, molecular biology) just as it does the softer (more descriptive) domains of inquiry (social psychology, neuroscience). But it is my impression that this is not what we find. Rather we find the problem areas more concentrated, roughly in those domains of inquiry where, to be blunt, we do not know that much about the fundamental mechanisms at play. Put another way, the problem is not merely (or even largely) the bad incentives. The problem is that we often cannot distinguish between those domains that are sciency from those domains that are scientific. What’s the difference? The latter have results (i.e. serious theory) that describe non-trivial aspects of the basic mechanisms where the former has methods (i.e. ways of “correctly” gathering and evaluating data) that are largely deployed to descriptive (vs explanatory) ends. As Suppes said over 60 years ago: “It’s a paradox of scientific method that the branches of empirical science that have the least theoretical developments have the most sophisticated methods of evaluating evidence.” It should not surprise that domains where insight is weakest are also domains where shortcuts are most accessible.

And this is why it’s important to distinguish these domains. If we look, it seems that the perverse incentives R&E identifies are most apt to do damage in those domains where we know relatively little. Fake data, non-replicability, bad statistical methods leading to forking paths/P-hacking, research biases, these are all serious problems especially in domains where nothing much is known. In domains with few insights and where all we have is the data then screwing with the data (intentionally or not) is the main source of abuse. And in those domains when incentives for abuse rise, then enticements to make the data say what we need them to say heightens. And when the techniques for managing the data are capable of being manipulated to make them say what you want them to say (or at least whose proper deployment eludes even the experts in the field (see here and here), then opportunity allows enticement to flower into abuse.

The problem then is not just perverse incentives and hypercompetition (these general factors hold in the mature sciences too) but the fact that in many fields the only bulwark against scientific malpractice is personal integrity. What we are discovering is that as a group, scientists are just as prone to pursuing self-interest and career advancement as any other group. What makes scientists virtuous is not their characters but their non-trivial knowledge. Good theory serves as a conceptual break against shoddy methods. Strong priors (which is what theory provides) is really important in preventing shoddy data and slipshod thinking from leading the field astray. If this is right, then the problem is not with the general sociological observation that the world is in many ways crappier than before, but with the fact that many parts of what we call the sciences are pretty immature. There is far less science out there than advertised if we measure a science by the depth of its insights rather than the complexity of its techniques (especially, its data management techniques).

There is, of course, a reason for why the term ‘science’ is used to cover so much inquiry. The prestige factor. Being “scientific” endows prestige, money, power and deference. Science stands behind “expertise,” and expertise commands many other goodies. There is thus utility in inflating the domain of “science” and this widens the possibilities for and advantages of the kinds of problems that R&E catalogue.

R&E ends with a discussion of ways to fix things. They seem like worthy, albeit modest fixes. But I personally doubt they will make much of a difference. They include getting a better fix on the perverting incentives, finding better ways to measure scientific contribution so that reward can be tuned to these more accurate metrics, and implementing more vigilant policing and punishment of malefactors. This might have an effect, though I suspect it will be quite modest. Most of the suggestions revolve around ways of short-circuiting data manipulation. That’s why I think these suggestions will ultimately fail to do much. They misdiagnose the problem. R&E implicitly takes the problem to mainly reside with current perverse incentives to pollute the data stream for career advancement. The R&E solution amounts to cleaning up the data stream by eliminating the incentives to dirty it. But the problem is not only (or mainly) dirty data. The problem is our very modest understanding for which yet more data is not a good substitute.

Let me end with a mention of another paper. The other paper (here) takes a historical view on metrics and how it affected research practice in the past. It notes that the problems we identify as novel today have long been with us. It notes that this is not the first time people are looking for some kind of methodological or technological fix to slovenly practice. And that is the problem; the idea that there is a kind of methodological “fix” available, a dream for a more rigorous scientific method and a clearer scientific ethics. But there is no such method (beyond the trivial do your best) and the idea that scientists qua scientists are more noble than others is farfetched. Science cannot be automated. Thinking is hard and ideas, not just data collection, matter. Moreover, this kind of thinking cannot be routinized and insights cannot be summoned not mater how useful they would be. What the crisis R&E identifies points to, IMO, is that where we don’t know much we can be easily misled and easily confused. I doubt that there is a methodological or institutional fix for this.


[1] It is worth pointing out that real GDP in 2016 is over five times higher than it was in 1960 (roughly 3 trillion vs 17 trillion) (see here). In real terms then, there is a lot more money today than there was then for science research. Though the denominator went down, the numerator really shot up.
[2] Again, what is odd is the dearth of comparative data on these measures. Are findings less replicable today than in the past? Are data sets more biased than before? Was statistical practice better in the past? I confess that it is hard to believe that any of these measures have gotten worse if compared using the same yardsticks.
[3] This is odd and I’ve complained about this myself. However, it is also true that in the good old days science was a far more restricted option for most people (it is far more open today and many more people and kinds of people can become scientists). However, what seems more or less right is that immediately after the war there was a lot of government money pouring into science and that made it possible to pursue research that did not show immediate signs of success. What would be nice to see is evidence that this made for better more interesting science, rather than more comfortable scientists. 

Sunday, October 30, 2016

More on the collapse of science

The first dropped shoe announced the “collapse” of science. It clearly dropped with a loud bang as this “news” has become a staple of conventional wisdom. The second shoe is poised and ready to drop. It’s ambition? To explain why the first shoe fell. Now that we know that  science is collapsing we all want to know why exactly it is doing so and whether there is anything we can do to bring back the good old days.

So why the fall? The current favorite answer appears to be a combination of bad incentives for ambitious scientists and statistical tools (significance testing being the current bête noir) that “gave scientists a mathematical machine for turning baloney into breakthroughs, and flukes into funding” ((now that’s a rhetorical flourish!) cited here p. 12). So, powerful tools in ambitious hands lead to scientific collapse. In fact, ambition may be beside the point, academic survival alone may be a sufficient motive. Put people in hyper competitive environments and give them a tool that “lets” them get their work “done” in a timely manner and all hell breaks loose.[1]

I have just read several papers that develop this theme in great detail. They are worth reading, IMO, for they do a pretty good job of identifying real forces in contemporary academic research (and not limited to the sciences). These forces are not new. The above “baloney” quote is from 1998 and there are prescient observations relating to somewhat similar (though not identical) effects made as early as 1948. Here’s Leo Szilard (cited here):

Answer from the hero in Leo Szilard’s 1948 story “The Mark Gable Foundation” when asked by a wealthy entrepreneur who believes that science has progressed too quickly, what he should do to retard this progress: “You could set up a foundation with an annual endowment of thirty million dollars. Research workers in need of funds could apply for grants, if they could make a convincing case. Have ten committees, each composed of twelve scientists, appointed to pass on these applications. Take the most active scientists out of the laboratory and make them members of these committees. ...First of all, the best scientists would be removed from their laboratories and kept busy on committees passing on applications for funds. Secondly the scientific workers in need
of funds would concentrate on problems which were considered promising and were pretty certain to lead to publishable results. ...By going after the obvious, pretty soon science would dry out. Science would become something like a parlor game. ...There would be fashions. Those who followed the fashions would get grants. Those who wouldn’t would not.”

The papers I’ve read come in two flavors. The first are discussions of the perils of p-values. Those who read the Andrew Gelman blog are already familiar with many of the problems. The main issue seems to be that phishing for significance is extremely hard to avoid, even by those with true hearts and noble natures (see the Simonsohn (a scourge of p-hacking) quote here). Here (and the more popular here) are a pair of papers that go into how this works in ways that I found helpful. One important point the author (David Colquhoun (DC)) makes is that the false discovery (aka: the false positive) problem is quite general, and endemic to all forms of inductive reasoning. It follows from the “obvious rules of conditional probabilities.” So this is not just a problem for Fisher and significance testing, but applies to all modes of inductive inquiry, including Bayesian modes.

Assuming this is right and that even the noble might be easily mislead statistically, is there some way of mitigating the problem? One rather pessimistic paper suggests that the answer is no. Here (with a popular exposition here) is a paper that gives an evolutionary model of how bad science must win out over good in our current academic environment. It is a kind of Gresham’s law theory where quick successful bad work floods less quick, careful good work. In fact, the paper argues that not even a culture where replication is highly valued will stop bad work from pushing out the good so long as “original” research remains more highly valued than “mere” replication.

The authors, Smaldino and McElreath (S&M), base these grim projections on an evolutionary model they develop which tracks the reward structure of publication and the incentives that these impose on individual and labs. I am no expert in these matters, but the model looks reasonable enough and the forces it identifies and incorporates seem real enough. The solution: shift from a culture that rewards “discovery” to one that rewards “understanding.”

I personally like the sound of this (see below), but I am skeptical that it is operationalizable, at least institutionally. The reason is that valuing understanding requires exercising judgment (it involves more than simple bookkeeping) and this is both subjective (and hence hard to defend in large institutional settings) and effortful (which makes it hard to get busy people to do). Moreover, it requires some very non-trivial understanding of the relevant disciplines and this is a lot to expect even within small departments, let alone university wide APT committees or broad based funding agencies. A tweet by a senior scientist (quoted in S&M p.2) makes the relevant point: “I’ve been on a number of search committees. I don’t remember anybody looking at anybody’s papers. Number and IF [impact factor] of pubs are what counts.” I don’t believe that this is only the result of sloth and irresponsibility. In many circumstances it is silly to rely on your own judgment. Given how specialized so much good work has become, it is unreasonable to think that we can as individuals make useful judgments about the quality of work. I don’t see this changing, especially above the department level anytime soon.

Let me belabor this. It is not clear how people above the department level would competently judge work outside their area of expertise. I know that I would not feel competent to read and understand a paper in most areas outside of syntax, especially if my judgment carried real consequences. If so, who can we get to judge whose judgments would be reasonable?  And if there is no one then what can one do but count papers weighted by some “prestige” factor? Damn if I know. So, I agree that it would be nice if we could weight matters towards more thoughtful measures that involved serious judgment, but this will require putting most APT decisions in the hands of those that can make these judgments, namely leave them at effectively the department level, which will not be happening anytime soon (and which has its own downsides if my own institution is anything to go by).

An aside: this is where journals should be stepping in. However, it appears that they are no longer reliable indicators of quality. Many are very conservative institutions whose stringent review processes tend to promote “safe” incremental findings. Many work hard to protect their impact factors to the point of only very reluctantly publishing work critical of previously published work. Many seem just a stones throw removed from show business where results are embargoed until an opening day splash can be arranged. At any rate, professional journals is a venue in which responsible judgment could be exercised, but, it appears, that it is difficult even here.

So, there are science (indeed academy) wide forces imposing shallow measures for evaluation and reward that bad statistical habits can successfully game. I have no problem believing this. But I still do not see how these forces suffice to explain the “crisis” before us. Why?  Because such explanations are too general and the problems appear to hold not in general but in localizable domains of inquiry. More exactly, the incentives S&M cites and the problems of induction that DC elaborates are pervasive. Nonetheless, the science (more particularly, replication) crisis seems localized in specific sub-areas of investigation, ones that I would describe as more concerned with establishing facts than in detailing causal mechanisms. [2] Here’s what I mean.

What’s the aim of inquiry? For DC it is “to establish facts, as accurately as possible” (here, 1). For me, it is to explain why things are as they are.[3] Now, I concede that the second project relies on the first. But I would equally claim that the first relies on the second. Just as we need facts to verify theories, we need theories to validate facts. The main problem with lots of “science” (and I am sure you won’t be surprised to hear me write this) is that it is theory free. Thus, the only way to curb its statistical enthusiasm is by being methodologically pristine. You gotta get the stats exactly right for this is the only thing grounding the result. In most cases of drug trials, for example, we have no idea why they work, and for practical purposes we may not (immediately) care. The question is do they, not how. Sciences stuck in the “does it” stage rather than the “how does it do it and why” stages, not surprisingly, have it tough. Fact gathering in the absence of understanding is going to really hard even with great stats tools. Should we be surprised that in areas where we know very little that stats can and do regularly mislead?

Note that the real sciences do not seem to be in the same sad state as psych, bio-med and neuroscience. You don’t see tons of articles explaining how the physics of the last 20 years is rotten to its empirical core. Not that Nobel winning results are not challenged. They can be and are. Here’s a recent example in which dark energy and the thesis that the universe is expanding at an accelerating rate is being challenged (see here) based on more extensive data. But in this case, evaluation of the empirical possibilities heavily relies on a rich theoretical background. Here’s a quote from one of the lead critics. Note how the critique relies on an analysis of an “oversimplified theoretical model” and how some further theoretical sophistication would lead to different empirical results. This interplay between theory and data (statistically interpreted data by and large) is not available in domains where there is no “fundamental theory,” (i.e. non-trival theory).

'So it is quite possible that we are being misled and that the apparent manifestation of dark energy is a consequence of analysing the data in an oversimplified theoretical model - one that was in fact constructed in the 1930s, long before there was any real data. A more sophisticated theoretical framework accounting for the observation that the universe is not exactly homogeneous and that its matter content may not behave as an ideal gas - two key assumptions of standard cosmology - may well be able to account for all observations without requiring dark energy. Indeed, vacuum energy is something of which we have absolutely no understanding in fundamental theory.'

So, IMO, the problem with most problematic “science” is that it is not yet really science. It has not moved from the earliest data collection stage to the explanation stage where what’s at issue are not facts but mechanisms. If this is roughly right, then the “end of science” problems will dissipate as understanding deepens (if it ever does (no guarantee that it will or should)) in these domains. So understood, the demise of science that replication problems herald is more a problem for the particular areas identified (and more an indication of how little is known here) than for science as a whole.[4]

That said, let me end with one or two caveats. The science-in-crisis narrative rests on the slew of false discoveries regularly churned out. Szilard’s worry mooted in the quote above is different. His worry is not false discoveries but the trivialization of research as big science promotes quantity and incrementalism over quality and concern for the big issues. Interestingly, this too is a recurrent theme. Szilard voiced this worry over 60 years ago. More recently (the last 15 years or so), Peter Lawrence voiced similar concerns in two pieces that discuss Szilard’s problem in the context of how scientific work is evaluated for granting and publication (here and here). And the problem is discussed in very much the same terms today. Here (and here) are two papers in Nature from 2016 which address virtually the same questions in virtually the same terms (i.e. how institutions reward more of the same reserach, punish thinking about new questions, look at publication numbers rather than judge quality etc.). What is striking is that this is all stuff noted and lamented before and the proposed fixes are pretty much the same: calls for judgment to replace auditing.

I agree that this would be a good idea. In fact, I believe that one of the reasons for the disparagement of theory in linguistics is a reflection of the same demands it makes on judgment for adequate evaluation. It is easier to see if a story “captures” the facts than to see if it offers an interesting explanation. So I am all in favor of promoting judgment as an important factor in scientific evaluation. However, to repeat, I am skeptical this is actually doable as judgment is not something that bureaucracies do well and like it or not, today science is big and so, not surprisingly, it comes with a large bureaucracy attached. Let me explain.

Today science is conducted in big settings (universities, labs, foundations, funding agencies). Big settings engender bureaucratic oversight, and not for entirely bad reasons. Bureaucracies arise in response to real needs where the actions of large numbers of people require coordination. And given the size of modern science, bureaucracy is inevitable. Unfortunately, bureaucracies by necessity favor blunt metrics over refined judgment (i.e. quantitative auditable measures over nuanced hard to compare evaluations). And all of this fosters the problems that Szilard and Lawrence and the Nature comments worry about. As noted, I think that this is simply unavoidable given the current economics of research. The hopeful (e.g. Lawrence) think that there are ways of mitigating these trends. I hope they are right. However, given the fact that this problem recurs regularly and the same solutions get suggested just as regularly, I have my doubts.

Let me end on a more positive note. It may not be possible to inject judgment into the process in a systematic way. However, it may be possible to find ways to promote unconventional research by having a sub-part of the bureaucracy looking for it. In the old days when money was plentiful, “whacky” research got institutional support because everything did (think of the early days of GG funding, or early CS). When money gets scarcer we need to still put aside some for work to support the unconventional. This is a problem in portfolio management: put most of your cash on safe stuff and 10% or so on unconventional stuff. The latter will mostly fail, but when it pays off, it pays off big. The former rarely fails, but its payoffs are small. Maybe the best we can do right now is allow our institutions to start thinking about the wild 10% just a little bit more.

So, the replication crisis will take care of itself as it is largely a reflection of the primitive nature of most of the “science” that it infects. The trivialization problem, IMO, is more serious and here, IMO, the problem is and will remain much harder to solve.



[1] I have long thought that stats should be treated a little like the Rabbis treated Kabbalah. The Rabbis banned its study as too dangerous until the age of forty, i.e. explosive in the hands of clever but callow neophytes.
[2] The collapse seems to be restricted. In psych, it is largely restricted to social psych. Perception and cognition, for example, seem relatively immune to the non-replicability disease. In bio-medicine, the bio part also seems healthy. Nobody is worrying about the findings in basic cell biology or physiology. The problem seems limited to non “basic” discoveries (e.g. is cholesterol/fat bad for you, does such and such drug work as advertised, and so on). In neuroscience the problems also seem largely restricted to fMRI results of the sort that make it into the NYTs. If one were inclined to be skeptical, one might say that the problems arise not in those areas where we know something about the underlying mechanisms but in those domains where we know relatively little. But who would be so skeptical?
[3] The search for explanation ends up generating novel data (facts). But the aim is not to establish new facts but to understand what is going on. In the absence of theory it might even be hard to know what a “fact” is.  Is it a fact that the sun rises in the East and sets in the west? Well, yes and no. It depends.
[4] It also reflects the current scientism of the age. Nothing nowadays is legit unless wrapped up in scientific looking layers. Not surprisingly much trivial insight is therefore statistically marinated so that it can look scientific.