Comments

Tuesday, April 18, 2017

An upcoming conference of some importance

Angel Gallego and Dennis Ott are running a kind of Athens II on the future of GG (I assume that syntax is the main focus, but I could be wrong). Here is the website. Unlike Athens which aimed to get graybeards talking about the state of play, this conference is aimed at the up and comers. I hope that they have better success than we did. I think Athens did a credible job at identifying where we have come from and what we have (sorta) accomplished. However, I do not think that a clear future direction emerged and, IMO, the confab left me with the impression that the field takes theoretical syntax as something very much to be avoided. My lasting impression is that the consensus view was that Minimalism has largely been a failure and that any work that does not resemble what we have always done is not really worth doing. In short, I found the whole thing, in retrospect, quite disheartening.

But, forget the once bitten twice shy thing. Gallego and Ott are going to give it another try, this time with people with more skin in the game; namely younger up and comers.  This is a great idea and I hope they fare better than we did. Here is the "manifesto."

For the last sixty years, Generative Grammar (GG) has been the dominant approach in the formal study of human language within the cognitive sciences. Building on classical ideas developed in a new context, GG provided the basis for a new wave of investigations that gave rise to significant theoretical and empirical discoveries, establishing a fertile ground for synergies with other disciplines.
This workshop extends an invitation, especially -- but not exclusively -- to young researchers to pause and reflect on the current state of the field.
  • Have the recent developments yielded a corresponding increase in the explanatory depth of the theory? 
  • Has cross-fertilization with other theories beneffited the research?
  • Where have we made real progress and what questions should most urgently be addressed?

The workshop will be structured around the contributions of a number of invited participants who will submit a position paper ahead of time and give short informal talks, which will serve to kickstart open discussion.

Students of all levels as well as professional linguists of any persuasion are encouraged to actively participate in the event, either in person or by using other modes of communication.
Based on these and other questions of this general kind, our goal is to assess where we are headed, why there, and what the best way is of getting there (or nearby).

Note that all are invited and also note that I am willing to bet that Barcelona in June is quite delightful. So the venue should be conducive to real interesting discussion. If I get hold of the position papers I will link to them here at FoL. So go, argue and let's make progress.







Monday, April 10, 2017

Is the examined academic life worth it?

Every year grad departments select an incoming class, funding agencies decide on who is going to get how much and journals decide who is going to get published. Each of these decision involves a scare resource issue: a (usually small) finite number of positions, a (usually much too small) finite number of dollars and a (usually small) finite number of pages.[1] The decisions amount to how to allocate these scare resources. The reason the resources are scarce is that there are often more applicants/submissions than there are places to allocate. But the real reason resources are scarce is that quality alone, at least evident quality, does not suffice to perfectly fit the applicants/submissions to the allocations available. In other words, there are many ways to allocate the money to high quality applicants/submissions, at least prima facie.  And that’s the problem.

How do academics solve this problem? In general we do this by applying every more refined (i.e. recondite?) markers of quality to winnow out the really truly deserving from the merely truly deserving from the merely deserving. It is well known that even blunt methods suffice to lop off the clearly undeserving from the rest. In other words, academics embrace the conceit that if we try hard enough we can find the very very very…very best, and can thus optimize our choices; the decision procedure being a simple easily defensible one: select the most deserving and allocate to them the scarce resources. So, every year departments (including mine, of course) select the very very…very best of the applicant pool, funding agencies fund the very very very…very best proposals and journals publish the very very very…very best papers. Those readers that find that the last sentence rings true, please follow me as I have bridge for sale for you to look at.

Ok, everyone knows that the above story is more than a tad tendentious. We all recognize that our powers of discernment, even if applied with the utmost seriousness (which is seldom the case), are loaded with piles of false positives and negatives. We know this, but as an institution we also believe that overall looking for the best and rewarding it is the superior strategy. But is it? Here is a recent piece that questions this in interesting ways and suggests an alternative.

Before getting into the nub of the proposal, here are some points that the piece makes that (largely, though not perfectly) fit with my experience (see some comments below). The author Shahar Avin (SA), only discusses research funding, but I would add that the same holds for journal pages, and grad slots.

1.     Currently, “the monetary cut-off point still tends to be way above the quality cut-off point.”
2.     “…expert reviewers spend a lot of time allocating grant money by trying to identify the best work. But the truth is that they’re not very good at it, and that the process is a huge waste of time.”
3.     Peer review, rather than adding information, often adds “another layer of irrationality” to the decision process.
4.     The application process asks that you to be more sure of where you are going and how you will get there than it is rational to expect someone who is aiming at original research to be.
5.     Our capacities to identify those novel ideas most likely to succeed is much more limited than we believe.
6.     The belief that we can make this kind of rational decision “demand[s] excessive amounts of information from applicants, and waste a colossal amount of their time.”
7.     For some areas, we can even quantify how time consuming this is. SA cites a study in Nature which calculates that “[i]n Australia, during a recent annual funding round for medical research, scientists spent the equivalent of 400 years writing applications that were eventually rejected.” That’s a lot of effort.
8.     “…‘expert reviewers’ are not fungible commodities. One reviewer is not the same as another, and their judgements tend to be highly personal. Of the nearly 3,000 medical research proposals submitted for public funding in Australia in 2009, nearly half would have received the opposite decision if the review panel had been different…”
9.     “…the process isn’t just ineffective – it’s systematically biased. There’s evidence that women and minorities have lower chances of securing grants than people who are male or white, respectively.”
10.  Unorthodox proposals are at a disadvantage in the current funding structures.

As I said, most of these observations seem right to me. Points 1/2 perfectly fits my little world. In my experience the quality of the typical applicant pool, after the first triage, leaves many excellent candidates to fill too few spots. I would also add that choosing the best among the crop of very good is not something that I have found academics are good at.  Certainly much of what comes to be recognized as the best fails to find much support much of the time. And this is especially so if the work is original or the applicant has an unusual pedigree.

Let me expand on this for a moment in the context of grad admissions. In assessing applicants we are going, reasonably enough, on track record (what else is there). But, I have found that what makes you a good undergrad is not necessarily what makes you a good grad student, nor is what makes you a good grad student necessarily what makes you a fecund researcher. Undergrads are, rightly, much more coddled. They are assessed by how well they solve problems that have answers in the back of the book (or, now, online). Grad students are largely assessed on how well they solve problems without available answers, as are researchers. But, in addition, the latter are judged by how well they find new problems and generate novel research. The best are those whose work create enlightenment that generates puzzlement and so underwrites many years of further work.  In my experience, the qualities that allow success in any one of these endeavors does not reliably signal success in any of the others. So, even if we are conscientious in reviewing the available material, it is not clear what we are looking for. And, I should add, an hour or two interview, does not, in my experience, add much useful information. Luckily, how well we decide does not really matter as in a large number of cases any selection will be ok. We don’t need to find the best of the bunch because even the “worst” of the “best” is very good.

Another reaction: I was surprised at how wasteful the process can be! See point 7 above. 400 years is a long time, even if there are collateral benefits of writing a grant that does not get funded or applying to schools that don’t accept you or writing papers that never see the light of day. And I would question how large these collateral benefits actually are. The claims that there are benefits of doing the work even if it is unrewarded often strike me as self-serving (often touted by those at the top of the food chain). I can see the benefits of thinking through what one is working on as a valuable exercise. What I am less convinced of is doing this in a grant application format or a grad student admissions form is a good way of thinking things through. Maybe it is. But it is not obvious to me that it is clearly a superior way of doing this useful activity. And if it is not very useful, then it is questionable given that it clearly is very time consuming.

I have made noises echoing point 4 on FoL many times. I think that this is especially true for “theoretical” work where laying out the timeline of research is a silly exercise. If a problem is really good, then how you plan to solve it is not terribly apparent. The best you can do is motivate a conjecture, and funding agencies (or at least the NSF) finds this very insufficient. I suspect that this is so beyond theory. In my experience, grants seldom fund the research proposed. Rather a grant submission often presents finished research which if funded is ready to go “public” and the generated new grant is used to fund novel research that is as yet indeterminate (and only lightly described in the proposal). Ok, maybe I am being too cynical here. But I don’t think I am being way too cynical. At any rate, SA seems to agree.

Last point wrt 8/9 and then the “solution.” I also agree that the current process is unstable in that choose your reviewers and things will change dramatically. This suggests that if there really are objective bases of quality that could be used to rank alternatives then this factor is epistemologically elusive, at least over a large domain. Moreover, this elusiveness allows for systematic bias to creep in part because the process is believed to squeeze the bias out by focusing on “quality” alone.  Sadly, we all know how this works. The point AH makes is that it might be working more effectively in a context in which we are aiming to choose the best.[2]

Ok, the solution:

“Fortunately, there’s a simple solution to many of these problems. We should tell the experts to stop trying to pick the best research. Instead, they should focus on filtering out the worst ideas, and admit the rest to a lottery. That way, we can make do with shorter proposals, because the decision to accept or reject a ticket to a random draw requires less information – and highly specific proposals are unrealistic anyway. So instead of asking reviewers to make unreasonable predictions, they can turn their minds to weeding out cranks and frauds. Bias will still occur in the filtering stage, of course, but many more proposals will make it through to a lottery, which is inherently unbiased.”

I confess to finding this idea appealing. I believe that weeding out the clearly undeserving is a lot easier than identifying the best. So, I believe, a lottery among the triaged could save lots of time and effort. It might also, as AH notes, eliminate some bias, against unconventional ways of thinking, biases of various sorts and allow some things outside the prevailing fashion (though not too far out, I suspect).

But I think that there is a possible second very salutary effect of adopting a lottery system. It might force academics to acquire a bit more modesty. Academics are natural meritocrats. We reward the rewarded and denigrate the failures. We tend to downplay how much luck plays in the process of “success.” One nice feature of a lottery system is that it will make weaken this fantasy by making crystal clear that luck plays a non-negligible role. It does this by institutionalizing luck as an overt feature of the process. This might also have the beneficial side benefit of forcing administrators to consider more refined ways of measuring academic quality than just counting entries on a CV (weighted by venue quality). No more just listing publications in “major” journals or “large” grants funded. Maybe some thought will need to go into the process. Ok, I admit that there are problems with this too. But right now the mechanization of the whole acadmic evaluation process has, I believe, gotten out of hand and some push back is required. This could help by, as AH says, making it official, that the whole process is (in large part) a lottery anyhow. So not only might recognizing this and institutionalizing it, make “the whole process… cheaper, fairer and more efficient.” It might also help to make it more honest, with the attendant benefits that honesty often promotes.



[1] In the case of journal space, web publishing offers a plausible solution to the “too few pages” problem. However, one benefit of the page limits is that it helps manage information overload for the reader. If page numbers expand then the selection of “quality”/”relevance” will be off loaded to the reader. Right now many rely on editors and journals to cull the good stuff. It is unclear that it works well, but managing the tidle wave of material out there is an important problem.
[2] I don’t know if the review process in linguistics leads to obvious bias, but I would not be surprised if it did. I just don’t know. Anyone with info to share is invited to do so.

A derivation "towards LF"? Hardly. (Lessons from the Definiteness Effect.)

A while ago, I took part in a very interesting discussion over on Linguistics Facebook. The discussion was initiated by Yimei Xiang, who was asking the linguistics hivemind for examples that demonstrate how semanticists and syntacticians approach certain problems differently. I chimed in with a suggestion or two; but the example that I find most compelling came from Brian Buccola, who brought up the Definiteness Effect.

At first approximation, the Definiteness Effect refers to the fact that low subjects in the expletive-associate construction allow only a subset of the determiners that are allowed in canonical subject position (the observation goes back to Milsark's 1974 dissertation):

(1) There was/were {a/some/several/*the/*every/*all} wolf/wolves in the garden.

Now, I must admit I'm not as familiar as I should be with the semantic literature on this topic. What I do know is that there is a long tradition (going back to Milsark himself) of attributing this effect in one way or another to "existential force." The idea is that sentences like (1) assert the existence of a wolf/wolves who satisfy the predicate in the garden, and that we should seek an explanation of the Definiteness Effect in terms of the (in)compatibility of the relevant determiners with this assertion of existence.

This is plainly wrong, as we will see shortly. But since this is an unusually narrow thing to be writing about on Norbert's blog, let me say a bit more about why I think this is an interesting/illuminating test case.

There's a persistent intuition, which has pervaded work within the Principles & Parameters / Government & Binding / Minimalist Program tradition, whereby syntax is a derivation "towards LF." In other words, insofar as syntax has a telos, that telos is assembling a structure to be handed off to (semantic) interpretation. The other interface, externalization to PF, is something of an "add-on." (See, for instance, this call for papers, which takes this (supposed) asymmetry between LF and PF as its point of departure.)

Now, the general claim that LF has a privileged role might be correct regardless of what we find out about the true nature of the Definiteness Effect in particular. But the only way I can envision reasoning about the general claim is by looking carefully at a series of test cases until a coherent picture emerges. With that in mind, it is interesting to consider the Definiteness Effect precisely because it looks, at first blush, like an instance where semantics is "driving the bus": a semantic property (the interaction of existential force with a certain class of determiners) dictates a syntactic property (where certain noun phrases can or cannot go). What I'd like to show you is that this is not actually how the Definiteness Effect works.

The crucial data come from Icelandic (I know, try to contain your shock). One important way in which Icelandic differs from English is that the element that will move to subject position, in the absence of an expletive or some other subject-position-filling element, is simply the structurally closest noun phrase – regardless of its case. To see why this matters, let's start with a sentence like (2). This sentence behaves the same in Icelandic as it does in English, but it forms the baseline for the critical case, later on.

(2) There seems to be {a/*the} wolf in the garden.

In ex. (2), the noun phrase [DET wolf] is part of the infinitival complement to seem, rather than in post-copular position as it is in (1). But the semantics-based explanation of the Definiteness Effect cannot afford to treat the similarity between (1) and (2) as a coincidence if it has any hope of remaining viable, and so whatever one says about "existential force" in (1) must extend to [DET wolf] in (2).

Now consider sentences of the form in (3), an English version of which is given in (4):

(3) EXPL seems [DET1 experiencer (dative)] to be [DET2 thing (nominative)] in the garden.

(4) There seems to the squirrels to be {a/*the} wolf in the garden.

The semantics-based explanation must now extend the same treatment to [DET2 wolf] in (4). But here's where things start to go awry. In Icelandic, it is not DET2 that is subject to the Definiteness Effect in a structure like (3); the restriction in Icelandic affects DET1, while DET2 can be whatever you want.

Here's some Icelandic data showing this, from Sigurðsson's (1989) dissertation:


In exx. (14a-b) we see that, in the absence of a dative experiencer, Icelandic behaves like English (cf. (2), above): the Definiteness Effect applies to the nominative subject of the embedded infinitive. In exx. (15a-b), however, we see that when there is a dative experiencer (mér me.DAT), it is the experiencer that is subject to the Definiteness Effect, whereas the nominative subject of the experiencer can now be definite (barnið child.the.NOM) even while remaining in its low position.

Where does this leave the semantics-based explanation? Insofar as "existential force" is responsible for the Definiteness Effect, it has to be the case that existential force shifts from the downstairs subject to the dative experiencer only when the experiencer is present (cf. (14b) vs. (15a)), and only in Icelandic (not in English; cf. (4) vs. (15a)). This seems to me like a reductio ad absurdum of the "existential force" approach.

There is a much simpler, syntax-based alternative to all of this. The Definiteness Effect can be accounted for if any DP headed by a strong determiner must attempt to move to subject position (even 'the garden' in (1-2, 4) must attempt to do so!). The rules on what can actually successfully move to subject position in different languages are different, and, consequently, the question of which noun phrases can and cannot be definite in-situ will have different answers in different languages, too.

Another thing to note is that the expletive plays no role, here. The ungrammatical variants of (1-2, 4, 14, 15), showing the Definiteness Effect, all have expletives in them. But if your language happens to allow other things – e.g. an adjunct like 'today' – to occupy the preverbal subject position, you can get the same effect with no expletive at all. Here's some more Icelandic data, this time from Thráinsson (2007), showing this:


On the syntactic approach to the Definiteness Effect sketched above, what these data mean is that adjuncts get to move to preverbal subject position only if no nominal has done so; meaning if there is a nominal that must attempt movement to subject position (as all nominals headed by strong determiners must do), and which is in a position where such an attempt would be successful (e.g. allir kettirnir all cats.the.NOM in (6.52d)) – it preempts any adjunct from being able to move there. Definiteness Effect sans expletives.

So the narrow take-home message is that the Definiteness Effect has nothing to do with "existential force." What does this mean for the relation between syntax and semantics? Obviously, weak and strong determiners do differ semantically; that is a truism. But a noun phrase will not exhibit the Definiteness Effect unless it is a position where it is a candidate for movement-to-subject. And which positions these are is a matter that is subject to morphosyntactic variation of a kind that has nothing to do with semantics. Basically, some noun phrases bear a diacritic that forces them to try to move to subject position; whether they bear this diacritic or not seems to be grounded in an interpretive property (strong vs. weak determiners); but whether they succeed in moving or not seems to have no effect on their interpretation (see, e.g., barnið child.the.NOM in (15a), happily interpreted as definite in its low position). The Definiteness Effect, then, is not about semantics except insofar as the presence of the diacritic [+must try to move to subject position] is semantically grounded. Hardly a derivation "towards LF"; the diacritic must be present on the relevant determiners to begin with, before (the relevant part of) the derivation even starts.

What are the consequences of this for the broader question concerning LF as the telos of the derivation? In one sense, not much: this is but one case study; so it turns out that the Definiteness Effect does not match the relevant profile of LF-as-telos. It just means one less entry in the relevant column. Not exactly earth-shattering, there. But in another sense, I think the profile of the Definiteness Effect is the norm, not the exception: syntax pays attention to some (interestingly, not all) semantic distinctions, but it in no way "serves" those distinctions. Definite noun phrases don't move "in order to achieve a definite interpretation" – they just move or don't move (perhaps in a way that depends on their definiteness and the morphosyntactic properties of the language in question), and then they are interpreted however they are interpreted, regardless of where they ended up. I've argued that the exact same thing is true of the relation between specificity and Object Shift. And I suspect that the same is true of almost every single case in which it looks like syntax "serves" interpretation: it is an illusion. There are certain syntactic features that are interpretively grounded (definiteness, specificity, plurality, person features, etc.); and these features can drive certain syntactic operations. But what the syntactic derivation is doing is not constructing a representation that more closely matches the target interpretation. It's doing its own thing. Sometimes this will line up with semantic properties of the target interpretation – like "existential force" – but in those narrow instances where it does, it's really something of an accident. The next time someone tells you that syntax is about constructing "meaning with sound," take it with a boulder of salt.

––––––––––––––––––––

UPDATE: As some commenters (esp. Ethan Poole over on facebook, and David Basilico down here in the comments) have pointed out, the data in (14-15) are confounded in some non-trivial ways. I was attempting to be cute and show that the essential observations have been around for close to 30 years – which I still believe to be true – but I now see that it would also have been helpful to include some less confounded data. In service of this, here is some data from my own 2014 monograph that hopefully demonstrates the same points more clearly:



Wednesday, April 5, 2017

For those with free time; an excellent way to fill it

David Poeppel and Michael Gazzaniga put together a great ideas symposium at the last CNS meetings. I heard about the even from Ellen Lau who thought it great fun, as well as provocative and instructive.  Boy, was she dead on. I watched the presentations (here) and there is a little commentary and useful links here. The latter also links to the videos if you want one stop shopping. The videos are short, 20 minutes, and so easily watchable. I particularly urge you to look at the first two, Gallistel and Ryan and the last one by Krakauer (I really liked it and he was a hoot). But they are all excellent and given a good sense of what is going on now.

A word about Randy's presentation: the thing that struck me (again) is how coherent the picture he is painting is. There are is a deep story here and it starts from what seem innocuous starting points but quickly get one into deep waters. Curiously (and importantly), both he and Ryan (who seems to disagree with most everything Randy put forward (though I really didn't see how he did or even that he did))agree that the idea that memory is coded in synaptic weights in neural nets is a DEAD idea. Ryan noted that it has been categorically shown to be wrong. There may be something to synaptic connections, but it is NOT weights and adjustments to them. This seems like a real big deal to me if this is the cog-neuro consensus now. It puts a very deep nail into standard connectionist-associationist conceptions of mind/brain.

A few more remarks: First, it looks like modularity is back big time. Everyone buys into it. Krakauer noted that even (especially) the deep learning people are buying into this big time, shut when cog-neuro types were running away from the idea. He notes the irony here. But it looks like this idea is back and everyone loves it again.

Second, instinct is back and so are biologically rich constraints on learning. See Krakauer again on "cost functions" and how they are intrinsic to the systems and where all the action is. Yes, I was nodding my head in agreement.

So modules, innate cost functions, classical computations and even recursion. All topics here and well received. It is a good time to look to tie GG to cog-neuro. Watch the videos and enjoy.

Monday, April 3, 2017

A weapon of math destruction

A very interesting pair of papers crossed my inbox the other day on the perils of a statistical education (here1 and here2). The second paper is a commentary on the first. Both make the point that statistical training can be very bad for your scientific mindset. However, what makes these papers interesting is not that very general conclusion. It is easy for experts to agree that a little statistics (i.e. insufficient or “bad” training) can lead one off the path of scientific virtue. No, what makes these papers interesting is that it focuses on the experts (or should-be experts) and shows that their statistical training can “blind [them] to the obvious.” And this is news.[1] For it more than suggests that the problem is not with the unwashed but with the soap that would clean them. It seems that even in the hands of experts the tools systematically mislead.

Even more amusing, it seems that lack of statistical training leaves you better off.[2] So statistically non-trained boobs better succeeded on the tests that the experts flubbed. As Gelman and Carlin (G&C) pithily puts it: “it’s not that we are teaching the right thing poorly; unfortunately, we’ve been teaching the wrong thing all too well.”

Imo, though a great turn of phrase (and one that I hope to steal and use in the future) this somewhat misidentifies the problem. As both papers observe, the problem is that the limits of the technology are ignored because there is a strong demand for what the technology cannot deliver. In other words, the distortion is systematic because it meets a need, not because of mis-education. What need? I will return to this. But first, I want to talk about the papers a bit.

The first paper (McShane and Gal (M&G)) is directed against null hypothesis significance testing models (NHST). However, as both M&G (and G&C seconds this) makes clear, the problem extends to all flavors of statistical regimentation and inference, not just the NHST. What’s the argument?

The paper asks experts (or people that we would expect to be so and would likely consider themselves as such e.g. the editorial board members of the New England Journal of Medicine, Psychological Science, the American Journal of Epidemiology, the American Economic Review, the Quarterly Journal of Economics, the Journal of Political Economy) to answer some questions given a scenario that includes some irrelevant statistical numbers, half of which “suggest” the numbers reported are significantly different, some that they are not. What M&G finds is that these irrelevant numbers effect both how these experts report (actually, misreport) the data (and the inferences they (wrongly) draw from it:[3]

In Study 1, we demonstrate that researchers misinterpret mere descriptions of data depending on whether a p-value is above or below 0 05. In Study 2, we extend this result to the evaluation of evidence via likelihood judgments; we also show the effect is attenuated but not eliminated when researchers are asked to make hypothetical choices. (1709)

Moreover, as I gleefully mentioned, the statistically untrained are not similarly mislead:

We have shown that researchers across a variety of fields are likely to make erroneous statements and judgments when presented with evidence that fails to attain statistical significance (whereas undergraduates who lack statistical training are less likely to make these errors; for full details, see the online supplementary materials). (1714)

The question now becomes how the stats mislead. And the answer given is interesting. It misleads because p-values are treated as “magic numbers.” More particularly, M&G suggests that the problem lies with dichotomization:

…assigning treatments to different categories naturally leads to the conclusion that the treatments thusly assigned are categorically different. (1709)

G&C cites a related malady characteristic of the statistically inclined: “the habit of demanding more certainty than their data can legitimately supply” (1). Or, to put this another way: “statistical analysis is being asked to do something that it simply can’t do, to bring the signal from any data, no matter how noisy” (2). In short, the effectiveness of statistical methods have been “triumphantly” (this is G&C’s word) overhyped.

Say that this is true, why does this happen? There are clear extrinsic advantages to overhyping (career advancement being the most obvious (and the oft cited source of the replication crisis)) but both papers point to something far more interesting as the root of the problem; the belief that the role of statistics is to test theories and this is understood to mean to rule the bad ones out. Here’s how G&C frames this point.

To stop there, though, would be to deny one of the central goals of statistical science. As Morey et al. (2012) write, “Scientific research is often driven by theories that unify diverse observations and make clear predictions. . . . Testing a theory requires testing hypotheses that are consequences of the theory, but unfortunately, this is not as simple as looking at the data to see whether they are consistent with the theory.” To put it in other words, there is a demand for hypothesis testing. We can shout till our throats are sore that rejection of the null should not imply the acceptance of the alternative, but acceptance of the alternative is what many people want to hear. (4)

The “demand for hypothesis testing,” more specifically the idea that theory proposes but data disposes and that theories that fail the empirical test must be dumped (aka, falsificationism) requires a categorical distinction between theory and data. The latter is “hard” the former “soft,” the former is flighty and fanciful whereas the latter is clear and dispositive. The data are there to prevent theoretical hanky-panky and they can do their job only if they are not theoretically infected or filled with subjective bias. In other words, the sex appeal of stats to many is that it presents the world straight in a “just the facts ma’m” sort of way. On this view, the problem is not in the technology of stats, but in the idea that statistical training inculcates precisely the dichotomy between facts and hypotheses that appear to be systematically problematic.[4]

Let me put this another way: statistics is a problem because it generates numbers that create the illusion of precision and certainty. Furthermore, the numbers generated are the product of “calculation” not imaginative reasoning. These two features lead practitioners to think that the resulting numbers are “hard” and foster the view that properly massaged, data can speak for itself (i.e. that statistically curated data offers an unvarnished view of reality unmediated by the distorting lens of theory (or any other kind of “bias”)). This is what gives data a privileged position in the scientific quest for truth. This is what makes data authoritative and justifies falsificationism. And this (get ready for it) is just the Empiricist world view.

Furthermore, and crucially, as both M&G and G&C note, it has become part of the ideology of everyday statistical practice. It’s what allows statistics to play the role of arbiter of the scientifically valuable. It’s what supports the claims of deference to the statistical. After all, shouldn’t one defer to the precise and the unbiased if one’s interest is in truth? Importantly, as the two papers make clear, this accompanying conception is not a technical feature inherent to formal statistical technology or probabilistic thinking. The add-on is ideological. One can use statistical methods without the Empiricist baggage, but the methods invite misinterpretation, and seem to do so methodically, precisely because they fit with a certain conception of the scientific method: data vets theories and if theories fail the tests of experiment then too bad for them.  This, after all, is what demarcates science from (gasp!) religion or postmodernism or Freudian psychology or astrology or Trumpism. What stands between rationality and superstition are the facts, and stats curated facts are diamond hard, and so the best.

In this context, the fact that statistical methods yield “numbers” that, furthermore, are arrived at via unimaginative calculation is part of the charm. Statistical competence promotes the idea that precise calculation can substitute for airy judgment. As Trollope put it, it promotes the belief that “whatever the subject might be, never [think] but always [count].” Calculating is not subject to bias and thus can serve to tame fancy. It gives you hard precise numbers, not soft ideas. It’s a practical tool for separating wheat from chaff.

The only problem is that it is none of these things. Statistical description and inference is not merely the product of calculation and the numbers it delivers are not hard in the right way (i.e. categorical). As M&G notes, “the p-value is not an objective standard” (1716).

M&G and G&C are not concerned with these hifalutin philosophical worries. But their prescriptions for remedying the failures they see in current practice (e.g. the “replication crisis”) nonetheless call for a rejection of this naïve Empiricist worldview. Here is M&G:

…we propose a more holistic and integrative view of evidence that includes consideration of prior and related evidence, the type of problem being evaluated, the quality of the data, the effect size, and other considerations. (1716)

Similarly G&C note that it’s harder than believed to test a theory, because the theories of interest are “open-ended” and so it requires judgment to decide how far the calculations relevant to a particular statistical hypotheses reflect the basic ideas of the underlying scientific theory.

There is a larger problem of statistical pedagogy associating very specific statistical “hypotheses” with scientific hypotheses and theories, which are nearly always open-ended.

In other words, both papers argue that statistical methods are not a substitute for judgment but one factor relevant to exercising it. Amen! The facts do not and cannot speak for themselves, no matter how big the data sets, no matter how carefully curated, no matter how patiently amassed.[5]

Let me end on a point closer to home. Linguists have become enthralled with stats lately. There is nothing wrong with this. It is a tool and can be used in interesting ways. But there is nothing magical about stats. The magical thinking arises when the stats are harnessed to an Empiricist ideology. Imo, this is also happening within linguistics, sometimes minus the stats. I have long kvetched about the default view of the field which privileges “data” over “theory.” We can all agree with the truism that a good story needs premises that have an interesting deductive structure and that also has a non-trivial empirical reach. This is a truism. But none of this means that the facts are hard and the theory soft or that counter-examples rule while hypotheses drool. There are no first and second fiddles, just a string section. A question worth asking is whether our own preferred methods of argumentation and data analysis might be blinding us to substantial issues concerning our primary aim of inquiry: understanding the structure of FoL. As you probably have guessed, I think that there is. GG has established a series of methods over the last 60 years that we have used to investigate the structure of language and, more importantly, FL. It is very much an open question, at least to me, whether these methods are all still useful and how much rich linguistic description reveals about the underlying mechanisms of FL. But this is a topic for another time.



[1] Well, sort of. The conclusion is similar to the brouhaha surrounding the Monty Hall Problem (here) and the reaction of the cognoscenti to Marilyn Vos Savant’s correct analysis. She was pilloried by many experts as an ignoramus despite the fact that she was completely correct and they were not. So this is not the first time that probabilistic reasoning has proven to be difficult even for the professionals.
[2] This is unlike the Monty Hall Problem.
[3] There was a paper floating around a number of years ago which showed that putting a picture of an fMRI colored brain on a cog-neuro paper enhanced its credibility even when the picture had nothing to do with any of the actually reported data. I cannot recall where this paper is, but the result seems very similar to what we find here.
[4] Or maybe those mose inclined to become expert are one’s predisposed to this kind of “falsificationist” world view.
[5] M&G notes in passing that what we really want is a “plausible mechanism” to ground our statistical findings (1716), and these do not emerge from the data no matter how vigorously massaged.