Comments

Showing posts with label Stats. Show all posts
Showing posts with label Stats. Show all posts

Friday, June 8, 2018

Science without theory

Sometimes we just don't know much and this puts us in an odd epistemic position. Not knowing much comes with an imperative of intellectual modesty: one should have relatively little confidence in one's descriptions of what is going on and even less in projections of what might happen counterfactually. Being ignorant is a real bummer, especially scientifically.

Now, all of this should be obvious. And to many it is. But is is also something that working scientists have a professional interest in "bracketing" (a terms that roughly means "setting aside" that I learned as a philo grad student and that has come in very handy over the years as it sounds so much better than "ignore" (something an honest intellectual should not do (purportedly)) which is more or less what it amounts to), and so, not surprisingly, they largely do. Moreover, as nobody gets kudos for advertising their ignorance ((well, there was Socrates, I guess) or less kudos and occasionally a hemlock milkshake), scientists, especially given the current incentive system, are less restrained in making the flimsiness of their conclusions apparent than perhaps they should be. And this is a problem, for it is really hard to say anything useful when you have no idea what the hell is going on.

Why do I mention this? Rob Chametzky sent me a recent paper on this topic (here). The post (the author is Denny Borsboom (here), so I will dub the post DB) makes the reasonable point (reasonable to me as I have been making it as well for a while now) that the absence of "unambiguously formalized theory in psychology" lies behind much of the "replication" crisis in psych (p.1).[1]It, moreover, suggests that this is more or less endemic to the discipline because“psychology has the hardest subject matter ever studied”(p. 3). I do not know whether this last point is correct, but it is certainly true that much of what psychologists insist on studying is almost surely the product of many many interacting systems (e.g. almost any social psych topic!) and it is well-known that interaction effects are very hard to disentangle, very very hard. So, it is not surprising that fit hese topics are the sorts of things that psychologists study then the level of non-trivial theory that exists is close to zero. That is what one would expect (and what one finds). DB traces out some of the implications of this for the replication crisis. Here are some consequences:

·      The field is heavily stats dependent, as stats methods substitute for theoretical infrastructure.
·      The role of stats can grow so great as to induce theoretical amnesia on the practitioners (a mental state wherein those in the field “no longer know what a theory is”(p.2).
·      Progress in atheoretical psych is necessarily very slow given that experiments are always trying to factor out poorly understood context dependent variables.
·      The discipline is susceptible to “fads”based on “poorly tested generalizations”that serve to make research manageable (at best) and allows for a kind of predatory free-riding (at worst).

Needless to say, this is not a pretty state of affairs. The solution for this? Well, of course, more care with the stats and a kind of re-education system for psychologists: “It would be extremely healthy if pshychologists received more education in fields which do have some theories, even if they are empirically shaky ones”so that the discipline can “try to remember what a theory is and what it is good for, so that we don’t fall into theoretical amnesia”(p.3). 

I cannot say that I find DB's description all that far off base. However, I think that a few caveats are in order. Here are some.

First, why is theoretical amnesia a bad thing if the field is doomed to be forever theoryless given its endemic difficulty? It is useful to understand how theory functions in a field where it does so usefully if this utility can be imported into one's own. Then being able to recognize it and value it is important. But if this is impossible then why bother?  

I suspect that DB's real gripe is that there is theory to be had (maybe by reshaping the topics studied) but that psychologists have been trained to ignore it and to substitute stats methods for theoretical insight. If this DB's point, then it is an important one. And it applies to many domains where empirical methods often overrun their useful boundaries. If this is DB's point, then there is a better way to put it: stats, no matter how technically fancy, cannot substitute for theory. Or, to put this in lay terms: lots of data carefully organized is not a theory, and thinking it is is just a confusion.

Second, the general point that RB makes (and I agree with) is not at all idiosyncratic. Gellman has made the point here (again) recently. The last paragraph is a good summation of RB's basic point:

…hypotheses in psychology, especially social psychology, are often vague, and data are noisy. Indeed, there often seems to be a tradition of casual measurement, the idea perhaps being that it doesn’t matter exactly what you measure because if you get statistical significance, you’ve discovered something. This is different from econ where there seems there’s more of a tradition of large datasets, careful measurements, and theory-based hypotheses. Anyway, psychology studies often (not always, but often) feature weak theory + weak measurement, which is a recipe for unreplicable findings.

Having little theory is real problem even if one's aim is to get a decent stats description of the lay of the land. The problems one finds in theoryless fields is what one should expect and the methodological sloppiness comes with the territory. As Gellman puts it:

p-hacking is not the cause of the problem; p-hacking is a symptom. Researchers don’t want to p-hack; they’d prefer to confirm their original hypotheses. They p-hack only because they have to.

If this is right, however, I think that both Gellman's and RB's optimism that this can be solved by better methodological hygiene is probably unfounded. The problem is that there is a real cost to not knowing what the hell is going on, and that cost is not “merely”theoretical but “observational”as well. Why? 

Here is a good place to play the Einstein card (roll the drums): theory is implicated in determining what counts as observational! Here's a quote: “It is the theory which decides what we can observe.”[2]To know what to count and how to count it (that's what stats does) you need a way of determining what should be counted and how (that's what theory does). So, no theory, no observations in the relevant scientific sense either. And if this is so, then when you really have no idea what the hell is going on, then you are in deep doo-doo whether you know it or not. Can stats help? I doubt it. Being very careful and very cautious might help. But really the only thing to do in these cases is pray to the God of scientific traction for a bit of luck in getting you started. There is a reason why researchers who figure out anything are heros. They are the people whose ideas allow us to get the ball rolling. Once it's rolling it's a whole new game. Until then. Nada! So, I doubt that the techno optimism that Gellman and RB point to, the idea that sans theory stats can step in and allow us to do another kind of useful science, will really fly. But, then I am a pessimist in general.

Third, I don't think that what holds for social psych is characteristic of the whole endeavor. There are large parts of psych broadly understood (e.g. large parts of learning theory (see Gallistel's work on this), development (see Carey and Spelke and Baillargeon and R. Gelman a.o. for example), perception, math capacities, language, where there is quite a bit of decent non-trivial theory that usefully guides inquiry). The problem is that social psych is where the fame and money are. You get on NPR for work on power poses, but not on Weber's law. 

Fourth, this is really not the state of play in most of linguistics. We really do have some decent theory to fall back on in many parts of the core discipline (syntax, phonology, parts of semantics) and that is why we have been able to make non-trivial progress. The funny thing is that if RB and Gellman are correct, the brouhaha over linguistic data foisted upon the field by the stats inclined has things exactly backwards for if they are right the methods adopted by fields without any theoretical ideas are bad models for those that have some theoretical sub-structure.

Fifth, what RB and Gellman describe is really what we should expect. There is a long-standing hope that there exists a mechanical way of doing science (i.e. gaining insight). If we were just careful enough in how we gathered data, if we only got rid of our pre-conceptions, if only our morals were higher, we could just look and see the truth. The problem this simple method fails is because we don’t do it right. This is, of course, a reflection of the old Empiricist dream (see the previous post for the logical Positivist version of this). It repeatedly fails. It will always fail. 

That’s it. Thx to Rob for sending me the RM piece. 


[1]The adjective “formalized” actually understates the problem. There is precious little non-trivial theory in large parts of psych (especially social psych, the epicenter of the crisis), formalized or not. The problem is not “formalization”per se.  
[2]Quoted in What is realby Adam Becker, p. 29. The philosopher Gorver Maxwell made a similar point oh so many years ago: “It is theory…which tells us what is or is not…observable” (Becker 184). If this is right, then the idea of grounding science on some a prioriconception of the observable independent of any theoretical assumptions is pretty much a non-starter. We indulge in “theory” either explicitly or tacitly. The optimism arises, I believe, by ignoring this or by hoping that for much of what we look at the tacit theory is pretty solid. Given that such theory will, when inexlicit, revert to “common sense,” this hope strikes me as idle given that common sense is precisely what scientific insight almost always overturns.

Monday, November 5, 2012

On Norvig on Chomsky on Stats




Stats can be used in (at least) three ways in linguistics; to test theory, as part of theory or in place of theory.  I would like to briefly discuss each in reference to Peter Norvig’s piece on stats in language, which I just recently (re)read.  For discussion of some of these issues by Chomsky see here.

Nobody could object in principle to the first use.  To date, the primary tool for investigating grammaticality has been native speakers’ acceptability judgments. Though often called “grammaticality judgments” (e.g. Norvig follows fashion in failing to distinguish ‘grammaticality’ (a theoretical term) from  ‘acceptability’ (a term describing judgment data), linguists have long known that such data reflect more than just grammatical concerns.  Sentences can be more or less acceptable depending on various factors (e.g. sentence length, presence of long distance dependencies, frequency of the lexical items used, conventionality of the thought expressed, etc.) only one of which is the grammatical structure of the sentence.  As has also long been known, grammatical structures need not be very acceptable (e.g. that that that Bill kissed Mary is fine is certain is deplorable) and acceptable sentences need not be grammatical (e.g. more people visited Rome last year than I did).  Nonetheless, acceptability judgments have proven to be reliable empirical probes for grammatical structure, despite the occasional disagreement about the status of one or another sentence among practitioners.  Nobody could object in principle to using more refined stats based techniques to monitor the reported data (this is done all the time in psycholinguistics for various quality control reasons), though as a matter of fact it appears to be largely overly fastidious (if not a waste of time and effort) to do so as the informal methods linguists have long used have proven to be wildly reliable (c.f. Sprouse &Almeida for a good review).[1]

The second role for stats is as a defining feature of a theory.  For example, probabilistic context free grammars (PCFG) are all the rage nowadays in parsing.  They provide a way of systematically combing frequency effects with grammatical parsing. Again, there can be no principled objection to this though reasons for mixing , in my view, are often misinterpreted.  For example, adding statistical features to grammars to allow them to be sensitive to various frequency properties within a sentence or discourse context does not imply that grammatical competence, (viz. a speaker’s grammar), is inherently statistical.  All it implies is that parsing, a process that uses a speaker’s linguistic knowledge to assign a structure to an incoming phonetic string, can use statistically gleaned information to facilitate this task.  Norvig suggests that the variability of judgment data indicates that grammars must be statistical (though I may be over-interpreting here as Norvig does not distinguish between the disparate features that contribute to linguistic behavior).  For him language is one big messy “complex, random, contingent” thingy subject to the “whims of evolution and cultural change” and so must be “analyzed with probabilistic models.”[2] Norvig might mean by this that grammars must be inherently probabilistic (confession: I don’t have the foggiest idea what he really means beyond his conviction that language use in all its glory is very complicated, something with which I agree, but from which I draw rather different conclusions) but this conclusion simply does not follow from the variability of language use, even variability in acceptability judgments, for one can model variability as an interaction effect of the various components that go into making an acceptability judgment and still keep the grammar completely categorical. 

Further, it is hard to detect the workings of probability in large parts of the acceptability data.  For example, there is in general no uncertainty about the products of the grammar.  Native speakers know with probability 1 that John loves Mary does not mean Mary loves John, that John is eager to please means that John is the pleaser while in John is easy to please John is the pleasee (and that John cannot be pleasee in the ‘eager’ case nor the pleaser in the ‘easy’ case), that who kissed Mary cannot be appropriately answered with Mary kissed Sue, that he thinks that everyone is tall does not have a paraphrase everyone thinks that he is tall, etc.  Speakers know these facts with certainty, which is not to say that they might not misspeak. But should they slip up, speakers will acknowledge as much when it is pointed out to them. What speakers don’t do is retort 18% of the time “oh no, I really meant John to be the pleasee when I said John is eager to please”. It is very hard to make the argument (at least at present though not in principle) that variable judgments and shifting linguistic behavior implies that grammars must be statistically loaded. Though grammars might be adorned with various probabilistic doo-dads to date there is no logical (or even overwhelming empirical) road starting at linguistic variablilty and ending with statistical grammars.  Addressing the issue requires (at least) carefully distinguishing degree of grammaticality from degree of acceptability (Norvig, recall runs these two concepts together) and showing that the analysis of the second requires something like the first. To date, I know of nothing that speaks to this convincingly, though should it prove correct (as many including Chomsky have speculated over the years) that would be fine with me.[3]

The third use of stats is pernicious.  It expresses a belief, one that I believe Norvig shares, that statistical analysis of lots of data can substitute for  of theory construction.  This represents a return of a certain unfortunate empiricism, both in its scientific methodology and its susceptibility to associationist conceptions of mind.  The program of statistically massaging enormous amounts of data (SMEAD) stands in place of looking for underlying causal powers of surface phenomena. As noted in an earlier post, this view has a home in classical Empiricism, which takes scientific inquiry as the hunt for regularities that describe behavior rather than the powers/natures/capacities whose complex interaction generate it.  Empiricist methodology swerves into associationist psychology when what is sampled and statistically measured is unanalyzed experiential data (rather than experimentally created phenomena, see below), the idea being that fast machines with vast memory resources can crunch lots of very simple surfacy data and in so doing uncover underlying causal mechanisms.  This is classical empiricist dogma and in the (what I hope will soon become) the immortal words of Sydney Brenner it’s “a form of insanity,” a species of “low input, high throughput, no output science.” 

There are at least two things wrong with SMEAD. First, it very much misconstrues the relation between data and theory.  And it does so in two ways.

First, scientific observation takes a lot of very careful stage setting. As Cartwright notes:

Outside the supervision of a laboratory or the closed casement of a factory-made module, what happens in one instance is rarely a guide to what will happen in others.  Situations that lend themselves to generalizations are special…(86).

In other words, most interesting data is manufactured, not natural.  Phenomena are factitious and are barely visible in the interactive swamp that constitutes regular experience. The scientific attitude rests not on observing the world but on putting it in a very tight straightjacket and then prodding and poking (if not worse; how would you like to sent flying at velocities near the speed of light to smash into a target possibly moving at YOU at a similar speed?)) until it reveals some small piece of useful data.  Nancy Cartwright sums this up well:

For anyone who believes that induction provides the primary building tool for empirical knowledge, the methods of modern experimental physics must seem unfathomable.  Usually the inductive base for the principles under test is slim indeed,…and in the best experimental designs…one single instance can be enough…Clearly in these physics experiments we are prepared to assume that the situation before us is of a very special kind: it is a situation in which the behavior that occurs is repeatable.  Whatever happens in this situation can be generalized (85).

This is no less true in linguistics than it is in any other domain of inquiry with theoretical aspirations.  Linguists have been luckier than most in that acceptability judgments (a very crude form of experiment) have been reliable probes to the structure of UG.[4]  Linguistic judgments made in “reflective equilibrium” seem capable of abstracting away from many interfering factors and allow a stable basis for theory construction.  In contrast to simply looking at linguistic behavior, judgment queries can be targeted to structures of interest (you don’t have to wait for someone to say what you are interested in), can alleviate attention and memory pressures (leading to slips of the tongue and “brainos”), and, most importantly, can query the status of unacceptable structures.  Within linguistics it is more often than not the dog that doesn’t bark that is the key to figuring out the fine structure of UG. 

Second, ‘data’ is a misleading term for what scientists use to advance insight for it suggests small isolated points, like data on a graph ready for statistical smoothing.  What drives understanding in the sciences are ‘effects’ and anomalies: e.g. the ultraviolet catastrophe, the perihelion of mercury, the Doppler effect, the retrograde motion of Mars, the Compton effect, etc. Linguistics is blessed with a dozen or so of these effects; e.g. intervention effects, island effects, fixed subject effects, strong and weak crossover effects, filled gap effects etc. These effects are only visible with the greatest contrivance and detectable when mere “observation” has been left far behind.[5]

Last of all, without theory of some sort even stats can’t get off the ground.  Stats are fancy methods of counting. Theory tells you what to count.  Even simple data does not organize itself. Things are counted with respect to some properties. SMEAD appears to believe that large data sets can organize themselves, pull themselves up by their large data base bootstraps. This is false (recall: it is impossible to pull oneself up by the bootstraps). Categories are always necessary and in the absence of some prior theory, surface distinctions become the default categories and these have a tendency of breeding associationist sympathies.  Indeed, if your aim is to collect a lot of data quickly and crunch lots of it fast then visible surface distinctions are what will entice you.  Regularities/associations among the observables is just a step away.  If we have learned anything in the last 50 years, it’s that this method of research will teach us nothing of interest as regards any non-trivial cognitive capacity.

There is a myth abroad in the land that generative grammar has a principled antipathy to stats. Wrong. We have a principled antipathy to useless and misleading and pernicious stats and the technical virtuosos who think that counting anything and everything can replace hard thinking.  All the rest, like all other tools, must prove its worth empirically, which in this case means, can be used to shed light on the structure of UG.


[1] Sprouse and colleagues have argued that these techniques can move beyond quality control and provide a novel kind of data for the investigation of UG and its interactions. This, of course, is a welcome development.
[2] Compare Quine’s empiricist description: language is nothing but “a fabric of sentences variously associated to one another and to nonverbal stimuli by the mechanism of conditioned response.”  Substitute ‘frequencies’ (what Norvig’s probabilities track) for ‘conditioned response’ and the two seem animated by the same unfortunate ethos.
[3] Within syntax degree of grammaticality has been modeled in terms of simple counting: violating two conditions is worse than one, violating some constraints is more serious than others.  In systems this simple, stats are not required.
[4] I suspect that the utility of such a crude procedure speaks volumes about the robustness of the language instinct.
[5] Those interested in these issues might like to look at this in addition to Cartwright’s above mentioned work.