Comments

Showing posts with label acceptability vs grammaticality. Show all posts
Showing posts with label acceptability vs grammaticality. Show all posts

Tuesday, September 15, 2015

Judgments and grammars

A native speaker judges The woman loves himself to be odd sounding. I explain this by saying that the structure underlying this string is ungrammatical, specifically that it violates principle A of the Binding Theory. How does what I say explain this judgment? It explains it if we assume the following: grammaticality is a causally relevant variable in judgments of acceptability. It may not be the only variable relevant to acceptability, but it is one of the relevant variables that cause the native speaker to judge as s/he does (i.e the ungrammaticality of the structure underlying the string causes a native speaker to judge the sentence unacceptable). If this is correct (which it is), the relation between acceptability and grammaticality is indirect. The aim of what follows is to consider how indirect it can be while still leaving the relation between (un)acceptability and (un)grammaticality direct enough for judgments concerning the former to be useful as probes into the structure of the latter (and into the structure of Gs which the notion of grammaticality implicitly reflects).

The above considerations suffice to conclude that (un)acceptability need not be an infallible guide to (un)grammaticality. As the latter is but one factor, then it need not be that the former perfectly tracks the latter. And, indeed, we know that there are many strings that are quite unacceptable but are not ungrammatical. Famous examples include self-embedding (e.g. That that that Mary saw Sam is intriguing is interesting is false), ‘buffalo’ sentences (e.g. Buffalo buffalo buffalo buffalo buffalo buffalo buffalo) and multiple negation sentences (No eye injury is too insignificant to ignore). The latter kinds of sentences are hard to process and are reliably judged quite poor despite their being grammatical. A favorite hobby of psycho-linguists is to find other cases of grammatical strings that trouble speakers as this allows them to investigate (via how sentences are parsed in real time) factors other than grammaticality that are psycho-linguistically important. Crucially, all accept (and have since the “earliest days of Generative Grammar”) that unacceptability does not imply ungrammaticality.

Moreover, we have some reason to believe that acceptability does not imply grammaticality. There are the famous cases like More people visited Rome than I did which are judged by speakers to be fine despite the fact that speakers cannot tell you what they mean. I personally no longer think that this shows that these sentences are acceptable. Why? Precisely because there is no interpretation that they support. There is no interpretation for these strings that native speakers consistently recognize so I conclude form this that they are unacceptable despite “sounding” fine. In other words, “sounding fine” is at best a proxy for acceptability, one that further probing may undermine. It often is good enough and it may be an interesting question to ask why some ungrammatical sentences “sound fine” but the mere fact that they do is not in itself sufficient reason to conclude that these strings are acceptable (let alone grammatical).[1]

So are there any cases of acceptability without grammaticality? I believe the best examples are those where we find subliminal island effects (see here for discussion). In such cases we find sentences that are judged acceptable under the right interpretation. Despite this, they display the kinds of super-additivity effects that characterize islands. It seems reasonable to me to describe these strings as ungrammatical (i.e. violate island conditions) despite their being acceptable. What this means is that for cases such as these the super-additivity profile is a more sensitive measure of grammaticality than is the bare acceptability judgment. In fact, assuming that the sentence violates islands restrictions explains why we find the super-additivity profile.  Of course, we would love to know why in these cases (but not in many other island violating examples) ungrammaticality does not lead to unacceptability. But not knowing why this is so, does not in and of itself compromise the conclusion that sentences these acceptable sentences are ungrammatical.[2]

So, (un)acceptability does not imply (un)grammaticality, nor vice versa. How then can the former be used as a probe into the latter? Well, because this relation is stable often enough. In other words, over a very large domain acceptability judgments track grammaticality judgments, and that is good enough. In fact, as I’ve mentioned more than once, Sprouse, Almeida, and Schutze have shown that these data are very robust and very reliable over a very wide range, and thus are excellent probes into grammaticality. Of course, this does not mean that they such judgments are infallible indicators of grammatical structure, but then nobody thought that they ever were. Let me elaborate on this.

We’ve known for a very long time that acceptability is affected by many factors (see Aspects:10-15 for an early sophisticated discussion of these issues), including sentence length, word frequencies, number of referential DPs employed, intonation and prosody, types of embedding, priming, kinds of dependency resolutions required, among others. These factors combine to yield a judgment of (un)acceptability on a given occasion. And these are expected to be (and acknowledged to be) a matter of degree. One of the things that linguists try to do in probing for grammaticality is to compensate for these factors by comparing sentences of similar complexity to one another to isolate the grammatical contribution to the judgment in a particular case (e.g. we compare sentences of equal degree of embedding when probing for island effects). This is frequently doable, though we currently have no detailed account of how these factors interact to produce any given judgment. Let me repeat this: though we don’t have a general theory of acceptability judgments, we have a pretty good idea what factors are involved and when we are careful (and even when we are not, as Sprouse has shown) we can control for these and allow the grammatical factor to shine thorough a particular judgment. In other words, we can set up a specific experimental situation that reliably tests for G-factors (i.e. we can test whether G-factors are causally relevant in the standard way that experiments typically do, by controlling the hell out of the other factors). This is standard practice in the real sciences, where unpacking interaction effects is the main aim of experimentation. I see no reason why the same should not hold in linguistics.[3]

It is worth noting that the problem of understanding complex data (i.e data that is reasonably taken to be the result of many causally interacting factors) is not limited to linguistics. It is a common feature of the real sciences (e.g. physics). Geoffrey Joseph has a nice (old) paper discussing this, where he notes (786):[4]

Success at the construction and testing of theories often does not proceed by attempting to explain all, or even most, of the actually available data. Either by selecting appropriate naturally occurring data or by producing appropriate data in the laboratory, the theorist implicitly acknowledges …his decomposing the causal factors at work into more comprehensible components. A consequence of this feature of his methodology is that we are often in the position of having very well-confirmed fundamental theories at hand, but at the same time being unable to formulate complete deductive explanations of natural (complex) phenomena.

This said, it is interesting when (un)acceptability and (un)grammaticality diverge. Why? Because, somewhat surprisingly, as a matter of fact the two track one another so closely (in fact, much more closely than we had any reason to expect a priori). This is what makes it theoretically interesting when the two diverge. Here’s what I mean.[5]

There is no reason why (un)acceptability should have been such a good probe into (un)grammaticality. After all, this is a pretty gross judgment that we ask speakers to make and we are absolutely sure that many factors are involved. Nonetheless, it seems that the two really are closely aligned over a pretty large domain. And precisely because they are, it is interesting to probe where they diverge and why. Figuring out what’s going on is likely to be very informative.

Paul Pietroski has suggested an analogy from the real sciences. Celestial mechanics tracks the actual position of planets in space in terms of their apparent positions. Now, there is no a priori reason why a planet’s apparent position should be a reliable guide to its actual one. After all, we know that the fact that a stick in water looks bent does not mean that it is bent. But at least in the heavens, apparent position data was good enough to ground Kepler’s discoveries and Newton’s. Moreover, precisely because the fit was so good, their apparent divergence in a few cases was rightly taken to be an interesting problem to be solved. The solutions required a complete overhaul of Newton’s laws of gravitation (actually, relativity ended up deriving Newton’s laws as limit cases, so in an important sense they these laws were conserved). Note that this did not deny that apparent position was pretty good evidence of actual position. Rather it explained why in the general case this was so and why in the exceptional cases the correlation failed to hold. It would have been a very bad idea for the history of physics had physicists drawn the conclusion that the anomalies (e.g. the perihelion of Mercury) showed that Newton’s laws should be trashed. The right conclusion was that the anomaly needed explanation, and they retained the theory until something came along that explained both the old data and also explained the anomalies.

This seems like a rational strategy, and it should be applied to our divergent cases as well. And it has been to good effect in some cases. The case of unacceptability-despite- grammaticality has generated interesting parsing models that try to explain why self- embedding is particularly problematic given the nature of biological memory (e.g. see work by Rick Lewis and friends). The case of acceptability-despite-ungrammaticality has led to the development of somewhat more refined tools for testing acceptability that has given us criteria other than simple acceptability to measure grammaticality.

The most interesting instance of the divergence, IMO, is the case of Mainland Scandinavian where detectable island violations (remember the super-additivity effects) do not yield unacceptability. Why not? Dunno.[6] But, as in the mechanics case above, the right attitude is not that the failure of acceptability to track grammaticality shows that there are no island effects and that a UG theory of islands is clearly off the mark. Rather the divergence indicates the possibility of an interesting problem here and that there is something that we still do not understand. Need I say, that this latter observation is not a surprise to any working GGer? Need I say that this is what we should expect in normal scientific practice?

So, grammaticality is one factor in acceptability and a reliable robust one at that. However, like most measures, it works better in some contexts than in others, and though this fact does not undermine the general utility of the measure, it raises interesting research questions as to why.

Let me end by repeating that all of this is old hat. Indeed, Chomsky’s discussion in chapter 1 of Aspects is still a valuable intro to these issues (11):

…the scales of grammaticalness and acceptability do not coincide. Grammaticalness is only one of many factors that interact to determine acceptability. Correspondingly, although one might propose various operational tests for acceptability, it is unlikely that a necessary and sufficient operational criterion might be invented for the much more abstract and important notion of grammaticalness.

So, can we treat grammaticalness as some “kind” of acceptability? No, nor should we expect to. Can we use acceptability to probe grammaticalness? Yes, but as in all areas of inquiry there is no guarantee that these judgments are infallible guides. Should we expect to one day have a solid theory of acceptability? Well, some hope for this, but I am skeptical. Phenomena that are the result of the interactions of many factors are usually theoretically elusive. We can tie loose ends down well enough in particular cases, but theories that specify in advance which looses ends are most relevant are hard to come by, and not only in linguistics. There are no general theories of experimental design. Rather there are rules of thumb of what to control for in specific cases informed by practice and some theory. This is true in the “real” sciences, and we should expect no less in linguistics. Those who demand more in the latter case are methodological dualists, holding linguistics to uniquely silly standards.



[1] Other similar examples involve cases where linear intervening non-licensing material can improve a sentences acceptability. There is a lot of work on this involving NPI licensing by clearly non-c-commanding negative elements. This too has been widely discussed in the parsing literature. Again, interpretations for these improved sentences are hard to come by and so it is unclear whether these sentences are actually acceptable despite their tonal improvements.
[2] So far as I can tell, similar reasoning applies to some recent discussion of binding effects in a recent Cognition paper by Cole, Hermon and Yanti that I hope to discuss more fully in the near future.
[3] As Nancy Cartwright observes in her 1983 book (p. 83), the aim of an experiment is to find “quite specific effects peculiarly sensitive to the exact character of the causes [you] want to study.” Experiments are very context sensitive set ups developed to find these effects.  And they often fail to explain a lot. See Geoffrey Joseph quote below.
[4] See his “The many sciences and the one world,” Journal of Philosophy 1980: 773-791.
[5] I own what follows to some discussion with Paul. He is, of course, fully responsible for my misstatements.
[6] But I believe that Dave Kush and friends have provided the right kind of answer. See chapter 11 here for an example).

Monday, November 5, 2012

On Norvig on Chomsky on Stats




Stats can be used in (at least) three ways in linguistics; to test theory, as part of theory or in place of theory.  I would like to briefly discuss each in reference to Peter Norvig’s piece on stats in language, which I just recently (re)read.  For discussion of some of these issues by Chomsky see here.

Nobody could object in principle to the first use.  To date, the primary tool for investigating grammaticality has been native speakers’ acceptability judgments. Though often called “grammaticality judgments” (e.g. Norvig follows fashion in failing to distinguish ‘grammaticality’ (a theoretical term) from  ‘acceptability’ (a term describing judgment data), linguists have long known that such data reflect more than just grammatical concerns.  Sentences can be more or less acceptable depending on various factors (e.g. sentence length, presence of long distance dependencies, frequency of the lexical items used, conventionality of the thought expressed, etc.) only one of which is the grammatical structure of the sentence.  As has also long been known, grammatical structures need not be very acceptable (e.g. that that that Bill kissed Mary is fine is certain is deplorable) and acceptable sentences need not be grammatical (e.g. more people visited Rome last year than I did).  Nonetheless, acceptability judgments have proven to be reliable empirical probes for grammatical structure, despite the occasional disagreement about the status of one or another sentence among practitioners.  Nobody could object in principle to using more refined stats based techniques to monitor the reported data (this is done all the time in psycholinguistics for various quality control reasons), though as a matter of fact it appears to be largely overly fastidious (if not a waste of time and effort) to do so as the informal methods linguists have long used have proven to be wildly reliable (c.f. Sprouse &Almeida for a good review).[1]

The second role for stats is as a defining feature of a theory.  For example, probabilistic context free grammars (PCFG) are all the rage nowadays in parsing.  They provide a way of systematically combing frequency effects with grammatical parsing. Again, there can be no principled objection to this though reasons for mixing , in my view, are often misinterpreted.  For example, adding statistical features to grammars to allow them to be sensitive to various frequency properties within a sentence or discourse context does not imply that grammatical competence, (viz. a speaker’s grammar), is inherently statistical.  All it implies is that parsing, a process that uses a speaker’s linguistic knowledge to assign a structure to an incoming phonetic string, can use statistically gleaned information to facilitate this task.  Norvig suggests that the variability of judgment data indicates that grammars must be statistical (though I may be over-interpreting here as Norvig does not distinguish between the disparate features that contribute to linguistic behavior).  For him language is one big messy “complex, random, contingent” thingy subject to the “whims of evolution and cultural change” and so must be “analyzed with probabilistic models.”[2] Norvig might mean by this that grammars must be inherently probabilistic (confession: I don’t have the foggiest idea what he really means beyond his conviction that language use in all its glory is very complicated, something with which I agree, but from which I draw rather different conclusions) but this conclusion simply does not follow from the variability of language use, even variability in acceptability judgments, for one can model variability as an interaction effect of the various components that go into making an acceptability judgment and still keep the grammar completely categorical. 

Further, it is hard to detect the workings of probability in large parts of the acceptability data.  For example, there is in general no uncertainty about the products of the grammar.  Native speakers know with probability 1 that John loves Mary does not mean Mary loves John, that John is eager to please means that John is the pleaser while in John is easy to please John is the pleasee (and that John cannot be pleasee in the ‘eager’ case nor the pleaser in the ‘easy’ case), that who kissed Mary cannot be appropriately answered with Mary kissed Sue, that he thinks that everyone is tall does not have a paraphrase everyone thinks that he is tall, etc.  Speakers know these facts with certainty, which is not to say that they might not misspeak. But should they slip up, speakers will acknowledge as much when it is pointed out to them. What speakers don’t do is retort 18% of the time “oh no, I really meant John to be the pleasee when I said John is eager to please”. It is very hard to make the argument (at least at present though not in principle) that variable judgments and shifting linguistic behavior implies that grammars must be statistically loaded. Though grammars might be adorned with various probabilistic doo-dads to date there is no logical (or even overwhelming empirical) road starting at linguistic variablilty and ending with statistical grammars.  Addressing the issue requires (at least) carefully distinguishing degree of grammaticality from degree of acceptability (Norvig, recall runs these two concepts together) and showing that the analysis of the second requires something like the first. To date, I know of nothing that speaks to this convincingly, though should it prove correct (as many including Chomsky have speculated over the years) that would be fine with me.[3]

The third use of stats is pernicious.  It expresses a belief, one that I believe Norvig shares, that statistical analysis of lots of data can substitute for  of theory construction.  This represents a return of a certain unfortunate empiricism, both in its scientific methodology and its susceptibility to associationist conceptions of mind.  The program of statistically massaging enormous amounts of data (SMEAD) stands in place of looking for underlying causal powers of surface phenomena. As noted in an earlier post, this view has a home in classical Empiricism, which takes scientific inquiry as the hunt for regularities that describe behavior rather than the powers/natures/capacities whose complex interaction generate it.  Empiricist methodology swerves into associationist psychology when what is sampled and statistically measured is unanalyzed experiential data (rather than experimentally created phenomena, see below), the idea being that fast machines with vast memory resources can crunch lots of very simple surfacy data and in so doing uncover underlying causal mechanisms.  This is classical empiricist dogma and in the (what I hope will soon become) the immortal words of Sydney Brenner it’s “a form of insanity,” a species of “low input, high throughput, no output science.” 

There are at least two things wrong with SMEAD. First, it very much misconstrues the relation between data and theory.  And it does so in two ways.

First, scientific observation takes a lot of very careful stage setting. As Cartwright notes:

Outside the supervision of a laboratory or the closed casement of a factory-made module, what happens in one instance is rarely a guide to what will happen in others.  Situations that lend themselves to generalizations are special…(86).

In other words, most interesting data is manufactured, not natural.  Phenomena are factitious and are barely visible in the interactive swamp that constitutes regular experience. The scientific attitude rests not on observing the world but on putting it in a very tight straightjacket and then prodding and poking (if not worse; how would you like to sent flying at velocities near the speed of light to smash into a target possibly moving at YOU at a similar speed?)) until it reveals some small piece of useful data.  Nancy Cartwright sums this up well:

For anyone who believes that induction provides the primary building tool for empirical knowledge, the methods of modern experimental physics must seem unfathomable.  Usually the inductive base for the principles under test is slim indeed,…and in the best experimental designs…one single instance can be enough…Clearly in these physics experiments we are prepared to assume that the situation before us is of a very special kind: it is a situation in which the behavior that occurs is repeatable.  Whatever happens in this situation can be generalized (85).

This is no less true in linguistics than it is in any other domain of inquiry with theoretical aspirations.  Linguists have been luckier than most in that acceptability judgments (a very crude form of experiment) have been reliable probes to the structure of UG.[4]  Linguistic judgments made in “reflective equilibrium” seem capable of abstracting away from many interfering factors and allow a stable basis for theory construction.  In contrast to simply looking at linguistic behavior, judgment queries can be targeted to structures of interest (you don’t have to wait for someone to say what you are interested in), can alleviate attention and memory pressures (leading to slips of the tongue and “brainos”), and, most importantly, can query the status of unacceptable structures.  Within linguistics it is more often than not the dog that doesn’t bark that is the key to figuring out the fine structure of UG. 

Second, ‘data’ is a misleading term for what scientists use to advance insight for it suggests small isolated points, like data on a graph ready for statistical smoothing.  What drives understanding in the sciences are ‘effects’ and anomalies: e.g. the ultraviolet catastrophe, the perihelion of mercury, the Doppler effect, the retrograde motion of Mars, the Compton effect, etc. Linguistics is blessed with a dozen or so of these effects; e.g. intervention effects, island effects, fixed subject effects, strong and weak crossover effects, filled gap effects etc. These effects are only visible with the greatest contrivance and detectable when mere “observation” has been left far behind.[5]

Last of all, without theory of some sort even stats can’t get off the ground.  Stats are fancy methods of counting. Theory tells you what to count.  Even simple data does not organize itself. Things are counted with respect to some properties. SMEAD appears to believe that large data sets can organize themselves, pull themselves up by their large data base bootstraps. This is false (recall: it is impossible to pull oneself up by the bootstraps). Categories are always necessary and in the absence of some prior theory, surface distinctions become the default categories and these have a tendency of breeding associationist sympathies.  Indeed, if your aim is to collect a lot of data quickly and crunch lots of it fast then visible surface distinctions are what will entice you.  Regularities/associations among the observables is just a step away.  If we have learned anything in the last 50 years, it’s that this method of research will teach us nothing of interest as regards any non-trivial cognitive capacity.

There is a myth abroad in the land that generative grammar has a principled antipathy to stats. Wrong. We have a principled antipathy to useless and misleading and pernicious stats and the technical virtuosos who think that counting anything and everything can replace hard thinking.  All the rest, like all other tools, must prove its worth empirically, which in this case means, can be used to shed light on the structure of UG.


[1] Sprouse and colleagues have argued that these techniques can move beyond quality control and provide a novel kind of data for the investigation of UG and its interactions. This, of course, is a welcome development.
[2] Compare Quine’s empiricist description: language is nothing but “a fabric of sentences variously associated to one another and to nonverbal stimuli by the mechanism of conditioned response.”  Substitute ‘frequencies’ (what Norvig’s probabilities track) for ‘conditioned response’ and the two seem animated by the same unfortunate ethos.
[3] Within syntax degree of grammaticality has been modeled in terms of simple counting: violating two conditions is worse than one, violating some constraints is more serious than others.  In systems this simple, stats are not required.
[4] I suspect that the utility of such a crude procedure speaks volumes about the robustness of the language instinct.
[5] Those interested in these issues might like to look at this in addition to Cartwright’s above mentioned work.