Comments

Showing posts with label acceptability and gramamticality. Show all posts
Showing posts with label acceptability and gramamticality. Show all posts

Tuesday, January 15, 2019

Movement, islands and the ECP

Some papers reset the research agenda. This one by Lu and Yoshida (L&Y), I believe, is one of those (here is a slide conveying the basic point. The paper is under submission at LI and I assume it will be accepted and rapidly published (if not this is will tell us more about LI than it will about the quality of this paper)). The topic is the island status of Wh-in-situ (WIS) constructions in Chinese. The finding is that using judgment studies of the Sprouse experimental syntax (ES) variety provides evidence for two stunning conclusions: (i) that WISs respect islands and (ii) that there is no evidence for an argument/adjunct distinction wrt WISs. Both data points are theoretically pregnant and this post will largely concentrate on drawing out some of the implications. Many of these are mentioned in the paper (yup, I have a draft), so are not original with me. Let’s start.

L&Y is motivated by the premise that ES provides a useful tool for the refining linguistic judgments. The idea, as Sprouse has convincingly argued, is that grammatical complexity should induce a super-additivity effect in well-constructed judgment experiments (see, e.g. here and here for discussion and here for a nice review of the methodology). Importantly, super-additivity profiles arise in cases where less involved rating studies find nothing indicating un-grammaticality.

Before pressing on, let’s make an important and obvious point: all GGers distinguish (or should distinguish) acceptability from grammaticality. Acceptability is a probe for grammaticality. Acceptability is an observable property of utterances. Grammaticality is an abstract property of I-linguistic mental representations. Grammaticality is inferred from acceptability under the right conditions (given the right controls as realized by the appropriate minimal pairs). All of this is old hat, but a still very stylish and durable hat. 

Happily, for most of what we have done in GG, acceptability closely tracks grammaticality, but we also know that the two notions can and do diverge (see here for some discussion). ES is particularly useful for cases where this happens and the simple judgment elicitation procedure (e.g. ask a native speaker) indicates all is well. Diogo Almeida has dubbed cases of the latter “subliminal.” One of ES’s important contributions to syntax has been the discovery of such subliminal effects (SE), SEs being cases where ES procedures reveal super-additivity effects while more regular elicitation suggests grammaticality. So, for example, we now have many examples where standard elicitation has indicated that a certain dependency in a certain language shows no island sensitivity (i.e. the sentences are judged (highly) acceptable) while ES techniques indicate sensitivity to these same island effects (i.e. the relevant data display super-additivity effects).

We also find the converse: standard techniques indicating a profound difference in acceptability, while ES techniques showing nothing at all.[1]All in all then, ES has provided a useful additional kind of data, one that is often more sensitive to G structure than the quick and dirty (and largely accurate and hence very useful) standard judgment techniques, which sometimes fail to track these. 

So, back to the main point: L&Y is an ES study of WISs in Chinese and it has two important findings: that allWISs in Chinese exhibit relative clause island effects (henceforth RCI) (i.e. they alldisplay the super-additivity profile) and that there is no ES evidence that long “why” movement from an RCI is appreciably worse than long “why” movement absent an RCI (i.e. these cases when contrasted do not show a super-additvity profile). The first result argues that WISs are island sensitive and the second argues that there is no additional ECP effect distinguishing WISs like who/what from WISs like why. If correct, this is very big news, and, IMO, very welcome news. Let me say why.

First, as L&Y emphasizes, this result rules out most of the standard approaches to WIS constructions. In particular the result rules out two kinds of theories: (i) accounts that distinguish between overt movement vs covert movement (e.g. Huang’s) and treat island effects as effectively reflexes of overt movement (say, via a chain condition at SS) and (ii) theories that postulate two different kinds of operations (Movement vs Binding) to license WISs with movement subject to islands and binding exempt from them (as in, say, a Rizzi-Cinque approach to ECP effects). Both such kinds of theories will have problems with the apparent fact that WISs induce super-additivity effects.

It is worth noting, furthermore, that the sensitivity of WISs to islands is not the only example of apparent non-movement generated structures being island compliant. The same holds wrt resumptive pronoun (RP) constructions. These also appear to respect islands despite the absence of the main hallmark of movement (i.e. a gap in the “movement” site).[2]Both this RP data and now the WIS data point to the same conclusion: that island effects are notPF effects.[3]From my reading of the literature, this is the most popular current approach to islands and it has some terrifically interesting evidence in support (in particular the fact that some ellipsis (i.e. sluicing) obviates island violations). However, if L&Y are right, then we may have to rethink this assumption (see note 3 however).

Indeed, I would go further (and here it is NH speaking rather than L&Y). There have long been two general approaches to islands. 

First, we have Chomsky’s view of subjacency elaborated in ‘On wh movment’ that treats islands as reflecting bounds on the computational procedure. Island effects reflect the subjacency condition (aka PIC), which bounds the domain of computation (an idea motivated by the reasonable assumption that bounding a domain of computation makes doing computations more tractable).[4]

The second approach can be traced back to Ross’s thesis (islands restrict chopping rules) but has been developed as part of the linearization industry spurred by Kayne’s seminal work and mooted most explicitly by Uriagereka.[5]

The L&Y results argue pretty strongly, IMO, for Chomsky’s original conception precisely because they appear to hold whether or not the construction involves an obvious phonetic gap (gaps being problematic as they undo linearizations). If this is so, then it argues against linearization based approaches to the problem (leaving, of course a very big question: what to do about sluicing).[6]

We can go further still. The L&Y results also argue for a Merge only syntax. Here is what I mean. IMO, the central empirical thesis of the Minimalist Program (MP) is the Merge Hypothesis (MH). MH is the claim that the only specifically linguistic operation of FL is Merge. This entails that allG dependencies are merge mediated. The strong version of the thesis excludes operations like long distance Agree, which spans the same domains as I-merge but is a different operation. Note that it is natural to suppose that I-merge is movement and Agree is some kind of binding or feature sharing. At any rate, the classical conceptions gain empirical benefit from the “observation” that WISs do not display island effects. Why? Because, we might say, they are licensed by Agree not by I-merge and only the latter (being the MP analogue of movement) is subject to subjacency (or its current analogue). But as L&Y indicates this is precisely the wrong conclusion. WISs are subject to islands. A merge only syntax insists that all A’-dependencies are formed in the same way, via I-Merge, as this as the only way to establish any non-local grammatical dependency. So if WISs are G licensed, then they must be G licensed via I-merge and so will form a natural class with overt Wh movement. And this is what L&Y find. Both show super-additivity effects across islands. Thus, L&Y’s findings are what we should expect from a merge only syntax and it cautions against larding this best of MP theories with Agree/Probe-Goal titivations.[7]

We can milk a second important conclusion from L&Y. It solves a giant problem for MP. Which problem? The problem of unifying subjacency with the ECP. I have suggested elsewhere (see here) that the argument/adjunct asymmetries at the heart of the ECP are very MP problematic. This is so for a variety of reasons. The three that move me most are the fact that the ECP is a trace licensing condition and MP eschews traces, the huge theoretical redundancy between ECP and subjacency, and the “ugliness” of the basic technical machinery required to allow the ECP to track the argument/adjunct asymmetry. One of the nice implications of L&Y is that we need not worry about the problems that the ECP generates for MP because the theoretical apparatus is based on a mistaken description of the data. If L&Y is right, then there is no argument/adjunct asymmetry. Poof, the MP problem disappears and with it the ad-hoc theoretically unmotivated (within MP) technical apparatus required to track it.  

Of course, this overstates matters. It behooves us to go over the ECP data more carefully and see how to resolve the difference in acceptability that the standard literature identified. Why after all if all Whs are created equal do long distance adjuncts resist movement more fiercely than do long distance arguments?[8]L&Y offer a suggestion (no spoiler from me, read the squib when it comes out). But whatever the right answer is, it does not rest on making an invidiousgrammaticaldistinction between the two kinds of dependencies. And this is just what MP needs in order to start distinguishing ECP effects from the ECP theoretical apparatus in GB.

Let me hit this a bit harder. The ugliest parts of theories of A’-dependency within GB arise in response to argument/adjunct asymmetry effects. The technical machinery in Barriers (built on Lasnik and Saito foundations), though successful empirically (IMO, Lasnik and Saito’s theory was considerably more empirically effective than Barriers) had little of the virtual conceptual necessity MPers pine for. Nor were alternative theories (Generalized Binding, Connectedness) much prettier. Nonetheless, we put up with that stuff and developed it theoretically because it appeared to be empirically called for. The right aesthetic conclusion should have been (and actually was) that it was too contrived to be correct. L&Y provides courage for our aesthetic convictions. We should have judged these theories as suspect because ugly, though we would have been empirically premature in drawing that conclusion. Given L&Y, the facts are not what we took them to be despite reflecting very different acceptability profiles. 

There is a moral here, and you can all guess what it is but I cannot resist making it explicit anyhow. L&Y provide evidence for a methodological precept that we fail to respect enough: facts can change, no less than theory can. Or to put this another way: just as we can make theoretical wrong turns that we come to revise, we can adopt empirical generalizations that turn out to be misleading. The standard view is that data is hard and theory is fluffy and when the two clash it is best to revise the theory than rethink the data. L&Y provides a case where this is reversed. And I say hooray!

Let me make one more point and I will end this overly long post (long, and yet, filled with endlessly many loose ends). L&Y exemplifies something that I think is important. It is an empirical paper whose purpose is to directlytest a core theoretical assumption. This is not something that we generally see within syntax. Most papers are not out to test theoretical assumptions. Most papers use theory to explore more data. Theory might be tested but it is generally a by-product of better descriptive coverage. L&Y works differently. It starts from the theory and constructs an empirical intervention to probe it. Moreover, the question is quite precise and the assumptions required to answer it are clear within the confines of the project. This has all the look and smell of an honest to god experiment, a process whereby we query the theory using a relatively well-understood probe. Both empirical methods of exploration are worthwhile, but they are different and it is only relatively recently, I think, that we are seeing examples of the second experimental kind gaining traction. 

Curiously (perhaps), a feature of this second kind of paper is that the paper is short. L&Y is a squib. Empirical explorations in linguistics often read like novellas. L&Y is very definitely a very very short story. I would like to suggest that experimental papers like L&Y reflect the scientific health of linguistics. It is now possible to ask a sharp question, and give a sharp answer. We need more of these kinds of short pointed experimental forays into the data starting from well-formulated theoretical starting points.

That’s it. I have gone on far too long. L&Y is terrific. If correct, it is very important. I personally hope the results stand up. It would go a long way to cleaning up a particularly untidy part of syntactic theory and thereby further vindicating the promise of MP, indeed a particularly strong version of MP, one that endorses a Merge only conception of grammar.


[1]ES techniques have, for example, suggested that adjunct island effects might not be of a piece with other islands as they (often) fail to display super-additivity effects.
[2]To be slightly more careful and squinting at the ES results wrt resumptives the following is more accurate: resumptives uniformly ameliorate fixed subject constraint violations but note “mere” subjacency violations. Thus, resumptives inside islands seem to show the same super-additvity profiles as their moved counterparts.
[3]Perhaps a more felicitous way of putting matters is that RPs and WISs are also products of I-merge and so expected to be subject to islands. One reason for treating islands as PF effects is to capture the distinctionbetween these cases and more conventional examples of overt movement. If, however, they pattern the same then the motivation for treating islands as PF effects weakens. This said, I am pretty sure it is possible to model these cases as formed via movement/I-merge by, for example, treating them as cases of remnant movement, the moved Wh or Q morpheme starts as part of a doubled structure including the RP or WIS. 
[4]See Chomsky’s ‘On wh movement’ for discussion. See herefor more prose.
[5]Versions of this original idea were developed by many people including Hornstein, Lasnik and Uriagereka, and Fox and Pesetsky. The idea centers on the idea that the problem with movement is that it reorders elements and so can come into conflict with the ordering algorithm. In this sense, gaps are a big deal and what distinguish movement from other kinds of long distance dependencies like binding.
[6]Or again, it argues for treating the operation (e.g. movement) as in need of constraint rather than the output of the operation (a gap, or new linear order). What plausibly unifies cases of “overt” WH (as in English), “covert” WH (as in Chinese) and resumptive WH (as in Hebrew) is that they all involve relating an A’ element to a non-local syntactic position that can be arbitrarily far away. IT’s the span that seems to matter, not what sits at the tail of the chain.
[7]RP constructions must also be formed via I-merge and so too all forms of binding. This requirement fits well with the observation that RPs obey islands. Binding, especially pronominal binding, is likely to be more problematic. As they say at this point in a journal paper in reply to referee 2; these are topics for further research.
[8]As I have noted in other places, this description of the ECP contrast is not quite correct, as we have known since Rizzi’s work on minimality. The distinction seems less a matter of argument vs adjunct than object centered vs non object centered quantification. But this is a topic for another time.

Monday, March 14, 2016

The deep difference between acceptable and grammatical

Up comes a linguist in the street interviewer and asks: “So NH, what would grad students at UMD find to be one of your more annoying habits?” I would answer: my unrelenting obsession with forever banishing from the linguistics lexicon the phrase “grammaticality judgment.” As I never tire of making clear, usually in a flurry of red ball-point scribbles and exclamation marks, the correct term is “acceptability judgment,” at least when used, as it almost invariably is, to describe how speakers rate some bit of data. “Acceptability” is the name of the scale along which such speaker judgments array. “Grammaticality” is how linguists explain (or partly explain) these acceptability judgments. Linguists make grammaticality judgments when advancing one or another analysis of some bit of acceptability data. I doubt that there is an interesting scale for such theoretical assessments.

Why the disregard for this crucial difference among practicing linguists? Here’s a benign proposal. A sentence’s acceptability is prima facie evidence that it is grammatical and that a descriptively adequate G should generate it. A sentence’s unacceptability is prima facie evidence that a descriptively adequate G should not generate it. Given this, using the terms interchangeably is no big deal. Of course, not all facies are prima and we recognize that there are unacceptable sentences that an adequate G should generate and that some things that are judged acceptable nonetheless should not be generated. We thus both recognize the difference between the two notions, despite their intimate intercourse, and interchange them guilt free.

On this benign view, my OCD behavior is simple pedantry, a sign of my inexorable aging and decline. However, I recently read a paper by Katz and Bever (K&B) that vindicates my sensitivities (see here), which, of course, I like very much and am writing to recommend to you (I would nominate it for classic status).[1] Of relevance here, K&B argues that the distinction between grammaticality and acceptability is an important one and that blurring it often reflects the baleful influence of that most pernicious intellectual habit of mind, EMPIRICISM! I have come to believe that K&B is right about this (as well, I should add, about many other things, though not all). So before getting into the argument regarding acceptability and Empiricism, let me recommend it to you again. Like an earlier paper by Bever that I posted about recently (here), this is a Whig History of an interesting period of GG research. There is a sustained critical discussion of early Generative Semantics that is worth looking at, especially given the recent rise of interest in these kinds of ideas. But this is not what I want to discuss here. For the remainder, let me zero in on one or two particular points in K&B that got me thinking.

Let’s start with the acceptability vs grammaticality distinction. K&B spend a lot of time contrasting Chomsky’s understanding of Gs and Transformations with Zellig Harris’s. For Harris, Gs were seen as compact ways to cataloguing linguistic corpora. Here is K&B (15):

 …grammars came to be viewed as efficient data catalogues of linguistic corpora, and linguistic theory took the form of a mechanical discovery procedure for cataloguing linguistic data.

Bloomfieldian structuralism concentrated on analyzing phonology and morphology in these terms. Harris’s contribution was to propose a way of extending these methods to syntax (16):

Harris’s particular achievement was to find a way of setting up substitution frames for sentences so that sentences could be grouped according to the environments they share, similar to the way that phonemes or morphemes were grouped by shared environments…Discourse analysis was…the product of this attempt to extend the range of taxanomic analysis beyond the level of immediate constituents.

Harris proposed two important conceptual innovations to extend Structuralist taxonomic techniques to sentences; kernel sentences and transformations. Kernels are a small “well-defined set of forms” and transformations, when applied to kernels, “yields all the sentence constructions of the language” (17). The coocurrence restrictions that lie at the heart of the taxanomy are stated at the level of kernel sentences. Transformations of a given kernel define an equivalence class of sentences that share the same discourse “constituency.” K&B put this nicely, albeit in a footnote (16:#3):

In discourse analysis, transformations serve as the means of normalizing texts, that is of converting the sentences of the text into a standard form so that they can be compared and intersentence properties [viz. their coocuurences, NH] discovered.

So, for Harris, kernel sentences and transformations are ways of compressing a text’s distributional regularities (i.e. “cataloguing the data of a corpus” (12)).[2]

This is entirely unlike the modern GG conception due to Chomsky, as you all know. But in case you need a refresher, for modern GG, Gs are mental objects internalized in brains of native speakers and which underlie their ability to produce and understand an effectively unbounded number of sentences, most of which have never before been encountered (aka; linguistic creativity). Transformations are a species of rule these mental Gs contain that map meaning relevant levels of G information to sound relevant (or articulator relevant) levels of G information. Importantly, on this view, Gs are not ways of characterizing the distributional properties of texts or speech. They are (intended) descriptions of mental structures.

As K&B note, so understood, much of the structure of Gs is not surface visible. The consequence?

The input to the language acquisition process no longer seems rich enough and the output no longer simple enough for the child to obtain its knowledge of the latter by inductive inferences that generalize the distributional regularities found in speech. For now the important properties of the language lie hidden beneath the surface form of sentences and the grammatical structure to be acquired is seen as an extremely complex system of highly intricate rules relating the underlying levels of sentences to their surface phonetic form. (12)

In other words, once one treats Gs as mental constructs the possibility of an Empiricist understanding of what lies behind human linguistic facility disappears as a reasonable prospect and is replaced by a Rationalist conception of mind. This is what made Chomsky’s early writings on language so important. They served to discredit empiricism in the behavioral sciences (though ‘discredit’ is too weak a word for what happened). Or as K&B nicely summarize matters (12):

From the general intellectual viewpoint, the most significant aspect of the transformationalist revolution is that it is a decisive defeat of empiricism in an influential social science.  The natural position for an empiricist to adopt on the question of the nature of grammars is the structuralist theory of taxanomic grammar, since on this theory every property essential to a language is characterizable on the basis of observable features of the surface form of its sentences. Hence, everything that must be acquired in gaining mastery of a language is “out in the open”; moreover, it can be learned on the basis of procedures for segmenting and classifying speech that presupposes only inductive generalizations from observable distributional regularities. On the structuralist theory of taxanomic grammar, the environmental input to language acquisition is rich enough, relative to the presumed richness of the grammatical structure of the language, for this acquisition process to take place without the help of innate principles about the universal structure of language…

Give up the idea that Gs are just generalizations of the surface properties of speech, and the plausibility of Empiricism rapidly fades. Thus, enter Chomsky and Rationalism, exit taxonomy and Empiricism.

The shift from the Harris Structuralist, to the Chomsky mentalist, conception of Gs naturally shifts interest to the kinds of rules that Gs contain and to the generative properties of these rules. And importantly, from a rule-based perspective it is possible to define a notion of ‘grammaticality’ that is purely formal: a sentence is grammatical iff it is generated by the grammar. This, K&B note is not dependent on the distribution of forms in a corpus. It is a purely formal notion, which, given the Chomsky understanding of Gs, is central to understanding human linguistic facility. Moreover, it allows for several conceptions of well-formedness (phonological, syntactic, semantic etc.) that together contribute along with other factors to a notion of acceptability, but are not reducible to it. So, given the rationalist conception, it is easy and natural to distinguish various ingredients of acceptability.

A view that takes grammaticality to just be a representation of acceptability, the Harris view, finds this to be artificial at best and ill-founded at worst (see K&B quote of Harris p. 20). On a corpus-based view of Gs, sentences are expected to vary in acceptability along a cline (reflecting, for example, how likely they are to be found in a certain text environment). After all, Gs are just compact representations of precisely such facts. And this runs together all sorts factors that appear diverse from the standard GG perspective. As K&B put it (21):

Statements of the likelihood of new forms occurring under certain conditions must express every feature of the situation that exerts an influence on likelihood of occurrence. This means that all sorts of grammatically extraneous features are reflected on a par with genuine grammatical constraints. For example, complexity of constituent structure, length of sentences, social mores, and so on often exerts a real influence on the probability that a certain n-tuple of morphemes will occur in the corpus.

Or to put this another way: a Rationalist conception of G allows for a “sharp and absolute distinction between the grammatical and the ungrammatical, and between the competence principles that determine the grammatical and anything else that combines with them to produce performance” (29). So, a Rationalist conception understands linguistic performance to be a complex interaction effect of discrete interacting systems. Grammaticality does not track the linguistic environment. Linguistic experience is gradient. It does not reflect the algebraic nature of the underlying sub-systems. ‘Acceptability’ tracks the gradiance, ‘grammaticality’ the discrete algebra. Confusing the two threatens a return to structuralism and its attendant Empiricism.

Let me mention one other point that K&B makes that I found very helpful. They outline what a Structuralist discovery procedure (DP) is (15). It is “explicit procedures for segmenting and classifying utterances that would automatically apply to a corpus to organize it in a form that meets” four conditions:

1.     The G is a hierarchy of classes; lower units being temporal segments of speech event, the higher are classes or sequences of classes.
2.     The elements of each level are determined by their distributional features together with their representations at the immediately lower level.
3.     Information in the construction of a G flows “upward” from level to level, i.e. no information at a higher level can be used to determine an analysis at a lower level.
4.     The main distributional principles for determining class memberships at level Li are complementary distribution and free variation at level Li-1.

Noting the structure of a discovery procedure (DP) in (1-4) allows us to appreciate why Chomsky stressed the autonomy of levels in his early work. If, for example, the syntactic level is autonomous (i.e. not inferable from the distributional properties of other levels) then the idea that DPs could be adequate accounts of language learning evaporates.[3] And once one focuses on the rules relating articulation and interpretation the plausibility of a DP for language with the properties in (1-4) becomes very implausible, or, as K&B nicely put it (33):

Given that actual speech is so messy, heterogeneous, fuzzy and filled with one or another performance error, the empiricist’s explanation of Chomskyan rules, as having been learned as a purely inductive generalization of a sample of actual speech is hard to take seriously to say the very least.[4]

So, if one is an Empiricist, then one will have to deny the idea of a G as a rule based system of the GG variety. Thus, it is no surprise that Empiricists discussing language like to emphasize the acceptability gradients characteristic of actual speech. Or, to put this in terms relevant to the discussion above, why Empiricists will understand ‘grammaticality’ as the limiting case of ‘acceptability.’

Ok, this post, once again, is far too long. Look at the paper. It’s really good and useful. It also is a useful prophylactic against recurring Empiricism and, unfortunately, we cannot have too much of that.



[1] Sadly, pages 18-19 are missing from the online version. It would be nice to repair this sometime in the future. If there is a student of Tom’s at U of Arizona reading this, maybe you can fix it.
[2] I would note that in this context corpus linguistics makes sense as an enterprise. It is entirely unclear whether it makes any sense once one gives up this structuralist perspective and adopts a Chomsky view of Gs and transformations. Furthermore, I am very skeptical that there exist Harris-like regularities over texts, even if normalized to kernel sentences. Chomsky’s observation that sentences are not “stimulus bound” if accurate (and IMO they are) undermine the view that we can say anything at all about the distributions of sentences in texts. We cannot predict with any reliability what someone will say next (unless, of course, it is your mother), and even if we could in some stylized texts, it would tell us nothing about how the sentence could be felicitously used. In other words, there would be precious little generalizations across texts. Thus, I doubt that there is any interesting statistical regularities regarding the distribution of sentences in texts (at least understood as stretches of discourse).
            Btw, there is some evidence for this. It is well known that machines trained on one kind of corpus do a piss poor job of generalizing to a different kind of corpus. This is quite unexpected if figuring out the distribution of sentences in one text gave you a good idea of what would take place in others. Understanding how to order in a restaurant or make airline reservations does not carry over well to a discussion of Trump’s (and the rest of the GOP’s) execrable politics, for example. A long time ago, in a galaxy far far away I very polemically discussed these issues in a paper with Elan Dresher that still gives me chuckles when I read it. See Cognition 1976, 4, pp.32l‑398.
[3] That this schema looks so much like those characteristic of Deep Learning suggests that it cannot be a correct general theory of language acquisition. It just won’t work, and we know this because it was tried before.
[4] That speech is messy is a problem, but not the only problem. The bigger one is that there is virtually no evidence in the PLD for may of G properties (e.g. ECP effects, island effects, binding effects etc.). Thus the data is both degenerate and deficient.

Tuesday, September 22, 2015

Degrees of grammaticality?

In a recent post (here), I talked about the relation between two related concepts, acceptability and grammaticality and noted that the first is the empirical probe that GGers have standardly used to study the second more important concept. The main point was that there is no reason to think that the two concepts should coincide extensionally (i.e. that some linguistic object (LO) is (un)grammatical iff it is (un)acceptable. We should expect considerable slack between the two, as Chomsky noted long ago in Aspects (and Current Issues). Two things follow from this, one surprising the other not so much. The expected fact is that there are cases where the two concepts diverge. The more surprising one is that this happens surprisingly rarely (though this is more an impression than a quantifiable (at least by me, claim)) and that when it does happen it is interesting to try and figure out why. 

In this post, I’d like to consider another question; how important are (or, should be) notions like degree of acceptability and grammaticality. It is often taken for granted that acceptability judgments are gradient and it is often assumed that this is an important fact about human linguistic competence and, hence, must be grammatically addressed. In other words we sometimes find the following inchoate argument: acceptability judgments are gradient therefore grammatical competence must incorporate a gradient grammatical component (e.g. probabilistic grammars or notions like degree of grammaticality). In what follows I would like to address this ‘therefore.’  I have no problem thinking that sentences may be grammatical to some degree or other rather than simply +/- grammatical. What I am far less sure of is whether this possible notion has been even weakly justified.  Let’s start with gradient acceptability.

It is not infrequently observed that acceptability judgments (of utterances) are gradient while most theories of G provide categorical classifications of Los (e.g. sentences). Thus, though GGers (e.g. Chomsky) have allowed that grammaticality might be a graded notion (i.e. the relevant notion being multi-valued not binary) in practice GG accounts have been based on a categorical understanding of the notion of well-formedness. It is often further mooted that we need (or methodologically desire) a tighter fit between the two notions and because acceptability data is gradient therefore we need a gradient notion of grammaticality. How convincing is this argument?

Let me say a couple of things. First, it is not clear, at least to me, that all or even most acceptability judgment data is gradient.  Some facts are clearly pretty black or white with nary a shade of gray. Here’s an example of one such.

Take the sentences (1)-(3):

(1)  Mary hugged John
(2)  John hugged Mary
(3)  John was hugged by Mary

It is a fact universally acknowledged that English speakers in search of meanings will treat (1) and (2) quite differently. More specifically, all speakers know that (1) and (2) don’t mean the same thing, that in (1) Mary is the hugger and John the thing hugged while the reverse is true in (2), that (3) is a paraphrase of (1) (i.e. that the same hugger hugee relations hold in (3) as in (1)) and that (3) is not a paraphrase of (2). So far as I can tell, these judgments are not in the least variable and they are entirely consistent across all native speakers of English, always. Moreover, these kinds of very categorical data are a dime a dozen. 

Indeed, I would go further: if asked to ordinally rank pairs of sentences, over a very wide range, speaker judgments will be very consistent and categorically so. Thus, everyone will judge the conceptually opaque colorless green ideas sleep furiously superior to the conceptually opaque ideas furiously sleep green colorless, and everyone will judge who is it that John persuaded someone who met to hug him is less acceptable than who is it that persuaded someone who met John to hug him. As regards pair-wise relative acceptability, the data are very clear in these cases as well.[1]

Where things become murkier is when we ask people to assign degrees of acceptability to individual sentences or pairs thereof. Here we ask not if some sentence is better or worse than another, but ask people to provide graded judgments to stimuli (to say how much worse); rate this sentence’s acceptability between 1-7 (best to worst) on a scale.  This question elicits graded judgments. Not for all cases, as I doubt that any speaker would be anything but completely sure that (1) and (2) are not paraphrases and that (1) and (3) depict the same events in contras to (2). But, it might be that whereas some assign a 6 to the first wh question above and a 2-3 for the second, come might give different rankings. I am even ready to believe that the same person might give different rankings on different occasions. Say that this is true? What if anything should we conclude about the Gs pertinent to these judgments?

One conclusion is that the Gs must be graded because (some of) our acceptability judgments are. But why believe this? Why believe that grammaticality is graded simply because how we measure it using one particular probe is graded (conceding for the moment that acceptability judgments are invariably gradient). Surely, we don’t conclude that the melting point of lead (viz. 327.5 Celsius) is gradient just because every time we measure it we get a slightly different value. In fact, anytime we measure anything we get different values. So what? Thus, the mere fact that one acceptability measure yields gradient values implies very little about whether grammaticality (i.e. the thing measured by the acceptability judgment) is also gradient. We wouldn’t conclude this about the melting point of lead, so why conclude this about the grammaticality of John saw Mary or the ungrammaticality of who did John see Mary?

Moreover, we know a thing or two about how gradient values can arise from the interaction of non-gradient factors. Indeed, phenomena that are the product of many interacting factors can lead to gradient outputs precisely because they combine many disparate factors (think of continuous height which is the confluence of many interacting integral genetic factors). Doesn’t this suffice to accommodate graded judgments of acceptability even if grammaticality is quite categorical? We have known forever that grammaticality is but one factor in acceptability so we should not be surprised that the many many interacting factors that underlie any give acceptability judgment lead to a gradient response even if every one of the contributing factors is NOT gradient at all.

I would go further: what is surprising is not that we sometimes get gradient responses when we ask for them (and that is what asking people to rate stimuli on a 7 point scale is doing) but that we find it easy to consistently bin a large number of similar sentences into +/- piles. That is surprising. Why? Because it suggests that grammaticality must be a robust factor in acceptability for it outshines all the other factors in many cases even when collateral influences are only weakly controlled. That is surprising (at least to me): why can we do this so reliably over a pretty big domain? Of course, the fact that we can is what makes acceptability judgments decent probes into grammatical structure, with all of the usual caveats (i.e. usual given the rest of the sciences) about how measuring might be messy.

Note, I personally have nothing against the conclusion that ‘grammatical’ is not a binary notion but multi-valued. Maybe it is (Chomsky has constantly mentioned this possibility over the years, especially in his earliest writing (see quote at end of post)). So, I am not questioning the possibility. What I want to know is why we should currently assume this? What is the data (or theoretical gain) that warrants this conclusion? It can’t just be variable acceptability judgments for these can arise even if grammaticality is a simple binary factor. In other words noting that we can get speakers to give gradient judgments is not evidence that grammars are gradient.

A second question: how would it change things (theory or practice) were it so. In other words, how do we conceptually of theoretically benefit by assuming that Gs are gradient? Here’s what I mean. I take it as a virtual given that linguistic competence rests (in part) on having a G. Thus it seems a fair question as to what kinds of Gs humans have: what kinds of structures and dependencies are characteristic of human Gs. Once we know this, we can add that the Gs humans use to produce and understand the linguistic world around them code not only the kinds of dependencies that are possible, but the probabilities of usage concerning this or that structure or dependency. In other words, I have no problem probabilizing a G. What I don’t see is how this second process of adding numbers between 0 and 1 to Gs will eliminate the need to specify the class of Gs without their probabilistic clothing. If it won’t, then whether or not Gs carry probabilities will not much affect the question of what the class of possible Gs is.

Let me put this another way: probabilities requires a set of options over which these probabilities (the probability “mass” (I love that term)) are distributed. This means that we need some way of specifying the options. This is something that Gs do well: they specify the range of options to which probabilities are then added (as people like John Hale and his friends) do to get probabilistic Gs, useful for the investigation of various kinds of performance facts (e.g. how hard something is to parse). Doing this might even tell us something about which Gs are right (as Time Hunter has recently been arguing (see here)). This all seems perfectly fine to me. But, but, but, …this all presupposes that we can usefully divide the question into two parts: what is the grammar and how are probabilities computed over these structures. And if this is so, then even if there is a sense in which Gs are probabilistic, it does not in any way suggest that the search for the right Gs is any way off the mark. In fact, the probabilistic stuff presupposes the grammar stuff and the latter is very algebraic. Why? Because probabilities presuppose a possibility space defined by an algebra over which probabilities are then added. So the question: why should I assume that grammaticality is a gradient notion even if the data that I use to probe it is gradient (if in fact it is)? I have no idea.

Let me say this another way. One aim of theory is to decompose a phenomenon to reveal the interacting sub-parts. It would be nice if we could regularly then add these subparts up together to “derive” the observable effect. However, this is only possible in a small number of cases (e.g. physics tells us how to combine forces to get a resultant force, but these laws of combination are surprisingly rare and is not possible even in large areas of physics). How then do we identify the interacting factors? By holding the other non-interacting factors constant (i.e. by controlling the “noise”). When successful this allows us to identify the components of interest and even investigate their properties even though we cannot elaborate in any general way how the various factors combine or, even what all of them might be. Thus, being able to control the relevant interfering factors locally (i.e within a given “experiment”) does not imply that we can globally identify all that is relevant (i.e. identify all potentially relevant factors ahead of time). The demand that Gs be gradient sounds like the demand (i) that we eschew decomposing complex phenomena into their subparts or (ii) that we cannot study parts of a complex phenomenon unless we can explain how the parts work together to produce the whole.  The first demand is silly. The second is too demanding. Of course, everyone would love to know how to combine various “forces” to yield a resultant one. But the inability to do this does not mean that we have failed to understand anything about the interacting components. Rather it only implies what is obvious; that we still don’t understand exactly how they interact.[2] This is standard practice in the real sciences and demanding more from linguists is just another instance of methodological dualism.

Let me end by noting that Chomsky in his early work was happy to think that grammaticality was also a gradient notion. In particular, in Current Issues (9) he writes:

This [human linguistic NH] competence can be represented…as a system of rules that we can call a grammar of his language. To each phonetically possible utterance…the grammar assigns a certain structural description that specifies the linguistic elements of which it is constituted and their structural relations…For some utterances, the structural description will indicate…that they are perfectly well-formed sentences. To others, the grammar will assign structural descriptions that indicate the manner of their deviation from perfect well-formedness. Where the deviation is sufficiently limited, an interpretation can often be imposed by virtue of formal relations to sentences of the generated language.

So, there is nothing that GGers have against notions of grammatical gradience, though, as Chomsky notes, it will likely be parasitic on some notion of perfect well-formedness. However, so far as I can tell, this more elaborate notion (though mooted) has not played a particularly important role in GG. We have had occasional observations that some kinds of sentences are more unacceptable than others and that this might perhaps be related to their violating more grammatical conditions. But this has played a pretty minor role in the development of theories of grammar and, so far as I can determine, have not displaced the idea that we need some notion of absolute well-formedness to support this kind of gradient notion. So, as a practical matter, the gradient notion of grammaticality has been of minor interest (if any).

So do we need a notion like degree of grammaticality? I don’t see it. And I really do not see why we should conclude from the (putative, but quite unclear) fact that utterances are (all?) gradient in their acceptability that GG needs a gradient conception of grammar. Maybe it does, but this is a lousy argument. Moreover, if we do need this notion, it seems at this point like it will be a minor revision to what we already think. In other words, if true, it is not clear that it is particularly important. Why then the fuss? Because many confuse the tools they use with the subject matter being investigated. This is a tendency particularly prone to those bent in an Empiricist direction. If Gs are just statistical summaries of the input, then as the input is gradient (or so it is assumed) then the inferred Gs must be. Given this conception, the really bad inference from gradient utterances to gradient sentences makes sense. Moral: this is just another reason to avoid being bent in an Esih direction.


[1] What is far less clear is that we can get a well ordering of these pair-wise judgments, i.e. if A is better than B and B is better than C then A will be better than C. Sometimes this works, but I can imagine that sometimes it does not. If I recall correctly, Jon Sprouse discusses this in his thesis.
[2] Connectionists used to endorse the holistic idea that decomposing complex phenomena into interacting parts falsifies cognition. The whole brain does stuff and trying to figure out what part of a given effect was due to which cognitive powers was taken to be wrong-headed. I have no idea if this is still a popular view, but if it is, that it is not a view endorsed in the “real” sciences. For more discussion of this issue see Geoffrey Joseph’s “The many sciences and the one world,” J of Philosophy December 1980. As he puts it (786):
The scientist does not observe, hypothesize and deduce. He observes, decomposes, hypothesizes, and deduces. Implicit in the methodology of the theorists to whom we owe our knowledge of the laws of the various force fields is the realization that one must formulate theories of the components and leave for the indefinite future the task of unifying the resulting subtheories into a comprehensive theory of the world…A consequence of this feature of his methodology is we are often in the position of having very well-confirmed fundamental theories at hand, but at the same time being unable to formulate complete deductive explanations of natural (complex) phenomena.

If they cannot do this in physics, it would be surprising if we could do this in cognition. Right now aiming to have a comprehensive theory of acceptability strikes me as aiming way too high.  IMO, we will likely never get this, and, more importantly, it is not necessary for trying to figure out the structure of FL/UG/G that we do. Demanding this in linguistics while even physics cannot deliver the goods is a form of methodological sadism, best ignored except in the privacy of your lab meetings between consenting researchers.