Comments

Showing posts with label Hermon and Yanti. Show all posts
Showing posts with label Hermon and Yanti. Show all posts

Monday, October 19, 2015

What's in UG (part 3)

Here is the third and final post on the CHY paper (see here and here).

The second CHY argument goes as follows: (i) the clear categorical complementary distribution of BT-anaphors and pronominals that one finds in languages like English is merely a preference in other languages and (ii) ungrammaticality implies categorical unacceptability. In other words, mere preference (i.e. graded acceptability) is a sure indicator that the acceptability difference cannot reflect G structure.[1] This argument form is one that we’ve encountered before (see here) and it is no more compelling here than it was there, or so I will again argue. Let’s go to the videotape for details.

What is the CHY case of interest. It describes two dialects of Malay. In one the categorical judgments found in English are replicated (call this M-1). In the other, the same kinds of sentences evoke preference judgments rather than categorical judgments (call this M-2). The argument is that because Gs only license categorical judgments, M-2’s preferences cannot be explained Gishly. But as M-1 and M-2 are so similar, then whatever account offered for one must extend to the other. Thus, because the account of M-2 cannot be a Gish one, the account of M-1 can’t be either. That’s the argument. Not very good unless one accepts that categorical (un)acceptability is a necessary property of (un)grammaticality. Reject this and the argument goes nowhere. And reject this we should. Here’s why.

The basic judgment data that linguists use involve relative acceptability (usually under an interpretation). Sometimes, the relevant comparison class is obvious and the data is so clean that we can treat the data as categorical (as I argued here, I think that this is not at all uncommon). However, little goes awry if judgment data is glossed in terms of relative acceptability and virtually all the data can be so construed. Now, in these terms, the perception that some judgments (acceptability under an interpretation in this case) are preferable to others is perfectly serviceable. And it may (need not, but may) reflect underlying Gish properties. It will depend on the case at hand.

I mention this because, as noted, CHY describes M-1 and M-2 as making the same distinctions but with M-1 judgments being categorical and M-2 being preferences. CHY concludes that these should be treated in the same way. I agree. However, I do not see how this implies that the distinction is non-grammatical, unless one assumes that preferences cannot reflect underlying grammatical form.  CHY provides no argument for this. It takes it as obvious. I do not.

Is there anything to recommend CHY’s assumption? There is one line of reasoning that I can think of. How is one to explain gradient acceptability (aka preference) if one takes grammaticality to be categorical?  This is the question extensively discussed here (see the discussion thread in particular) and here. The problem in the domain of island phenomena is that even when we find the diagnostic marks of islandhood (super additivity effects) and we conclude that there is “subliminal” ungrammaticality, we are left asking why for speakers of one G the effects of ungrammaticality manifest themselves in stronger unacceptability judgments than for speakers of another G. In other words, why if island violations are ungrammatical do some find them relatively acceptable? The same question arises in the case that binding data that CHY discusses. And in both instances the question raised is a good one. What kind of answer should/might we expect?

Here’s a proposal: we should expect ungrammaticality ceteris paribus to get reflected in categorical judgments of unacceptability. However, ceteri are seldom paribused.  We know that lots goes into an acceptability judgment and it is hard to keep all things equal.  So for example, it is not inconceivable that sentences with many probable parses are more demanding performance wise than those without. More concretely, imagine a sentence where the visible functional surface vocabulary (FSV) fails to make clear what the underlying structure is. I am assuming, as is standard, that functional vocabulary can be a guide to underlying form (hence the relative “acceptability” of Jabberwocky). Say that in some languages the underlying surface morphology is more closely correlated to the underlying syntactic categories than in others. And say that this creates problems mapping from the utterance to the underlying G form. And say that this manifests itself in muddier (un)acceptability judgments. To say this another way; the less ambiguous the mapping from surface forms to underlying forms the more categorical the judgment will be. If something like this is right, then if we find a language where the morphology does not disambiguate BT-anaphors from exempt anaphors then we might expect acceptability to be less than categorical. Think of it as the acceptability judgment averaging over the two G possibilities (maybe a weighted average). On this scenario, then the absence of a “dedicated reflexive” form (see CHY p. 9) in M-2 will make it harder to apply BT than in a language where there is a dedicated form, as in M-1. Note, this is consistent with the assumption that in both languages the G distinguishes well-formed forms from ill-formed forms. However, it is harder to “see” this in M-2 given the obscurity of the surface FSVs than it is in M-1 where the distinction has been “grammaticalized.”[2]

I mention this option for it is consistent with everything that CHY discusses and, as I hope is evident, it leaves the question of the UG status of BT untouched. In short, in this particular case it is easy enough to cook up an explanation for why binding judgments in M-2 are murkier than those in M-1 without assuming that both reflect the operations of a common UG.[3] Thus, the CHY conclusion is not only based on a debatable premise, but in this particular case there is a pretty obvious way of explaining why the two dialects might provide different acceptability judgments. I should also add that this little story I’ve provided is more than CHY does. Here’s what I mean.

Curiously CHY does not explain how M-1 and M-2 are related except to say that M-1 has grammaticalized a distinction that M-2 has not. Which? M-1 has grammaticalized the notion notion of anaphor. What kind of process is “gramamticalizing”?  CHY does not say. It does not provide an account of what grammaticalization actually is, it only points to some of its effects in Malay and suggests that this goes on in creolization. Nor does CHY explain how undergoing grammataicalization renders preferences in the pre-grammaticalization period categorical in the post grammaticalization period? In CHY “grammaticalization” is Voltarian.[4] Let me offer a proposal of what gramamticalization is (actually this is implicit in CHY’s discussion).

Here’s one proposal: grammaticalization involves sharpening the FSV so that it more directly reflects the underlying G structure. In other words, grammaticalization is a process that aligns surface functional vocabulary with underlying grammatical forms. It might even be the case that language change is driven to sharpen this alignment (though I doubt that the force is very strong (personal opinion) Why? Because FSVs muddies overtime as well as sharpens and lots of FSV is very misleading). But if this is what grammaticalization is, then it can hardly challenge the UG nature of BT as it presupposes it. Grammaticalization is the process whereby the underlying categories of FL/UG act as attractors for overt functional morphology (i.e. LADs try to treat visible functional as reflecting underlying G categories and so over time the surface functional vocabulary will come to (more) perfectly delineate UG cleavages). In fact on this view, CHY, inadvertently, argues FOR the UG nature of BT for it assumes that grammaticalization is the operative process linking M-1 and M-2.

As noted, CHY does not explain what grammaticalization is (nor, to my knowledge has anybody else), though it does note what drives it. It is the usual suspect in such cases; the facilitation of processing (p. 17).  Unfortunately, even were this so (and I am skeptical that this actually means anything), it leaves unexplained how languages like M-2 could exist. After all, if processing ease is a good thing, then why should only M-1 partake?  The answer must be that something stops it from enjoying the fruits of parsing efficiency. What might this be? Well, how about the fact that the PLD only murkily maps the binding relevant FSVs (i.e. the surface forms of the anaphoric morphemes) onto the relevant underlying grammatical categories. But, as noted, if this is what grammaticalization is and what it does, then it is not merely compatible with the view that BT is part of UG and that it is innate, but virtually presupposes that something like this must be the case. Attractors cannot attract without existing.

Let me end here, with a diagnosis of what I take the fundamental error that drives the CHY discussion to be. It is not a new mistake, but one that is, sadly, endemic. It rests on the confusion between Greenberg and Chomsky Universals. CHY assumes that BT aims to catalogue surface distribution of overt morphemes. On this construal, BT is indeed not universal (as even a rabid nativist like me would concede). It is clearly not the case that languages all distinguish overt morphological categories subject to different BT principles. Some languages don’t clearly have a demarcated distinction between overt anaphors or pronominals among their FSVs, some don’t even have dedicated overt functional forms for reflexivization or pronominalization. If one understands UG as committing hostages to surface functional morphology, then CHY is right that BT is not universal. However, this is not how GGers ever understood (or, more accurately, ever should have understood) UG and universals. Chomsky universals are not Greenberg universals. They are more abstract and can be hard to discern from the surface (btw, this is what makes them interesting, IMO). Thus criticizing BT because it is wrong when understood in Greenbergian terms is not much of a fault given that it was not supposed to be so understood (i.e. another Dan Everett moment (see here)).  What is surprising is that the distinction between the two kinds of universals seems so difficult for linguists to grasp. Why is this?

Here’s an unfair (though I believe close to accurate) speculation: it results from the confluence of two powerful factors (i) the attraction of Empiricist conceptions of learning and (ii) the fascination with language diversity.

The first is a horse that I have hobbied on many times before. If you think that acquisition is largely inductive then universals without clear surface reflexes are a challenging concept. Being Eish with a taste for universals leads one to naturally erroneously understand Chomsky universals as Greenbrg universals (as Greenberg universals are the only ones that Eism tolerates).

The second force leading to the confusion between Greenberg and Chomsky universals comes from a fascination with linguistic variation (clearly something that is at the center of CHY). FL/UG rests on the idea that underlyingly there is very little real G variation. If one’s interest is in variation, then this notion of UG will seem way off track. Just look at all the differences! To be told that this just surface morphology will seem unhelpful at best and hostile at worst. The natural response is to look for helpful typological universals and these, not surprisingly will be Greenbergian. Here the generalizations concern surface patterns, as do typological differences if Chomsky’s conception of FL/UG is on the right track. Typological interests do not require embracing a Greenberg conception of universals (unlike a commitment to Eism, which does). However, it is, I believe, a constant temptation. The fact is that a Chomskyan conception of UG is consistent with the view that there are very few (if any) robust typological (i.e. surface true) universals. UG in Chomsky's sense doesn’t need them. It just needs a way of mapping overt forms to underlying forms. In other words, UG needs to be coupled with a theory of acquisition, but this theory does not require that there be surface true universals. Of course, there may be some but they are not conceptually required.

So that’s it. CHY’s conclusions only follow from a flawed understanding of what a universal is and what UG enjoins. The argument is not very good. Sadly, it might well be influential, which is why I spent so much effort trying to dismember it. It appeared in an influential cog sci journal. It will be read as undermining the notion of UG and the relevance of PoS reasoning. It will do so not because the arguments are sound but because this is a welcome conclusion to many. I strongly suggest that GGers educate their psycho counterparts and explain to them why Cognition has once again failed to understand what Chomskyan linguistics is all about. I also suggest that understanding a PoS argument be placed at the center of the field’s pedagogical concerns. It really helps to know how to construct one.


[1] I have no idea why this assumption is so robust among linguists. I don’t believe that anyone ever argued the case and many argued that it was not. So, for example, Chomsky explicitly denies this that one could operationalize grammaticality in terms of (categorical, or otherwise) judgments of acceptability (see, e.g. Current Issues: 7-9 and chapter 3). In fact, there is little reason to believe that there can be operational criteria of FL/UG notions of grammaticality, as holds true for any interesting abstract scientific notions (See Current Issues: 56-7).  If this is correct, then systematic preference judgments might be just as revealing of underlying grammatical form as categorical judgments. At any rate, the assumption that preferences exclusively reflect extra grammatical factors is tendentious. It really depends.
[2] I return on a moment to explicate this term.
[3] It is actually harder to come up with a good story for variable island effects, though as I’ve mentioned before I believe that Kush’s ambiguity hypothesis is likely on the right track.
[4] As in: why does this morphine put you to sleep? In virtue of its dormitive powers. 
I should add that CHY needs to offer an account of what the process consists in at pains of undermining its main argument. The argument is that M-1 and M-2 are sensitive to the same distinction. But this distinction cannot be a Gish one because Gish ones result on categorical judgments. But the M-2 judgments are not categorical therefore the distinction cannot be grammatical. How then to explain m-1 judgments? Well they are categorical because the non G distinction in M-2 has been grammaticalized in M-1. This invites the obvious question: what’s the output of the process of grammaticalization? It sounds like the end product is to render the distinction a grammatical one. But if this is so, then the premise that the distinction is the same in both M-1 and M-2 fails for what is a non-grammatical distinction in M-2 is a grammatical distinction in M-1. The only way to explicate what is going on in the argument is to specify what grammaticalization is and what it does. CHY does not do this.    

Monday, October 12, 2015

What's in UG (part 2)

In the previous post (here) I provided some relevant background for discussing a forthcoming paper in Cognition by Cole, Hermon and Yanti (CHY) that argues against a UG interpretation of the Binding Theory (BT), and, by extension, against most PoS forms of reasoning to UG. Here we get into some details. There is one more post to follow. Told you this was going to be involved.

The paper argues that there are languages where the relatively clean morphological distinctions found in English between anaphors and pronominals is considerably less clear. Thus, there are some languages where some “semantically dependent expressions” (SDE, I return in a moment to explicate this notion) are completely exempt from BT and that there are many where there is no clear morphological distinction between reflexives and pronouns (as their appears to be in English). None of this is news. In fact, as CHY notes, and has been known for quite a while even to typological illiterates like me, there are even languages where the morphological distinction between pronominals and anaphors is hard to discern.[1] In other words, there are languages where the functional surface vocabulary (FSV) relevant for binding does not cleanly reflect the underlying categories relevant for BT. Or, to put this another way, it’s easier to study BT in some languages than in others precisely because in some languages the overt FSV more clearly marks the underlying grammatical distinctions than in other languages. Isn’t this one of the reasons for doing comparative syntax? To find those languages where the system of interest is easiest to study? Does it matter that in some languages the system of interest is muddied by a lousy functional inventory? Not that I can see.[2]

However, CHY’s main argument is that this is a problem for BT and that it shows that UG is not key to understanding binding phenomena and that PoS problems do not support nativist conclusions in the domain of binding. Here’s how CHY puts it (p.3):

UG constitutes a general solution to the problem of the poverty of the stimulus, and would be expected to provide the solution to this dilemma in the realm of Binding, as well as other areas of syntax. If, however, Binding is not determined by UG, and, hence, must be learned by the language learner (as we claim below is in fact the case) then Binding must be learnable pace a priori claims to the contrary. For this reason, Binding constitutes a good test case for the claim that the facts of syntax generally cannot be accounted for without making the assumption that much of our knowledge of syntax is innate (i.e. determined by UG).

So the aim of CHY is to argue that BT facts can do without a UGish BT. However, what CHY actually shows is that it is often hard to fix which overt morphemes/expressions (if any) fall under which parts of BT, i.e. to fix the FSVs of a given language’s G. Thus, the fact that BT does not tell us how to categorize overt morphemes wrt binding categories is taken by CHY to show that a UG BT cannot play a part in solving PoS problems in the domain of binding. In other words, CHY argues that because BT is not by itself a full theory of acquisition, (in particular does not provide a general account of fixing FSV) then it does nothing at all. And this is a very poor argument against UG.

Let me be clear exactly what my criticism of CHY is: nowhere does CHY outline just what the PoS problem in the domain of binding is (as I did in the previous post). It doesn’t discuss how facts like (3) in the post here, could be learned purely on the basis of PLD. Nor does CHY argue that the binding principles would not be useful in categorizing morphemes. The CHY argument is entirely based on the observation that the BT categories are often opaque (i.e. that the mapping from overt FSVs to BT relevant categories is often unclear).[3] However, as noted above, this cannot be an argument against the UG nature of BT because the claim that BT is part of UG never assumed that it could solve this problem (all by itself, though it almost surely contributes towards solving it). And a good thing too for we know that whatever is in UG it had better not be a specification of which language particular morphemes are subject to BT-A, B or C. If anything is learned, it’s this.  So, the CHY conclusions do not, as a matter of logic, argue against the claim that BT is part of FL/UG. Let’s consider some further details.

The CHY argument comes in three parts. Part one is the observation that though many languages have a BT functional vocabulary that clearly separates BT anaphors from BT pronominals (see CHY section 2), this is not universally true.[4] In section 3 CHY notes that there are languages (a version of Javanese) with BT exempt anaphors, expressions that are anaphoric (in one sense) but do not fall under BT at all. CHY takes this to be problematic. In particular, CHY claims that the existence of BT exempt anaphors (p. 7):

…constitutes a serious challenge for UG-based approaches to Binding. The presence in a language of a form that is used anaphorically but which is exempt from the Binding requirements of UG would impose a considerable burden on the child acquiring the language…The existence of UG sanctioned categories for anaphora simplifies learning only if all anaphoric elements (in the non technical sense of “anaphoric” that includes both pronouns and reflexives) are subject to UG principles. Thus, in the context of a system containing a mixture of UG compliant and UG exempt elements, UG principles do not provide a solution to the poverty of stimulus problem. At the least, the distinction between exempt and non- exempt forms must be learned from experience…and it is doubtful that the data is sufficiently structured to make this distinction learnable on the basis of the distributional data that are available to the child.

What’s the problem? Not all forms that are “anaphoric” in the “non-technical sense” (i.e. they are semantic but not syntactic dependents) are anaphoric in the BT sense (i.e. subject to BT-A) and, CHY claims, there is no PLD that can help the PLD figure out if an expression belongs in one category or another. Let’s consider these claims much more closely, for they come close to inviting a form of magical thinking.

First, what is the “non-technical sense” of anaphoric? Linguists have two senses of “anaphoric.” The first is “semantic.” An expression is an anaphoric dependent of another expression if its interpretation presupposes/requires/relies on the interpretation of another. This is what I meant by SDEs above. English contains expressions like this. For example, ‘the others’ is an expression that has no interpretation sans a semantic antecedent (i.e. it must have a semantic antecedent). However, it is not subject to BT-A as it can have a non sentence internal antecedent. In fact, all non-deictic pronouns (e.g. cataphoric pronouns) are another example.

(4)  a. Three of the men are wearing tuxedos. The others are sporting jeans.
b. John ate an ice cream. Mary then kissed him.

In (4a), the others depends for its interpretation on three of the men in the previous sentence, as does him on John in (4b). These, then are semantically anaphoric, but BT-A exempt. Are these then a problem for a UG account of BT-A? Only if categorizing them as exempt from BT-A is problematic. So is it? Is there no available positive PLD that might inform the LAD that despite appearances, these expressions are not subject to BT-A (i.e. that despite being semantic anaphors they are not syntactic anaphors)? Well, how about the data in (4). Are these too exotic to be considered part of the PLD? Not from where I sit. Of course, I could be wrong, but it is certainly not “doubtful” that such data might be available to the LAD. In fact, I would go further: cross sentential dependencies of the sort in (4) are almost certainly part of the PLD. [5]

Is Javanese any different? Maybe, but as CHY notes (p. 6) the BT exempt anaphor it discusses can have discourse antecedents, in contrast to the BT-A forms that cannot. But this then is exactly the kind of evidence that would advise the LAD against categorizing them as syntactically anaphoric (i.e. subject to BT-A). If so, there is no PoS problem as regards the categorization of these overt expressions into more abstract syntactic categories (i.e. treating them (or not) as FSVs for the abstract grammatical category BT-anaphor).

Conclusion: CHY’s argument here fails. It is certainly right in concluding that the world would be a better place for the LAD if semantic anaphora were an infallible indicator of syntactic anaphora. A world where visible diagnostics are perfect indicators of underlying structure is a nice place (of course, this is no less true for measles and quarks than for BT-A anaphors). But, this does not mean that a world where the two pull apart (i.e. our own) implies that the distinction cannot be acquired on the basis of PLD. It all depends on the PLD and the learner, and CHY does not consider its own cited data in concluding that the PLD is sadly lacking. Is cross-sentential “anaphora” that exotic in the PLD that the LAD would never have access to it? Maybe, but I would like a lot more than simple assertion to that effect before concluding that it is so.

In fact, we can go further. CHY must assume that this categorization can be fixed on the basis of PLD (contrary to their apparent claim to the contrary). Why so? CHY argues that distinguishing BT-A from BT-A-exempt FSVs is not based on innate features of the LAD. That, in fact, is its main point (which, btw, must be correct given the obvious variation in surface forms). But if it’s not fixed via FL/UG then it must be learned. But to be learned there needs to be evidence that the LAD can use to learn it. Thus, there must be evidence in the PLD relevant to fixing the categorization. I mention this as CHY appears to deny this. Specifically CHY asserts (p. 7):

At the least, the distinction between exempt and non- exempt forms must be learned from experience…and it is doubtful that the data is sufficiently structured to make this distinction learnable on the basis of the distributional data that are available to the child

But as a matter of logic, this pair of claims is close to contradictory (That’s being weaselish. IMO it is a contradiction). If the categorization is learned, then the PLD must be “sufficiently structured” to allow it to be learned. And if the PLD is not “sufficiently structured” to allow it to be learned, then it must be innately determined. There is no third alternative.  There is no learning fairy that fits neatly between these possibilities.[6] Luckily, relevant data appears to be plausibly accessible, and so there is no clear PoS problem as regards binning potential “anaphoric” FSVs into different BT relevant categories.

Last point: recall, that even where it very hard to fix BT exemption on the basis of PLD it would not argue against the classical BT being part of UG. Why? Because CHY does not address how BT-A compliant FSVs acquire all of their distributional properties? In specific, CHY does not address how native speakers acquire competence wrt the data in (3) (see previous post). Recall that this is the PoS problem that BT was developed to address. And in this case, as we outlined in the previous post, there is a clear PoS problem. Thus, the LAD cannot inductively acquire knowledge of facts like (3) (in earlier post) as there is no relevant data in the PLD for fixing this knowledge. This argument stands regardless of how the problems with BT-A exempt expressions is resolves.  In short, it is a very odd criticism of any theory that because it (BT in this case) does not solve a problem that it was never intended to address that it also does not solve the problem that it was constructed to address. And, to further conclude, that because some other theory (not yet provided I may add) solves a problem BT was not constructed to address that the same solution obviously solves the problem that BT did address.

Conclusion: this argument comes nowhere close to showing anything about the UG status of BT. And, for the record, I don’t believe that CHY has shown that categorizing FSVs as BT exempt cannot be inferred from a judicious use of the PLD (and, as I noted above, CHY can’t really assume this either).[7]

The next and last post addresses CHY’s second argument.


[1] The most well-known cases involve languages with so called “long distance anaphors.” Are these really anaphors or just bound pronouns? What’s the difference? Are they both or neither? Tough questions. But does linguistic theory need to provide strict answers? Or does it only need to provide ways of classifying things to the degree that they can be so classified? Can you spot the rhetorical question?
[2] Nor is looking for the easiest window into the mechanism limited to linguists. Ask your favorite neuroscientist about squid axons sometime and see why they were a research favorite for such a long time.
[3] For example, on p 9 CHY states:

…the division of anaphora into reflexives and pronouns cannot be simply a matter of compliance with UG principles…[this] suggest strongly that the pattern modeled in Binding Theory is not primarily due to the interaction among principles of UG…

This suggests that CHY (wrongly) understands BT as aiming to explain how FSV gets mapped into BT categories. And CHY is right. BT does not do this. But it does other things, like explain the pattern in (3), which is not attested in the PLD.

[4] Oddly, IMO, CHY group English and Chinese together as well-behaved languages wrt BT. However, if English is the BT poster child (which it really shouldn’t be, see below) then Chinese makes things less transparent. So, for many speakers, ‘ziji’ can serve both as the overt form of the reflexive and of the long distance anaphor. True, ‘ta ziji’ distributes much like ‘himself,’ but Chinese has a simple ‘self’ form that English does not have and this might well make the categorization problem in Chinese harder than it is in English, at least for the canonical data. Of course, not even English is that well behaved as the well-known facts about picture NPs indicates.
[5] JL, my friendly local acquisitionist, tells me that I am correct in thinking that such data are available in the PLD.
[6] There is a third option, but it is irrelevant to CHY’s concerns. It could be that some feature is not specified by PLD or intrinsically fixed by LAD. In such a case, the “parameter” could be random in the population (see Han, Lidz and Musolino LI (2007) for a worked out example in Korean “dialects”). However, even in this case, there is an innate learning principle at work: if no data relevant to P then flip a coin and set P among UG available options. As I noted, this possibility is not relevant for the CHY argument.
[7] CHY argue against the hypothesis that the BT exempt anaphor is ambiguously either a local anaphor or a pronoun depending on context. I sort of liked this proposal. CHY argues against it. It argues instead that the exempt anaphor is under-specified wrt being an anaphor and a pronominal. CHY proposes that this under-specification allows such anaphors to be simultaneously interpreted anaphors and pronouns (as opposed to being one or the other in any given context). The argument against this involves strict vs sloppy readings under ellipsis. CHY assumes that BT-A anaphors only license sloppy readings under ellipsis. It further notes that the BT exempt anaphor in Javanese it discusses allows strict readings even when the anaphor in the ellipsis licensing antecedent would be locally bound (i.e. even when the ambiguity thesis should treat the relevant structure as one of unambiguous BT-A binding).

The form of this argument is fine (or sorta fine: it relies on an ad hoc stipulation that expressions that are underdetermined wrt anaphor and pronominal features is simultaneously both anaphoric and pronominal semantically. So far as I know, this finely-tuned assumption does not follow from anything else about binding or under-specification, and so is ad hoc.). However the premise seems to me empirically ill founded. Assume, as CHY does, that English is well behaved wrt BT-A then we predict that reflexive binding should never license strict readings under ellipsis in English. However this is false (or there is some challenging counter-evidence). One standard case where the strict reading is acceptable is in (i):
(i)             John1 defended himself1 more competently than his lawyer did (defend him1)
In (i) the “defend John” reading is fine. For me, ditto for sentences like (ii) and (iii).  
(ii)           John1 defended himself1 ably in court. His1 lawyers did not (defend him1 ably)
(iii)          John1 admires himself1 greatly. Nobody else does (admire him1).
Thus, the diagnostic tool CHY uses to argue against the ambiguity thesis is not correct in general as paradigmatic reflexive antecedents can license strict readings in ellipsis sites.

Why?  Well this requires some theory of ellipsis (which CHY does not provide). The technical issue of relevance will be what licenses ellipsis. One standard theory sees ellipsis as deletion under identity. What are the relevant identity parameters? Say that morpho-phonological identity is one such, then the fact that the same FSV in Javanese is ambiguous between a reflexive and a pronoun could allow the strict reading to be readily available (the “reflexive” antecedent would be morpho-phonologically identical to the “pronominal” in the ellipsis site). The expression in the ellipsis site interpreted as a pronoun would be morphologically identical to the expression in the antecedent clause with the expression interpreted as a reflexive. It is well known that such conditions are grammatically relevant (e.g. ATB and Parasitic Gap constructions often require morphological case identity to be licit).
            Last point: it is worth observing that the strict/sloppy dichotomy CHY uses as a diagnostic is not part of BT. In fact, BT is mum concerning these matters. It is a strictly empirical question, which CHY argues one way. My point is not that CHY is wrong, but that the premise is dubious and this leaves the ambiguity hypothesis alive and kicking.

Wednesday, October 7, 2015

What's in UG (part 1)?

This is the first of three posts on a forthcoming Cognition paper arguing against UG. The specific argument is against the Binding Theory. But the form is intended to generalize. The paper is written by excellent linguists, which is precisely why I spend three posts exposing its weaknesses. The paper, because it will appear in Cognition, is likely to be influential. It shouldn’t be. Here’s the first of three posts explaining why.

Let’s start with some truisms: not every property of a language particular G is innate. Here’s another one: some features of G reflect innate properties of the language acquisition device (LAD). Let’s end with a truth (that should be a truism by now but is still contested by some for reasons that are barely comprehensible): some of the innate LAD structure key to acquiring a G is linguistically dedicated (i.e. not cognitively general (i.e. due to UG)). These three claims should be obvious. True truisms. Sadly, they are not everywhere and always recognized as such. Not even by extremely talented linguists. I don’t know why this is so (though I will speculate towards the end of this note), but it is. Recent evidence comes from a forthcoming paper in Cognition (here) by Cole, Hermon and Yanti (CHY) on the UG status of the Binding Theory (BT).[1] The CHY argument is that BT cannot explain certain facts in a certain set of Javanese and Malay dialects. It concludes that binding cannot be innate. The very strong implication is that UG contains nothing like BT, and that even if it did it would not help explain how languages differ and how kids acquire their Gs. IMO, this implication is what got the paper into Cognition (anything that ends with the statement or implication that there is nothing special about language (i.e. Chomsky is wrong!!!) has a special preferential HOV lane in the new Cognition’s review process). Boy do I miss Jacques Mehler. Come back Jacques. Please.

Before getting into the details of CHY, let’s consider what the classical BT says.[2] It is divided into three principles and a definition of binding:

A.   An anaphor must be bound in its domain
B.    A pronominal cannot be bound in its domain
C.    An R-expression cannot be bound

(1)  An expression E binds an expression E’ iff E c-commands E’ and E is co-indexed with E’.

We also need a definition of ‘domain’ but I leave it to the reader to pick her/his favorite one. That’s the classical BT.

What does it say? It outlines a set of relations that must hold between classes of grammatical expressions. BT-A states that if some expression is in the grammatical category ‘anaphor’ then it must have a local c-commanding binder. BT-B states that if some expression is in the category ‘pronominal’ then it cannot have a local c-commanding binder. And BT-C states, well you know what it states, if

Now what does BT not say? It says nothing about which phonetically visible expressions fall into which class. It does not say that every overt expression must fall into at least one of these classes. It does not say that every G must contain expressions that fall into these classes. In fact, BT by itself says nothing at all about how a given “visible” morphologically/phonetically visible expression distributes or what licensing conditions it must enter into. In other words, by itself BT does not tell us, for example, that (2) is ungrammatical. All it says is that if ‘herself’ is an anaphor then it needs a binder. That’s it.

            (2) John likes herself

How then does BT gain empirical traction? It does so via the further assumption that reflexives in English are BT anaphors (and, additionally, that binding triggers morphologically overt agreement in English reflexives). Assuming this, ‘herself’ is subject to principle BT-A and assuming that John is masculine, herself has no binder in its domain, and so violates BT-A above. This means that the structure underlying (2) is ungrammatical and this is signaled by (2)’s unacceptability.

As stated, there is a considerable distance between a linguistic object’s surface form and its underlying grammatical one. So what’s the empirical advantage of assuming something as abstract as the classical BT? The most important reason, IMO, is that it helps resolve a critical Poverty of Stimulus (PoS) problem. Let me explain (and I will do this slowly for CHY never actually explains what the specific PoS problem in the domain of binding is (though they allude to the problem as an important feature of their investigation), and this, IMO, allows the paper to end in intellectually unfortunate places).

As BT connoisseurs know, the distribution of overt reflexives and pronouns is quite restricted. Here is the standard data:[3]

(3) a. John1 likes herself1/*2
b. John1 likes himself1/*2
c. John1 talked to Bill2 about himself1/2/*3
d. John1 expects Mary2 to like himself*1/*2/*3
e. John1 expects Mary2 to like herself*1/2/*3
f. John1 expects himself1/*2/*3 to like Mary2
g. John1 expects (that) he/himself*1/*2/*3 will like Mary2

If we assume that reflexives are BT-A-anaphors then we can explain all of this data. Where’s the PoS problem? Well, lots of these data concern what cannot happen. On the assumption that the ungrammatical cases in (3) are not attested in the PLD, then the fact that a typical English Language Acquisition Device (LAD, aka, kid) converges on the grammatical profile outlined in (3) must mean that this profile in part reflects intrinsic features of the LAD. For example, the fact that kids do not generalize from the acceptability of (3f) to conclude that (3g) should also be acceptable needs to be explained and it is implausible that the LAD infers that that this is an incorrect inference by inspecting unacceptable sentences like (3g), for being unacceptable they will not appear in the PLD.[4] Thus, how LADs come to converge to Gs that allow the good sentences and prevent the bad ones looks like (because it is) a standard PoS puzzle.

How does assuming that BT is part of UG solve the problem? Well, it doesn’t, not all by itself (and nobody ever thought that it could all by itself). But it radically changes it. Here’s what I mean.

If BT is part of UG then the acquisition problem facing the LAD boils down to identifying those expressions in your language that are anaphors, pronominals and R-expressions. This is not an easy task, but it is easier than figuring this out plus figuring out the data distribution in (3). In fact, as I doubt that there is any PLD able to fix the data in (3) (this is after all what the PoS problem in the binding domain consists in) and as it is obvious that any theory of binding will need to have the LAD figure out (i.e. learn) using the PLD which overt morphemes (if any) are BT anaphors/pronominals (after all, ‘himself’ is a reflexive in English but not in French and I assume that this fact must be acquired on the basis of PLD) then the best story wrt Plato’s Problem in the domain of binding is where what must obviously be learned is all that must be learned. Why? Because once I know that reflexives in English are BT anaphors subject to BT-A then I get the knowledge illustrated by the data in (3) as a UG bonus.  That’s how PoS problems are solved.[5] So, to repeat: all the LAD needs do to become binding competent is figure out which overt expressions fall into which binding categories. Do this and the rest is an epistemic freebie.

Furthermore, it’s virtually certain that the UG BT principles act as useful guides for the categorization of morphemes into the abstract categories BT trucks in (i.e. anaphor, pronominal, and R-expression).  Take anaphors. If BT is part of UG it provides the LAD with some diagnostics for anaphoricity. Anaphors must have antecedents. They must be local and high enough. This means that if the LAD hears a sentence like John scratched himself in a situation where John is indeed scratching himself then he has prima facie evidence that ‘himself’ is a reflexive (as it fits A constraints). Of course, the LAD may be wrong (hence the ‘prima facie’ above). For example, say that the LAD also hears pairs of sentences like John loves Mary. She loves himself too and ‘himself’ here is anaphoric to John, then the LAD has evidence that reflexives are not just subject to BT-A (i.e. they are at best ambiguous morphemes and at worst not subject to BT-A at all). So, I can see how PLD of the right sort in conjunction with an innate UG provided BT-A would help with the classification of morphemes to the more abstract categories using simple PLD in the.[6]  That’s another nice feature of an articulate UG.

Please observe: on this view of things UG is an important part of a theory of language learning. It is not itself a theory of learning. This point was made in Aspects, and is as true today as it was then. In fact, you might say that in the current climate of Bayesian excess that it is the obvious conclusion to draw: UG limns the hyporthesis space that the learning procedure explores. There are many current models of how UG knowledge might be incorporated in more explicit learning accounts of various flavors (see Charles Yang’s work or Jeff Lidz’s stuff for some recent general proposals and worked out examples).

Does any of this suppose that the LAD uses only attested BT patterns in learning to classify expressions? Of course not. For example, the LAD might conclude that ‘itself’ is a BT-A anaphor in English on first encountering it. Why? By generalizing from forms it has encountered before (e.g. ‘herself’, ‘themselves’). Here the generalization is guided not by UG binding properties but by the details of English morphology.  It is easy to imagine other useful learning strategies (see note 6). However, it seems likely that one way the LAD will distinguish BT-A from BT-B morphemes will be in terms of their cataphoric possibilities positively evidenced in the PLD.

So, BT as part of UG can indeed help solve a PoS problem (by simplifying what needs to be acquired) and plausibly provides guide-posts towards that classification. However, BT does not suffice to fix knowledge of binding all by itself nor did anyone ever think that it would.  Moreover, even the most rabid linguistic nativist (I know because I am one of these) is not committed to any particular pattern of surface data. To repeat, BT does not imply anything about how morphemes fall into any of the relevant categories or even if any of them do or even if there are any relevant surface categories to fall into.

With this as background, we are now ready to discuss CHY. I will do this in the next post.


[1] I have been a great admirer of both Cole and Hermon’s work for a long time. They are extremely good linguists, much better than I could ever hope to be. This paper, however, is not good at all. It’s the paper, not the people, that this post discusses.
[2] I will discuss the GB version for this is what CHY discusses. I personally believe that this version of BT is reducible to the theory of movement (A-chain dependencies actually). The story I favor looks more like the old Lees & Klima account. I hope to blog about the differences in the very near future.
[3] As GGers also know, the judgments effectively reverse if we replace the reflexive with a bound pronoun. This reflects the fact that in languages like English, reflexives and bound pronouns are (roughly) in complementary distribution. This fact results from the opposite requirements stated in BT-A and BT-B. The same effect was achieved in earlier theories of binding (e.g. Lees and Klima) by other means.
[4] From what I know, sentences like (3g) are unattested in CHILDES. Indeed, though I don’t know this, I suspect that sentences with reflexives in ECM subject position are not a dime a dozen either.
[5] I assume that I need not say that once one figures out which (if any) of the morphemes are pronominals then BT-B effects (the opposite of those in (3) with pronouns replacing reflexives) follow apace. As I need not say this, I won’t.
[6] Please note that this is simply an illustration, not a full proposal. There are many wrinkles one could add. Here’s another potential learning principle: LADs are predisposed to analyze dependencies in BT terms if this is possible. Thus the default analysis is to treat a dependency as a BT dependency. But this principle, again, is not an assumption properly part of BT. It is part of the learning theory that incorporates a UG BT.