Comments

Showing posts with label UG. Show all posts
Showing posts with label UG. Show all posts

Wednesday, October 7, 2015

What's in UG (part 1)?

This is the first of three posts on a forthcoming Cognition paper arguing against UG. The specific argument is against the Binding Theory. But the form is intended to generalize. The paper is written by excellent linguists, which is precisely why I spend three posts exposing its weaknesses. The paper, because it will appear in Cognition, is likely to be influential. It shouldn’t be. Here’s the first of three posts explaining why.

Let’s start with some truisms: not every property of a language particular G is innate. Here’s another one: some features of G reflect innate properties of the language acquisition device (LAD). Let’s end with a truth (that should be a truism by now but is still contested by some for reasons that are barely comprehensible): some of the innate LAD structure key to acquiring a G is linguistically dedicated (i.e. not cognitively general (i.e. due to UG)). These three claims should be obvious. True truisms. Sadly, they are not everywhere and always recognized as such. Not even by extremely talented linguists. I don’t know why this is so (though I will speculate towards the end of this note), but it is. Recent evidence comes from a forthcoming paper in Cognition (here) by Cole, Hermon and Yanti (CHY) on the UG status of the Binding Theory (BT).[1] The CHY argument is that BT cannot explain certain facts in a certain set of Javanese and Malay dialects. It concludes that binding cannot be innate. The very strong implication is that UG contains nothing like BT, and that even if it did it would not help explain how languages differ and how kids acquire their Gs. IMO, this implication is what got the paper into Cognition (anything that ends with the statement or implication that there is nothing special about language (i.e. Chomsky is wrong!!!) has a special preferential HOV lane in the new Cognition’s review process). Boy do I miss Jacques Mehler. Come back Jacques. Please.

Before getting into the details of CHY, let’s consider what the classical BT says.[2] It is divided into three principles and a definition of binding:

A.   An anaphor must be bound in its domain
B.    A pronominal cannot be bound in its domain
C.    An R-expression cannot be bound

(1)  An expression E binds an expression E’ iff E c-commands E’ and E is co-indexed with E’.

We also need a definition of ‘domain’ but I leave it to the reader to pick her/his favorite one. That’s the classical BT.

What does it say? It outlines a set of relations that must hold between classes of grammatical expressions. BT-A states that if some expression is in the grammatical category ‘anaphor’ then it must have a local c-commanding binder. BT-B states that if some expression is in the category ‘pronominal’ then it cannot have a local c-commanding binder. And BT-C states, well you know what it states, if…

Now what does BT not say? It says nothing about which phonetically visible expressions fall into which class. It does not say that every overt expression must fall into at least one of these classes. It does not say that every G must contain expressions that fall into these classes. In fact, BT by itself says nothing at all about how a given “visible” morphologically/phonetically visible expression distributes or what licensing conditions it must enter into. In other words, by itself BT does not tell us, for example, that (2) is ungrammatical. All it says is that if ‘herself’ is an anaphor then it needs a binder. That’s it.

            (2) John likes herself

How then does BT gain empirical traction? It does so via the further assumption that reflexives in English are BT anaphors (and, additionally, that binding triggers morphologically overt agreement in English reflexives). Assuming this, ‘herself’ is subject to principle BT-A and assuming that John is masculine, herself has no binder in its domain, and so violates BT-A above. This means that the structure underlying (2) is ungrammatical and this is signaled by (2)’s unacceptability.

As stated, there is a considerable distance between a linguistic object’s surface form and its underlying grammatical one. So what’s the empirical advantage of assuming something as abstract as the classical BT? The most important reason, IMO, is that it helps resolve a critical Poverty of Stimulus (PoS) problem. Let me explain (and I will do this slowly for CHY never actually explains what the specific PoS problem in the domain of binding is (though they allude to the problem as an important feature of their investigation), and this, IMO, allows the paper to end in intellectually unfortunate places).

As BT connoisseurs know, the distribution of overt reflexives and pronouns is quite restricted. Here is the standard data:[3]

(3) a. John1 likes herself1/*2
b. John1 likes himself1/*2
c. John1 talked to Bill2 about himself1/2/*3
d. John1 expects Mary2 to like himself*1/*2/*3
e. John1 expects Mary2 to like herself*1/2/*3
f. John1 expects himself1/*2/*3 to like Mary2
g. John1 expects (that) he/himself*1/*2/*3 will like Mary2

If we assume that reflexives are BT-A-anaphors then we can explain all of this data. Where’s the PoS problem? Well, lots of these data concern what cannot happen. On the assumption that the ungrammatical cases in (3) are not attested in the PLD, then the fact that a typical English Language Acquisition Device (LAD, aka, kid) converges on the grammatical profile outlined in (3) must mean that this profile in part reflects intrinsic features of the LAD. For example, the fact that kids do not generalize from the acceptability of (3f) to conclude that (3g) should also be acceptable needs to be explained and it is implausible that the LAD infers that that this is an incorrect inference by inspecting unacceptable sentences like (3g), for being unacceptable they will not appear in the PLD.[4] Thus, how LADs come to converge to Gs that allow the good sentences and prevent the bad ones looks like (because it is) a standard PoS puzzle.

How does assuming that BT is part of UG solve the problem? Well, it doesn’t, not all by itself (and nobody ever thought that it could all by itself). But it radically changes it. Here’s what I mean.

If BT is part of UG then the acquisition problem facing the LAD boils down to identifying those expressions in your language that are anaphors, pronominals and R-expressions. This is not an easy task, but it is easier than figuring this out plus figuring out the data distribution in (3). In fact, as I doubt that there is any PLD able to fix the data in (3) (this is after all what the PoS problem in the binding domain consists in) and as it is obvious that any theory of binding will need to have the LAD figure out (i.e. learn) using the PLD which overt morphemes (if any) are BT anaphors/pronominals (after all, ‘himself’ is a reflexive in English but not in French and I assume that this fact must be acquired on the basis of PLD) then the best story wrt Plato’s Problem in the domain of binding is where what must obviously be learned is all that must be learned. Why? Because once I know that reflexives in English are BT anaphors subject to BT-A then I get the knowledge illustrated by the data in (3) as a UG bonus.  That’s how PoS problems are solved.[5] So, to repeat: all the LAD needs do to become binding competent is figure out which overt expressions fall into which binding categories. Do this and the rest is an epistemic freebie.

Furthermore, it’s virtually certain that the UG BT principles act as useful guides for the categorization of morphemes into the abstract categories BT trucks in (i.e. anaphor, pronominal, and R-expression).  Take anaphors. If BT is part of UG it provides the LAD with some diagnostics for anaphoricity. Anaphors must have antecedents. They must be local and high enough. This means that if the LAD hears a sentence like John scratched himself in a situation where John is indeed scratching himself then he has prima facie evidence that ‘himself’ is a reflexive (as it fits A constraints). Of course, the LAD may be wrong (hence the ‘prima facie’ above). For example, say that the LAD also hears pairs of sentences like John loves Mary. She loves himself too and ‘himself’ here is anaphoric to John, then the LAD has evidence that reflexives are not just subject to BT-A (i.e. they are at best ambiguous morphemes and at worst not subject to BT-A at all). So, I can see how PLD of the right sort in conjunction with an innate UG provided BT-A would help with the classification of morphemes to the more abstract categories using simple PLD in the.[6]  That’s another nice feature of an articulate UG.

Please observe: on this view of things UG is an important part of a theory of language learning. It is not itself a theory of learning. This point was made in Aspects, and is as true today as it was then. In fact, you might say that in the current climate of Bayesian excess that it is the obvious conclusion to draw: UG limns the hyporthesis space that the learning procedure explores. There are many current models of how UG knowledge might be incorporated in more explicit learning accounts of various flavors (see Charles Yang’s work or Jeff Lidz’s stuff for some recent general proposals and worked out examples).

Does any of this suppose that the LAD uses only attested BT patterns in learning to classify expressions? Of course not. For example, the LAD might conclude that ‘itself’ is a BT-A anaphor in English on first encountering it. Why? By generalizing from forms it has encountered before (e.g. ‘herself’, ‘themselves’). Here the generalization is guided not by UG binding properties but by the details of English morphology.  It is easy to imagine other useful learning strategies (see note 6). However, it seems likely that one way the LAD will distinguish BT-A from BT-B morphemes will be in terms of their cataphoric possibilities positively evidenced in the PLD.

So, BT as part of UG can indeed help solve a PoS problem (by simplifying what needs to be acquired) and plausibly provides guide-posts towards that classification. However, BT does not suffice to fix knowledge of binding all by itself nor did anyone ever think that it would.  Moreover, even the most rabid linguistic nativist (I know because I am one of these) is not committed to any particular pattern of surface data. To repeat, BT does not imply anything about how morphemes fall into any of the relevant categories or even if any of them do or even if there are any relevant surface categories to fall into.

With this as background, we are now ready to discuss CHY. I will do this in the next post.


[1] I have been a great admirer of both Cole and Hermon’s work for a long time. They are extremely good linguists, much better than I could ever hope to be. This paper, however, is not good at all. It’s the paper, not the people, that this post discusses.
[2] I will discuss the GB version for this is what CHY discusses. I personally believe that this version of BT is reducible to the theory of movement (A-chain dependencies actually). The story I favor looks more like the old Lees & Klima account. I hope to blog about the differences in the very near future.
[3] As GGers also know, the judgments effectively reverse if we replace the reflexive with a bound pronoun. This reflects the fact that in languages like English, reflexives and bound pronouns are (roughly) in complementary distribution. This fact results from the opposite requirements stated in BT-A and BT-B. The same effect was achieved in earlier theories of binding (e.g. Lees and Klima) by other means.
[4] From what I know, sentences like (3g) are unattested in CHILDES. Indeed, though I don’t know this, I suspect that sentences with reflexives in ECM subject position are not a dime a dozen either.
[5] I assume that I need not say that once one figures out which (if any) of the morphemes are pronominals then BT-B effects (the opposite of those in (3) with pronouns replacing reflexives) follow apace. As I need not say this, I won’t.
[6] Please note that this is simply an illustration, not a full proposal. There are many wrinkles one could add. Here’s another potential learning principle: LADs are predisposed to analyze dependencies in BT terms if this is possible. Thus the default analysis is to treat a dependency as a BT dependency. But this principle, again, is not an assumption properly part of BT. It is part of the learning theory that incorporates a UG BT.

Sunday, October 13, 2013

The Merge Conspiracy [Part 3]

Part 2 and a half showed us the consequences of the MSO-Merge correspondence, both good and bad. In an ideal world, there should be a way to curtail the bad without limiting the good. Alas, ever since J.J. Abrams got his hands on Star Trek I have been uncertain about the degree of idealness of the world we inhabit.

Wednesday, October 3, 2012

Universal Grammar



Everyone believes that humans have a Universal Grammar (UG). Why?  Because it is a one step conclusion licensed by a trivial (and it is trivial) inference from one obvious factual premise (viz. humans are linguistically capable beings) and one major premise (viz. if humans are linguistically capable then there are some mental properties on which this capacity rests).  As UG is the name we give to these mental properties there cannot be a real debate about whether humans have a UG.  What has been contentious is what UG looks like. In what follows I discuss two features that generative linguists attribute to it: (i) UG is exclusive to humans and (ii) UG is linguistically specific.  Chomsky, for example, has claimed both properties for UG. Let’s consider these in turn.

First what does “species-specific” mean?  One thing it means is that UG is a property of humans the way that bipedalism or opposable thumbs are.  Human genetics insures that individual humans normally (i.e. exempting pathological cases) come equipped with the capacity to walk erect, to grab stuff and to acquire and use a language. So, just as tigers are biologically built to have stripes, and salmon to return to their birthplaces to spawn so too humans come biologically equipped to develop linguistic facility.  The empirical basis for this observation is overwhelming and not at all subtle. Anyone who observes language acquisition in the wild cannot fail to notice that human children (regardless of socio-economic status, religious affiliation, birth marks, head size, overall IQ or anything else short of pathology) when reared in a linguistic environment come to acquire linguistic competence in the native language they are exposed to. This observational truism leads smoothly to a related truism: that the capacity to develop such linguistic facility is due to mental equipment that individual humans share simply in virtue of being human. 

These truisms conceded, even if UG is species specific to humans it does not imply that UG is exclusively a human endowment. After all the observation that humans come genetically packaged with a four chamber heart does not imply that other animals do not.  Nonetheless, as a matter of fact, it appears that whatever humans have that allows them to develop linguistic competence is not widely shared.  Concretely, so far as we can tell (and take it from me, many investigators have tried to tell!), nothing does language like humans do. Apes don’t. Dolphins don’t. Parrots don’t. Or at least they don’t obviously. While it takes considerable effort to show that other animals show language-like behavior, nobody will win a Nobel for demonstrating that 5 year olds (or even 2 year olds) talk.  Thus, it’s a safe bet that whatever is going on in the human case is qualitatively different from what we see in other animals.

This said, it is worth noting that the program of describing UG wouldn’t change much were it established that other animals had one. Depending on which animals it was it might raise additional questions of how these UGs evolved (other apes? look for common ancestor; apes and dolphins? look for language as correlate of brain size; birds and bees? who knows). If other animals talked we could (at least in principle) investigate UG by studying how they acquired and used language, though the difficulties of studying UG in this way should not be underestimated. The biggest bonus would likely arise if we decided to treat these non-human talkers as possible targets for the kinds of experiments that are morally and legally forbidden on humans (cut them up, put them in Skinner boxes), though if they really talked like us we might be squeamish about treating them the way we treat white mice, chimps and cute bunny rabbits, though considering the unquenchable (blood thirsty?) desire for pure knowledge that homo sapiens regularly displays even pleading animals might not be safe from our inquisitive minds.  However, excluding such scenarios, finding another species that talked just like we do would not substantially change the research problem. In fact, it would not make it appreciably different from studying UG by investigating the grammatical properties of different languages (English, Chinese, ASL etc.), something that generative grammarians already do in spades. So though it appears as a matter of fact that UG is exclusively a feature of humans, if the aim is to describe UG it does not much matter that this is so.

Let’s now consider the suggestion that UG is a linguistically specific capacity, to be understood as the claim that UG’s cognitive mechanisms are sui generis, different from the cognitive mechanisms at work in other areas of cognition. There is a stronger and a weaker version of this claim. The stronger one is that all (or most or many) of the cognitive powers that go into linguistic competence differ from those that support other cognitive capacities (e.g. the capacity to identify objects, recognize and “read” other minds, understand causal interactions, navigate home, keep track of where and when you hid your food, etc.). The weak one is that UG enjoys at least one cognitively distinctive feature.  Current speculation among generativists leans towards the weak claim. Much current research in syntax (especially that which flies under the flag of the Minimalist Program) aims to reduce the linguistically specific mechanisms of UG to a small core.  Chomsky, for example, has argued that the only real distinctive feature of UG is (hierarchical) recursion (a product of the operation Merge), the property whereby the outputs of rules can be treated as inputs to these same rules. This allows for the generation of endlessly large linguistic objects, a fact that sits well with the observation that there appears to be no upper bound on the size of admissible phrases and sentences in natural languages. 

How reasonable is this second claim?  To my mind, it is almost ineluctable for the following reasons. First, if as discussed above, humans have UG but other animals do not then one plausible reason for this is that humans have at least one mental power that other animals don’t. The alternative is that human cognition is not qualitatively different from that of other animals but only quantitatively so; all animals share the same basic mechanisms just that humans have more horse-power under the cranial hood.  This
option is a favorite of those excited by general learning theories. They tend to be of an empiricist bent (something we will discuss in a later post).   On this view, the same cognitive powers are used in every area of cognition, including language.  There are two kinds of puzzles this empiricist conception runs into in the domain of language. First, the species specificity problem noted above; why do only humans talk?  The second is the separability of linguistic competence from other forms of cognition. It appears that linguistic competence is independent of most other kinds of cognitive competence, e.g. IQ, face recognition, etc. Why so if all involve the same general all purpose cognitive powers? Were linguistic competence a product of general cognitive factors it would be natural to expect that success in acquiring linguistic competence tightly correlated with other cognitive achievements. But it appears that it does not; both the rich and the poor, the high IQed and the low, those with good memories and bad all seem to acquire linguistic competence at roughly the same rate and roughly the same way.

The second reason for thinking that UG involves at least one special cognitive feature is that it would be quite surprising biologically if it did not.  Many animals have (almost) unique capacities (think echo location in bats, or navigation in ants).  These capacities supervene on distinctive cognitive powers (e.g. in ants, the built in capacity to form a compass oriented map using sun position as anchor). Why should it be any different with humans and language? We are not surprised to find that other animals are specifically built to do the special things they do, why should humans be any different? 

Third, the alternative, that all learning relies on general purpose mechanisms, is as coherent as the idea that all perception relies on a general purpose sensing mechanism.  Just as seeing involves mental mechanisms and cognitive apparatus different from hearing or smelling or touching or tasting so too learning language is different from learning faces or learning to recognize objects or fixing causal interactions.  Gallistel and King (in Memory and the Computational Brain: 221) make the point well:
…a very general truth about learning mechanisms [is] they do not learn universal truths.  The relevant universal truths are built into the structure of a learning mechanism. Indeed, absent some built-in relevant universal truths and the strong constraints they place on the form of the representation that can be extracted from a given experience, learning would not be possible.
As linguistic representations from all that we currently know are formally quite different from other cognitive objects it would be surprising if their peculiar properties did not require some built in language specific mental mechanisms to allow for their acquisition and use, just as in the case of honey-bees and ants with respect to navigation.

Many resist these conclusions about UG. The idea that UG involves at least one linguistic specific feature is considered particularly controversial. But for the general reasons noted above, I can’t see why anyone would assume anything different. This judgment is reinforced once one takes a look at the linguistic competence humans have in more detail. Linguistic objects are very distinctive. UG must be able to accommodate these distinctive properties. If this means that UG invokes special cognitive powers, we should not be in the least surprised, nor perturbed.  That’s the way biology works.

Friday, September 28, 2012

Why this blog?

-->


This blog is the direct result of an article by Tom Bartlett in the May 12, 2012 issue of the Chronicle of Higher Education. The article reports on a “debate” pitting Chomsky (“the discipline’s long-reigning king”) against Dan Everett (“the former missionary” and “true-blooded Chomskyan” whose belief in God and Chomsky “had melted away”). Everett’s claim is that Pirahã (an indigenous language spoken in Brazil) fails to display recursion and that this conclusively demonstrates that Chomsky’s conception of Universal Grammar (in which recursion is the defining property) is wrong.  Despite the fevered prose (Chomsky coverage is almost always breathless) it was pretty clear to me that given Everett’s reported views there could be no “debate” for the simple reason that Everett’s apparent understanding of ‘Universal Grammar’ had nothing to do with Chomsky’s (I contributed some comments on the website of the article to this effect under ‘nhornste’).  The “debate” was based on a misunderstanding and so a simple equivocation.  The article was apparently widely read and so a success for the Chronicle,  (a Chomsky take-down always makes for “good press”) but it had virtually no substance.

It did, however, have a consequence. The “debate” led me to appreciate how little linguistic outsiders (and even practitioners) know about the foundations and results of the Generative Enterprise initiated by Chomsky in the mid 1950s.  This blog is an attempt to rectify this. It will partly be a labor of hate; aimed squarely at the myriad distortions and misunderstandings about the generative enterprise initiated by Chomsky in the mid 1950s.  There is a common view, expressed in the Chronicle article, that Chomsky’s basic views about the nature of Universal Grammar are hard to pin down and that he is evasive (and maybe slightly dishonest) when asked to specify what he means by Universal Grammar (henceforth I’ll stick to the shorter ‘UG’ for ‘Universal Grammar’).  This is poodle poop!

The basic idea is simple and has not changed: Just as fish are built to swim and birds to fly humans are build to talk. Call the faculty responsible for this ability ‘the Faculty of Language,’ (FL for short).  The aim of the generative enterprise is to describe the fine structure of FL. The name we give to the proposed structure is ‘Universal Grammar’; ‘universal’ because it is intended to describe the capacity that all humans have and ‘grammar’ because grammars are compact ways of describing the words, morphemes, phrases, and sentences of a language. Over the years Chomsky and colleagues have made various proposals concerning the structure of UG.  It is not a daring hypothesis to propose that natural language grammars are recursive (it follows from the easily observed fact that there is no real upper bound on the size of a sentence) and so UG must allow for recursive grammars.  The interesting question is not whether there is recursion but the specific nature of the recursion that natural language grammars have. Studying the properties of natural language grammars should, we hope, shed light on how UG is constructed. So what’s UG?  It is the general recipe in FL that humans have to build grammars of natural languages.  What features does it have? Well, that is, as they say, an empirical question which this blog will discuss.  But reader be warned: as the central object of study within Chomskyan linguistics is the structure of UG and as the field is very active the details of the description change, or at least may appear to change to the untutored eye. I personally think that many of the central findings are pretty secure and that later theories have been conservative in that they have preserved the findings of earlier theories.  We intend to discuss some of this in the future.

The perspicuous reader will have noted the ‘we’ in the last sentence. I am one of many that will be writing here.  David Pesetsky is a co-conspirator. I have asked several others to contribute as well.  The contributors disagree on many issues.  However we all believe that there are no major empirical discoveries that have invalidated the Generative approach in linguistics initiated by Chomsky, no serious methodological failings concerning the practice of linguists and no conceptual incoherence in the leading assumptions upon which this practice is founded.  The details are all up for grabs. The basic perspective has more than proven its worth. 

Before ending some may be wondering about the equivocation that vitiated the “debate” in the Chronicle.  Well it’s this: Chomsky’s claim is that the distinctive characteristic of UG is that it contains recursion.  This is the defining property of FL, which, recall, is the human capacity to acquire language.  This does not imply that every human language grammar deploys recursion. It does imply that every human can learn a grammar that is recursive. The Pirahã may not deploy recursion when speaking Pirahã (though I should add here that Everett’s claim is likely false (c.f. Pirahã Exceptionality: A Reassessment in Language 2009:355-204) but Pirahã children have no trouble learning Brazilian Portuguese (an undisputedly recursive language) and so there is no evidence that their UGs are any different from anyone else’s.  Everett (and the Chronicle) interpreted UG to mean that every language must have recursive structures, while what Chomsky means is that recursion is a property of FL.  Whether Pirahã has recursion or not (and to repeat, it looks like it does) has no bearing on whether Pirahã speakers’ UGs have it or not.  This was the equivocation and this is why the “debate” was pointless.