Comments

Showing posts with label Levinson. Show all posts
Showing posts with label Levinson. Show all posts

Thursday, October 5, 2017

Pity the non-Chomskyan!

Pity the non Chomskyans! They don’t value their work except in opposition to what Chomsky does (or doesn’t). The only glory they prize is reflected, and they will go to great lengths to sun themselves in it, even to the point of (knowingly?) distorting the mirror in which they reflect themselves. Imagine the irony that someone like me perceives. I am constantly remonstrated with for not sufficiently valuing non GG work and then when I look at some I find that the practitioners themselves only prize their research to the degree that it overturns some GG nostrum and thereby “revolutionizes” the study of language (never a brick in the wall for them, always a complete overturning of the basics). It would appear that for these investigators Chomsky has indeed defined the limits of the interesting in the study of language (a view I have some sympathy with, I would add) and that anything that does not directly address a point that he has (allegedly) made is of little value. Indeed, compared to them, my insistence that one can study language with interests orthogonal to GG’s must seem disingenuous. To non-Chomskyans, Chomsky is everywhere and always and their research is nowhere and never unless it confronts his.

A recent addition to this literature of self-loathing is making quite a splash (here, here, here, here). Part of the splash can be traced to the PR-Academy complex that I mentioned in a previous post (here). Some of the co-authors have Max Planck affiliation and so the powerful PR Wurlitzer has been fully cranked up to spread the word.

However, part of the splash is due to the claim the paper makes that Chomsky is, once again, wrong, more specifically that culture rather than biology is what drives language structure. Of course, as you all know, this is one fork in the intellectual road that any sane person should immediately take. Is it culture or biology? Well, depending on the linguistic feature of interest it could be either, neither or both.  Language, we all know, is a complex thing and the confluence of many different interacting causal forces. Everybody knows this, so it is not news (though it is often intoned as if it were a great discovery, like people noting, sagely, that any given cognitive capacity is the combination of learned and innate factors (duh!)). What is news is finding out which factor predominates for any given property of interest and how it does so. But finding this out in any given case will not (and I can guarantee this) discredit the causal efficacy of other factors in other cases. And, moreover this can hardly be news. So, this is something that GG has acknowledged for a very very very …very long time.  Even if you think, as I do, that biology (widely construed) plays the lead role in restricting the class of Generative Procedures (GP) available to human Gs, you need not think that culture (widely construed) plays no role in determining what a given G looks like. For example, why I have the G that I have is not exclusively due to my having a human FL/UG. I have an Englishy G because I grew up in an English speaking environment, was culturally exposed to Howdy Doody, Captain Kangaroo, and Rocky and Bullwinkle, and read Bill Shakespeare in high school. Many of my G’s idiosyncrasies are similarly cultural (e.g. I am a proud speaker of a dialect in which Topicalization (aka Yiddish movement) runs rampant). But I very much doubt the fact that my Topicalization forming G displays ECP effects has much to do with Rocky, Shakespeare, Bullwinkle or Sholem Aleichem. Here I look to biology to explain why my G obeys the ECP (and for the familiar Poverty of Stimulus reasons which I could go on about for hours (and have)). So biology AND culture, with each playing a more prominent role depending in the phenomenon of interest.

Curiously, this most obvious position is tacitly denied by non-Chomsyans. They act as if Chomskyans must think that anything languagy must reflect innate features of the mind/brain and so that if anything is shown to not be such, then this shows that Chomsky was wrong. And their obsession with showing that Chomsky is wrong suggests that they believe that unless he is, then what they have shown about, say, the influence of cultural mechanisms on some languagy fact is inherently BORING, without any possible intrinsic interest. This, atleast, would neatly explain why non-Chomskyans consistently assume that Chomsky’s position consists in the absurd claim that anything involving language in any way must be innate.

You probably think that I am exaggerating here. But I am not, really. Here is the authors’ summary of the Dunn, Greenhill, Levinson, and Gray (DGLG) paper published in Nature:

Languages vary widely but not without limit. The central goal of linguistics is to describe the diversity of human languages and explain the constraints on that diversity. Generative linguists following Chomsky have claimed that linguistic diversity must be constrained by innate parameters that are set as a child learns a language (1, 2). In contrast, other linguists following Greenberg have claimed that there are statistical tendencies for co-occurrence of traits reflecting universal systems biases (3, 4, 5), rather than absolute constraints or parametric variation. Here we use computational phylogenetic methods to address the nature of constraints on linguistic diversity in an evolutionary framework (6). First, contrary to the generative account of parameter setting, we show that the evolution of only a few word-order features of languages are strongly correlated. Second, contrary to the Greenbergian generalizations, we show that most observed functional dependencies between traits are lineage-specific rather than universal tendencies. These findings support the view that—at least with respect to word order—cultural evolution is the primary factor that determines linguistic structure, with the current state of a linguistic system shaping and constraining future states.

Let’s engage in some initial parsing. The paper aims to show that language change (in particular word order changes in diachronically related languages) is path dependent, with different dependencies changing at different rates across different groupings of languages. DGLG concludes from this that the transitions between the languages is not driven by innate features of FL/UG, nor does it reflect systematic universal probabilistic biases. And they conclude form this that Chomsky and Greenberg must be wrong. I am not qualified to discuss Greenberg’s positions in any detail, but I would like to cast a very skeptical eye towards the claims made for Chomsky’s parameter views.

Let’s read the above précis a little more carefully.

First, DGLG focuses on “languages” and the diachronic changes between them. To be  GG/Chomsky relevant, we need to unpack this and relate it to grammars. With this translation we get the following:

Grammars (G) vary widely but not without limit. The central goal of linguistics is to describe the diversity of human Gs and explain the constraints on that diversity. Generative linguists following Chomsky have claimed that linguistic diversity must be constrained by innate parameters that are set as a child learns a language…

Second, I disagree with even this reworked version of DGLG’s claims about the aim of linguistics, at least as GG and Chomsky understand it. The ultimate aim is to describe the structure of FL/UG. A way station towards this end is to understand the structure of human Gs used by speakers of different languages. Hence, describing these (their commonalities and differences) is a useful proximate goal towards the ultimate end. Linguists have traced some of the differences between human languages to differences in the Gs that native speakers use. This implies diverse Gs and this further implies that FL must be capable of acquiring Gs at least as diverse as these (and maybe yet more diverse given that it is unclear that 7,000 (the purported number of languages out there) marks the limit of G diversity). So yes, describing the variety of human languages to the degree that it enables us to describe the variety of human Gs is a useful step in exploring the structure of FL/UG but the ultimate aim of linguistics is to understand the structure of FL/UG not to describe the diversity of human Gs, or “the diversity of human languages and the constraints on that diversity”.[1]

Third, the phrase “constraints on that diversity” is ambiguous. One reading is anodyne and correct. One aim of GG has been to describe principles of grammar invariant across Gs, the idea being that these will reveal the design features of FL/UG.[2] This does not imply that the language products of these Gs will manifest invariant patterns. Missing this is to again confuse the difference between Chomsky and Greenberg Universals. Two Gs may embody the very same principle and yet the products of those Gs might differ greatly. Thus, for example, Rizzi’s early proposal concerning the Fixed Subject Effect is that the ECP (which, let us say underlies the effect) holds in both Italian and English but the two Gs derive subject A’-movement in different ways so that only English perceptibly falls under the scope of the ECP.[3] In other words, Italian manages to derivationally escape the purview of the G invariant ECP and hence does not show Fixed Subject Effects. Note, crucially, Italian does embody the ECP but it does not show Fixed Subject Effects. The “languages” (English vs Italian) differ, the invariant principle (the ECP) is the same.  So the move from invariant principles to invariant language effects is not one that any GGer can or should blithely license. In sum, if you mean that a goal of Chomskyan linguistics is to describe the properties of Gs that arise as products of the design features of FL/UG then you would be correct. But this still leaves distance between this and invariant properties of languages.[4]

So rightly understood, describing G invariants is a proximate goal of GG inquiry. But this must be distinguished from a different project: explaining the limits of diversity. It is entirely possible that Gs have invariant properties without it being the case that there is a limit on G diversity. Let me explain. One of the innovations of LGB is its Principles and Parameters (P&P) architecture. The idea is that FL/UG specifies not only the invariant properties of Gs but all of the ways that Gs could possibly differ. These differences are coded as a finite number of two valued parameters with a given G being a (vector) specification of these specific values. As P&P was understood to have a finite number of parameters, say N, and as they could only bear one of two values this meant that there were at most 2N distinct Gs that FL/UG permitted. On this LGB/P&P conception there is a reasonable sense in which FL/UG could explain the limits of G diversity.  So, DGLG are correct in thinking that some version of GG, the LGB/P&P theory aimed to place strong limits on G diversity.

However, this theory did not address how parameters were set or how parameters changed over time as Gs changed. P&P theories must be supplemented with theories of learning/acquisition to provide a theory of language change. Or, even if you think that there are only a finite number of Gs because any G is simply a list of finite parameter values, you still need a theory of how parameters are fixed to explain how parameters change over time. Now, one theory of G change would be that it tracks some intrinsic structure/fault lines of the parameter space (e.g. parameter 1 links to 2 and to 4 so that if you change the value of 1 from A to a then you need to change the values of 2 from B to b and 4 from C to c). This is one possible theory. Call it an endogenous theory of parameter setting (EnPS). EnPS accounts would provide FL/UG internal paths along which G change would occur and would provide a very strong implicit theory on the dynamics of G diversity. It would not only explain what the range of possibility is, but would also specify the possible range of changes between Gs that is available. Note that this kind of view need not endorse the position that all G change is canalized by FL/UG. It is possible that some changes are and some are not. But the strongest view would aim to predict the dynamics of G change entirely from the endogenous structure of FL/UG.

To my knowledge nobody has ever proposed such a view. In fact, to my knowledge, such a view has been understood to be very problematic, the reason being that the degree to which the parameters are mutually dependent to that degree the problem of incremental parameter setting increases. In fact, were all the parameters to speak to one another (i.e. the value of any parameter being conditional on the value of every parameter) the problem of incremental parameter setting becomes effectively impossible.

Dresher and Kaye discussed this first a while ago as regards stress systems, and Fodor and Sakas have explored this in detail as regards syntactic parameters. The solution has been to try and identify linguistic “triggers,” types of data that relied exclusively on the value of a single parameter. Triggers, in other words, are ways of trying to finesse the intractability of incremental parameter setting without denying that parameters are inter-twined. The idea is that their intermingling need not appear everywhere in the PLD and that all that setting requires is that there be some PLD data that unambiguously reveals what value a given parameter has. In other words, the idea is that in some domains the parameters function as if independent of one another (do not interact) and this relieves the computational problem that intertwining presents.

Why do I mention this? Because, the bulk of work on parameters has not been in trying to limn their interdependencies, but to isolate them and render them relatively independent so that incremental parameter setting be possible. In other words, so far as I know, there has precious little work or commitment to a EnPS kind of theory within Chomskyan GG. Moreover, and this is the important bit, a Chomskyan theory of GG does not need such an account. In other words, if it is true, then it is very interesting, but there is nothing as regards the Chomskyan project that requires that something like this be available. It is entirely consistent with that project that there be no explicit or implicit dynamics coded in the parameter space. So asserting that the absence of such a theory of parameters challenges the Chomsky conception of FL/UG (which is what DGLG does) is just plain wrong. Or, to put this another way: claiming that there is a richly structured FL/UG is compatible with the claim that FL/UG does not determine how populations of speakers  move from one G in the space of possible Gs to another.[5]

We can in fact go further. As those who have read some linguistics since LGB know, the idea that FL/UG contains a finite list of parameters that delimit the range of possible Gs has been under debate. We have even talked about this on FoL (e.g. here, here, here, here). What’s important here is that the parameters part of P&P is not an intrinsic part of the Chomskyan problematic. It might be true, but then it might not be. There are theories like GB that endorse a P&P architecture and there are accounts like that in Aspects or LSLT (and, from the way I read it, current MP accounts) that do not. If FL/UG has a list of specified parameters, that would be an amazing and remarkable discovery. But the evidence is not overwhelming that this is the case (so far as I can tell as a non expert in these matters) and if it is not the case it does not mean that Chomsky is wrong in claiming that we have an innately specified FL/UG that limits the properties of human Gs. All it means is that there are some features of Gs about which FL/UG is mute. Happily, that would leave something for non syntacticians to do (e.g. provide theories of learning that would address how we go from PLD to Gs given FL/UG (e.g. as Yang and Lidz do, for example)).

In sum, DGLG’s claims about the implications of their results for the Chomsky enterprise are flatly wrong, and in two ways. First, nobody has proposed the kind of theory that DGLG’s data is meant to refute and second the Chomskyan conception does need a theory of the kind that DGLG’s data is meant to refute. So, whether or not DGLG is correct is at right angles to Chomsky’s central claims. And, what is more, I suspect that DGLG knows this (or at least should have). Let me say a word about this.

DGLG cites LGB and Baker’s book as the source of the ideas that the paper argues to be incorrect. However, DGLG cites no specific passages or pages for this claim. Why not? When Chomsky goes after someone critically he does so chapter and verse. He quotes exactly what his protagonist says before arguing that it is incorrect. This is not what Chomsky’s critics generally do. Rather they assert in very broad brushstrokes what Chomsky’s views are and then go on to state that they are inadequate.[6] The problem is that what they criticize is often not his views. The fact that this happens so often leads me to think that this is not accidental. Either critics do not care what his views are (they only care to discredit them so as to discredit him) or they are too lazy to do serious criticism. I am not sure which is worse, but they are both serious intellectual failings.

What I did not realize until recently is that Chomsky’s critics might well be motivated less by malice and sloth than by a deep intellectual insecurity. Many of Chomsky’s critics are upset by the possibility (fact?) that he does (might?) not care about what they are doing. What motivates some critics, then, is the suspicion (fear?) that what they are doing is of little value. Too assuage this desperation, they orient their conclusions as rebuttals of Chomsky’s putative views. Why? Because they are sure that what Chomsky does is interesting and so they reassure themselves that their work has value by arguing that it shows that his views are wrong. The implication is that if this were not the case (and very often it is in fact not the case as the empirical conclusions are generally irrelevant to Chomsky’s claims (Everett is the poster child for this)) then their own work was boring and of dubious interest.[7] And I thought that I was in thrall of Chomsky![8]

Let me end with one more point regarding the DGLG paper and then point you to a very good review.

DGLG focuses entirely on word order. What DGLG means by “linguistic structure” is word order properties of utterances/sentences. It says this very explicitly. But if this is the focus then DGLG must know that it will be of dubious relevance to Chomsky’s central claims which, as everyone knows by now, considers word order effects to be at best second order (and maybe even less relevant) as regards the central features of FL/UG. Word order effects are, of course, centrally relevant to Greenberg’s conceptions and there are GGers who have concentrated on this (e.g. Kayne) and we have even covered some of this in FoL (see Culbertson and Adger here). But as regards Chomsky’s views, word order effects are decidedly secondary. Indeed, from what I know of Chomsky’s views, he might agree that word order effects are entirely “cultural” in the sense of driven by the properties of the child’s ambient linguistic environment.[9] So far as I can tell, nothing Chomsky has said in recent years (or before) would be inconsistent with this. So the fact that DGLG knowingly focuses on the kinds of effects that the authors (or at least some of the authors (I am looking at you Mr Levinson)) know are at right angles to Chomsky’s concerns further buttresses my conclusion that DGLG thinks its results worthless unless they directly gainsay Chomsky’s views. Sadly, if this is DGLG’s position, then worthless it is. The paper even if completely correct scarcely bears on Chomsky’s central claims. Fortunately, the DGLG conclusions are worth thinking about, IMO, even if they do not bear on Chomsky’s views at all. Let me turn to this briefly.

I have spent a lot of time pooping on the DGLG paper’s claim that it overturns some central Chomskyan dogma. However, contrary to the authors, I am not as sure as they seem to be that its results are uninteresting unless they bear on Chomsyan/GG concerns. My tastes then are more catholic than DGLG’s. I believe that finding that G change is path dependent is potentially very interesting, even if not all that new. It is not a new idea as it is already embodied in the position that language contact can affect how Gs change (how Gs change (and maybe even their rates of change) is likely a function of the specific properties of Gs in contact. If so, change will be path dependent. Indeed, from what I know this is the standard view, which is why Bickerton’s contrary claim is so contentious.

Mark Liberman has an excellent discussion of the DGLG paper (here) that touches on this point as well as many others way above my pay grade (damn I wish I knew more stats). There is also some interesting response by Greenhill in the comment section of Mark’s post, though I personally think that he fails to engage the main point concerning the large number of ignored dimensions and the kind of structure they might contain. As Mark observes, there is lots of room in these ignored dimensions for an EnPS story should one care to make one. 

I also think that Liberman’s last point touches on something critical. As he points out, at least as regards GGers who work on G change (like Tony Kroch or David Lightfoot or Ian Roberts): “…features like “OBV” (the code for whether objects follow verbs) should be seen as superficial grammatical symptoms rather than atomic grammatical traits” (3). This points to a larger problem of the relevance of DGLG to GG research into diachrony; DGLG takes the project to be language change rather than G change (as I noted above, these are not the same thing). Greenhill responds to this that these categories are not his, but those that other people have identified, DGLG just aiming to test them. The problem Mark is pointing to is that they are the wrong things to test, at least if one’s interests lie with the structure of Gs.  Going from overt language to G rules/parameters is not straightforward (see Dresher and Kaye, and Fodor and Sakas). What is relevant to speakers qua Language Acquisition Devices is the features of the Gs, so abstracting from this is, as Mark observes, a problem.

Ok, should you read the DGLG paper? Sure. It is very short and potentially interesting (though, IMO, inevitably overhyped) and, as Mark Liberman notes, the product of a lot of hard work. But, it is also deeply misleading and, IMO, borders (well, IMO, crosses the border to) dishonest. The source of the dishonesty is likely overdetermined. I mentioned malice and sloth. But I suspect that intellectual insecurity is really what is driving the anti GG, anti Chomsky slant. Anti Chomskyans do not have the courage of their stated interests. So, when you are done reading DGLG, spend a second mourning the sad plight of the non Chomskyan. Only by being anti can they be at all. Sad really. And I would be greatly sympathetic regarding this insecurity were they not sullying the intellectual landscape in trying to convince themselves of the value of their research.



[1] An analogy: the aim of astronomy is not to describe the motion of the planets but to describe the forces that have the planets move as they do. Of course an excellent first step towards the latter is a decent description of the former. But the two goals are not identical. Moreover, as Newton discovered, describing the motion of non-planets here on earth is also a useful step in exposing the forces at work our there in the heavens.
[2] Though, IMO, looking for common detectable features across individual Gs is not as useful as generally supposed.
[3] Perceptible to the linguist that is, not the LAD as the evidence will be of the negative variety; fixed subject violations lead to unacceptability.
[4] My reading of Greenberg is that his project was to identify invariant properties of languages. From what I’ve seen (which is limited) my impression is that typologists are pretty skeptical that many of these exist. If this is right, then it seems as regards “surfacy” language properties, invariants are pretty hard to come by. This would be no surprise to a Chomskyan GGer.
[5] Which is not to say that this might not be a very interesting question to investigate, and GGers did so. See Berwick and Niyogi’s early work on this in a kind of Neo-Darwinian setting.
[6]  I want to emphasize the broad brushstroke nature here. Like any interesting view, Chomsky’s has several interacting sub-parts and is based on several assumptions. It is entirely possible (probable?) that he may be right in some ways and wrong in others (as is true of most everyone’s views). The goal of a critic is to isolate how someone is wrong, and then means getting into details. What specific assumption is false? What particular inference would we like to challenge? Quoting passage and verse forces one to zero in on the problem. Citing LGB as the source does not. So, not only is this lazy, but it allows a critic to let him/herself off the hook and allows him/her to avoid explaining in detail how someone is wrong. As criticism is valuable in allowing us to isolate troublesome assumptions, this kind of lazy citation promotes obfuscation. So what part of Chomsky’s assumptions are wrong? Well you know the LGB part. Really? Common.
[7] I personally would not conclude this. But it provides a reasonable motive for the otherwise inexplicable regular incapacity of Chomsky’s critics to get his views right.
[8] It is interesting that Chomsky does not feel equally threatened by work different from his own. He just gets on with it. Yes, he responds to critics. But he what he mainly does is define the project and get on with it. It would be nice if this were the norm.
[9] To be honest, I do not know what ‘cultural’ means. I am assuming it means to contrast with biological, memes and all that. In effect, fixed via something like learning rather than fixed by biological make-up.

Monday, May 20, 2013

Evans-Levinson: the sound and the fury


I confess that I did not read the Evans and Levinson article (EL) (here) when it first came out. Indeed, I didn’t read it until last week.  As you might guess, I was not particularly impressed. However, not necessarily for the reason you might think. What struck me most is the crudity of the arguments aimed at the Generative Program, something that the (reasonable) commentators (e.g. Baker, Freidin, Pinker and Jackendoff, Harbour, Nevins, Pesetsky, Rizzi, Smolensky and Dupoux a.o.) zeroed in on pretty quickly. The crudity is a reflection, I believe, of a deep seated empiricism, one that is wedded to a rather superficial understanding of what constitutes a possible “universal.” Let me elaborate.

EL adumbrates several conceptions of universal, all of which the paper intends to discredit. EL distinguishes substantive universals from structural universals and subdivides the latter into Chomsky vs Greenberg formal universals. The paper’s mode of argument is to provide evidence against a variety of claims to universality by citing data from a wide variety of languages, data that EL appears to believe, demonstrate the obvious inadequacy of contemporary proposals. I have no expertise in typology, nor am I philologically adept. However, I am pretty sure that most of what EL discuss cannot, as it stands, broach many of the central claims made by Generative Grammarians of the Chomskyan stripe. To make this case, I will have to back up a bit and then talk on far too long. Sorry, but another long post. Forewarned, let’s begin by asking a question.

What are Generative Universals (GUs) about?  They are intended to be in the first instance, descriptions of the properties of the Faculty of Language (FL). FL names whatever it is that humans have as biological endowment that allows for the obvious human facility for language. It is reasonable to assume that FL is both species and domain specific. The species specificity arises from the trivial observations that nothing does language like humans do (you know: fish swim, birds fly, humans speak!). The domain specificity is a natural conclusion from the fact that this facility arises in all humans pretty much in the same way independent of other cognitive attributes (i.e. both the musical and the tone deaf, both the hearing impaired and sharp eared, both the mathematically talented and the innumerate develop language in essentially the same way).  A natural conclusion from this is that humans have some special features that other animals don’t as regards language and that human brains have language specific “circuits” on which this talent rests. Note, this is a weak claim: there is something different about human minds/brains on which linguistic capacity supervenes. This can be true even if lots and lots of our linguistic facility exploits the very same capacities that underlie other forms of cognition. 

So there is something special about human minds/brains as regards language and Universals are intended to be descriptions of the powers that underlie this facility; both the powers of FL that are part of general cognition and those unique to linguistic competence.  Generativists have proposed elaborating the fine structure of this truism by investigating the features of various natural languages and, by considering their properties, adumbrating the structure of the proposed powers. How has this been done? Here again are several trivial observations with interesting consequences.


First, individual languages have systematic properties. It is never the case that, within a given language, anything goes.  In other words, languages are rule governed. We call the rules that govern the patterns within a language a grammar. For generativists, these grammars, their properties, are the windows into the structure of FL/UG. The hunch is that by studying the properties of individual grammars, we can learn about that faculty that manufactures grammars.  Thus, for a generativist, the grammar is the relevant unit of linguistic analysis. This is important. For grammars are NOT surface patterns. The observables linguists have tended to truck in relate to patterns in the data. But these are but way stations to the data of interest: the grammars that generate these patterns.  To talk about FL/UG one needs to study Gs.  But Gs are themselves inferred from the linguistic patterns that Gs generate, which are themselves inferred from the natural or solicited bits of linguistic productions that linguists bug their friends and collaborators to cough up. So, to investigate FL/UG you need Gs and Gs should not be confused with their products/outputs, only some of which are actually perceived (or perceivable).

Second, as any child can learn any natural language, we are entitled to conclude from the intricacies of any given language to powers of FL/UG capable of dealing with such intricacies.  In other words, the fact that a given language does NOT express property P does not entail that FL/UG is not sensitive to P. Why? Because a description of FL/UG is not an account of any given language/G but an account of linguistic capacity in general.  This is why one can learn about the FL/UG of an English speaker by investigating the grammar of a Japanese speaker and the FL/UG of both by investigating the grammar of a Hungarian, or Swahili, or Slave speaker. Variation among different grammars is perfectly compatible with invariance in FL/UG, as was recognized from the earliest days of Generative Grammar. Indeed, this was the initial puzzle: find the invariance behind the superficial difference!

Third, given that some languages display the signature properties of recursive rule systems (systems that can take their outputs as inputs), it must be the case that FL/UG is capable of concocting grammars that have this property. Thus, whatever G an individual actually has, that individual’s FL/UG is capable of producing a recursive G. Why, because that individual could have acquired a recursive G even if that individual’s actual G does not display the signature properties of recursion. What are these signature properties?  The usual: unboundedly large and deep grammatical structures (i.e. sentences of unbounded size). If a given language appears to have no upper bound on the size of its sentences, then it's a sure bet that the G that generates the structures of that language is recursive in the sense of allowing structures of type A as parts of structures of type A. This, in general will suffice to generate unboundedly big and deep structures. Examples for this type of recursion include conjunction, conditionals, embedding of clauses as complements of propositional attitude verbs, relative clauses etc.  The reason that linguists have studied these kinds of configurations is precisely because they are products of grammars with this interesting property, a property that seems unique to the products of FL/UG, and hence capable of potentially telling us a lot about the characteristics of FL/UG.

Before proceeding, it is worth noting that the absence of these noted signature properties in a given language L does not imply that a grammar of L is not basically recursive.  Sadly, FL seems to leap to this conclusion (443). Imagine that for some reason a given G puts a bound of 2 levels of embedding on any structure in L. Say it does this by placing a filter (perhaps a morphological one) on more complex constructions. Question: what is the correct description of the grammar of L?  Well, one answer is that it does not involve recursive rules for, after all, it does not allow unbounded embedding (by supposition).  However, another perfectly possible answer is that it allows exactly the same kinds of embedding that English does modulo this language specific filter.  In that case the grammar will look largely like the ones that we find in languages like English that allow unbounded embedding, but with the additional filter. There is no reason just from observing that unbounded embedding is forbidden to conclude that the grammar of this hypothetical language L (aka Kayardild or Piraha) has a grammar different in kind from the grammars we attribute to English, French, Hungarian, Japanese etc. speakers.  In fact, there is reason to think that the Gs that speakers of this hypothetical language have does in fact look just like English etc.  The reason is that FL/UG is built to construct these kinds of grammars and so would find it natural to do so here as well.  Of course L would seem to have an added (arbitrary) filter on the embedding structures, but otherwise the G would look the same as the G of more familiar languages. 

An analogy might help.  I’ve rented cars that have governors on the accelerators that cap speed at 65 mph.  The same car without the governor can go far above 90 mph. Question: do the two cars have the same engine?  You might answer “no” because of the significant difference in upper limit speeds. Of course, in this case, we know that the answer is “yes”: the two cars work in virtually identical ways, have the very same structures but for the governor that prevents the full velocity potential of the rented car from being expressed.  So, the conclusion that the two cars have fundamentally different engines would be clearly incorrect.  Ok: swap Gs for engines and my point is made.  Let me repeat it: the point is not that the Gs/engines might be different in kind, the point is that simple observation of the differences does not license the conclusion that they are (viz. you are not licensed to conclude that they are just finite state devices because they don’t display the signature features of unbounded recursion, as EL seems to).  And, given what we know about Gs and engines the burden of proof is on those that conclude from such surface differences to deep structural differences.  The argument to the contrary can be made, but simple observations about surface properties just doesn’t cut it.

Fourth, there are at least two ways to sneak up on properties of UGs: (i) collect a bunch and see what they have in common (what features do all the Gs display) and (ii) study one or two Gs in great detail and see if their properties could be acquired from input data. If any could not be, then these are excellent candidate basic features of FL/UG. The latter, of course, is the province of the POS argument.  Now, note that as a matter of logic the fact that some G fails to have some property P can in principle falsify a claim like (i) but not one like (ii).  Why? Because (i) is the claim that every G has P, while (ii) is the claim that if G has P then P is the consequence of G being the product of FL/UG. Absence of P is a problem for claims like (i) but, as a matter of logic, not for claims like (ii) (recall, If P then Q is true if P is false).  Unfortunately, EL seems drawn to the conclusion that PàQ is falsified if –P is true. This is an inference that other papers (e.g. Everett’s Piraha work) are also attracted to. However, it is a non-sequitur. 

EL recognizes that arguing from the absence of some property P to the absence of Pish features in UG does not hold.  But the paper clearly wants to reach this conclusion nonetheless. Rather than denying the logic, EL asserts that “the argument from capacity is weak” (EL’s emphasis). Why? Because EL really wants all universals to be of the (i) variety, at least if they are “core” features of FL/UG. As these type (i) universals must show up in every G if they are indeed universal, absence to appear in one grammar is sufficient to call into question its universality. EL is clearly miffed that Generativists in general and Chomsky in particular would hold a nuanced position like (ii). EL seems to think that this is cheating in some way.  Why might they hold this? Here’s what I think.

As I discussed extensively in another place (here), everyone who studies human linguistic facility appreciates that competent speakers of a language know more than they have been exposed to.  Speakers are exposed to bits of language and from this acquire rules that generalize to novel exemplars of that language.  No sane observer can dispute this.  What’s up for grabs is the nature of the process of generalization. What separates empiricists from rationalists conceptions of FL/UG is the nature of these inductive processes. Empiricists analyze the relevant induction as a species of pattern recognition. There are patterns in the data and these are generalized to all novel cases.  Rationalists appreciate that this is an option, but insist that there are other kinds of generalizations, those based on the architectural properties (Smolensky and Dupoux’s term) of the generative procedures that FL/UG allow. These procedures need not “resemble” the outputs they generate in any obvious way and so conceiving this as a species of pattern recognition is not useful (again, see here for more discussion).  Type (ii) universals fit snugly into this second type, and so empiricists won’t like them.  My own hunch is that an empiricist affinity for generalizations based on patterns in the data lies behind EL’s dissatisfaction with “capacity” arguments; they are not the sorts of properties that inspection of cases will make manifest. In other words, the dissatisfaction is generated by Empiricist sympathies and/or convictions which, from where I sit, have no defensible basis. As such, they can be and should be discounted. And in a rational world they would be. Alas…

Before ending, let me note that I have been far too generous to the EL paper in one respect.  I said at the outset that its arguments are crude. How so?  Well, I have framed the paper’s main point as a question about the nature of Gs. However, most of the discussion is framed not in terms of the properties of Gs they survey but in terms of surface forms that Gs might generate.  Their discussion of constituency provides a nice example (441).  They note that some languages display free word order and conclude from this that they lack constituents.  However, surface word order facts cannot possibly provide evidence for this kind of conclusion, it can only tell us about surface forms. It is consistent with this that elements that are no longer constituents on the surface were constituents earlier on and were then separated, or will become constituents later on, say on the mapping to logical form.  Indeed, in one sense of the term constituent, EL insists that discontinuous expressions are such for they form units of interpretation and agreement. The mere fact that elements are discontinuous on the surface tells us nothing about whether they form constituents at other levels. I would not mention this were it not the classical position within Generative Grammar for the last 60 years. Surface syntax is not the arbiter of constituency, at least if one has a theory of levels, as virtually every theory that sees grammars as rules that relate meaning with sounds assumes (EL assumes this too).  There is nary a grammatical structure in EL and this is what I meant be my being overgenerous. The discussion above is couched in terms of Gs and their features. In contrast, most of the examples in EL are not about Gs at all, but about word strings. However, as noted at the outset, the data relevant to FL/UG are Gs and the absence of Gish examples in EL makes most of EL’s cited data irrelevant to Generative conceptions of FL/UG.

Again, I suspect that the swapping of string data for G data simply betrays a deep empiricism, one that sees grammars as regularities over strings (string patterns) and FL/UG as higher order regularities over Gs. Patterns within patterns within patterns. Generativists have long given up on this myopic view of what can be in FL/UG.  EL does not take the Generative Program on its own terms and show that it fails. It outlines a program that Generativists don’t adopt and then shows that it fails by standards it has always rejected using data that is nugatory.

I end here: there are many other criticisms worth making about the details, and many of the commentators of the EL piece better placed than me to make them do so. However, to my mind, the real difficulty with EL is not at the level of detail. EL’s main point as regards FL/UG is not wrong, it is simply besides the point.  A lot of sound and fury signifying nothing.