Comments

Showing posts with label Berwick. Show all posts
Showing posts with label Berwick. Show all posts

Tuesday, May 27, 2014

A game I play

Every now and then I play this game: how would Chomsky respond?  I do this for a variety of reasons. First, I respect his smarts and I think it is interesting to consider how things would look from his point of view. Second, I ache found that trying to understand his position, even when it appears foreign to my way of thinking has been useful for me in clarifying my own ideas. And third, because given Chomsky's prominence in the field and his influence on how the world views the efforts of GG, it is useful to know how he would defend a certain point of view even if he himself doesn't (or hasn't) defended it in this way.  regarding the third point: it's been my experience that when one suggests that "GG assumes "X or "GG has property Y" people take this to mean that Chomsky said that "GG assumes X" or "GG has property Y."  I am not always delighted with this way of parsing things, but given the way the world is and given that Chomsky is wickedly smart and very often correct the game is worth the effort.

In an earlier post (here), I tried to explain why I did not find any of the current attacks on the POS argument in the literature compelling. part of this consisted in explaining why I thought that the standardly cited reanalyses had :"focused on the wrong data to solve the wrong problem" and that as a result there is no reason to think that more work along these lines would ever shed any useful light on POS problems. I suggested that this was how one should really understand the discussion over Polar Questions: the anti-POS "rebuttals" misconstrue the point at issue, get the data wrong and supply answers for the wrong questions. In short, useless.

Why do I mention all of this again? Because there is an excellent recentish paper (here) by Berwick, Chomsky and Piatelli-Palmarini (BCP) that makes the points that I tried to make quickly, extensively. It is chapter 2 of the book (which is ludicrously expensive and you should take out from your library) and it suggests that my interpretation of the problem was largely on the right track. For example, I suggested that the original discussion was intended as a technically simple illustration of a much more general point aimed at a neophyte audience. BCP confirms this interpretation stating that "the examples were selected for expository reasons, deliberately simplified so that they could be presented as illustrations without the need to present more than site trivial linguistic theory" (20). They further note that the argument that Polar questions are formed using a structure dependent operation is the minimum one could say. It is not itself a detailed analysis but a general conclusion concerning the class of plausible analyses.  I also correctly surmised that the relevant data goes far beyond the simple cases generally discussed and that any adequate theory would have to extend to these more complex cases as well.  To make a long story short: I nailed it!!!

However, for those who want to read a pretty short very good discussion of the POS issue once again, a discussion where Chomsky's current views are very much in evidence, I could not do better than suggest this short readable paper "Poverty of the stimulus stands: why recent challenges fail."

One last point: there is a nice discussion here too of the interplay between PP and DP. As BCP notes, the aim of MPish accounts is to try to derive the effects of UG laden accounts that answer the POS with accounts that exploit less domain specific innate machinery. As they also note, the game is worth playing just in case you take the POS problem seriously and address the relevant data and generalizations. Changing the topic (as Perfors et al does) or ignoring the data (as Clark does and Christiansen does) means that whatever results ensue are irrelevant to the POS question at hand.  I would not have thought that this is worth repeating but for the fact that it appears to be a contentious claim. It isn't. That's why, as BCP indicates, the extant replies are worthless.

Addendum May 28/2014:

In the comments, Noah Motion has provided the following link to a very cheap version of the BCP paper. Thanks Noah.

Friday, March 21, 2014

Let's pour some oil on the flames: A tale of too simple a story

Olaf K asks in the comments section to this post why I am not impressed with ML accounts of Aux-to-C (AC) in English. Here’s the short answer: proposed “solutions” have misconstrued the problem (both the relevant data and its general shape) and so are largely irrelevant. As this judgment will no doubt seem harsh and “unhelpful” (and probably offend the sensibilities of many (I’m thinking of you GK and BB!!)) I would like to explain why I think that the work as conducted heretofore is not worth the considerable time and effort expended on it. IMO, there is nothing helpful to be said, except maybe STOP!!! Here is the longer story. Readers be warned: this is a long post. So if you want to read it, you might want to get comfortable first.[1]

It’s the best of tales and the worst of tales. What’s ‘it’? The AC story that Chomsky told to explicate the logic of the Poverty of Stimulus (POS) argument.[2] What makes it a great example is its simplicity. To be understood requires no great technical knowledge and so the AC version of the POS is accessible even to those with the barest of abilities to diagram a sentence (a skill no longer imparted in grade school with the demise of Latin).

BTW, I know this from personal experience for I have effectively used AC to illustrate to many undergrads and high school students, to family members and beer swilling companions how looking at the details of English can lead to non-obvious insights into the structure of FL. Thus, AC is a near perfect instrument for initiating curious tyros who into the mysteries of syntax.

Of course, the very simplicity of the argument has its down sides. Jerry Fodor is reputed to have said that all the grief that Chomsky has gotten from “empiricists” dedicated to overturning the POS argument has served him right. That’s what you get (and deserve) for demonstrating the logic of the POS with such a simple straightforward and easily comprehensible case. Of course, what’s a good illustration of the logic of the POS is, at most, the first, not last, word on the issue. And one might have expected professionals interested in the problem to have worked on more than the simple toy presentation. But, one would have been wrong. The toy case, perfectly suitable for illustration of the logic, seems to have completely enchanted the professionals and this is what critics have trained their powerful learning theories on. Moreover, treating this simple example as constituting the “hard” case (rather than a simple illustration), the professionals have repeatedly declared victory over the POS and have confidently concluded that (at most) “simple” learning biases are all we need to acquire Gs. In other words, the toy case that Chomsky used to illustrate the logic of the POS to the uninitiated has become the hard case whose solution would prove rationalist claims about the structure of FL intellectually groundless (if not senseless and bankrupt).

That seems to be the state of play today (as, for example, rehearsed in the comments section of this). This despite the fact that there have been repeated attempts (see here) to explicate the POS logic of the AC argument more fully. That said, let’s run the course one more time. Why? Because, surprisingly, though the AC case is the relatively simple tip of a really massive POS iceberg (c.f. Colin Phillips’ comments here March 19 at 3;47), even this toy case has NOT BEEN ADEQUATELY ADDRESSED BY ITS CRITICS! (see. In particular BPYC dhere for the inadequacies).  Let me elaborate by considering what makes the simple story simple and how we might want to round it out for professional consideration.

The AC story goes as follows. We note, first, that AC is a rule of English G. It does not hold in all Gs. Thus we cannot assume that the AC is part of FL/UG, i.e. it must be learned. Ok, how would AC be learned, viz: What is the relevant PLD? Here’s one obvious thing that comes to mind: kids learn the rule by considering its sentential products.[3] What are these? In the simplest case polar questions like those in (1) and their relation to appropriate answers like (2):

(1)  a. Can John run
b. Will Mary sing
c. Is Ruth going home

(2)  a. John can run
b. Mary will sing
c. Ruth is going home

From these the following rule comes to mind:

(3)  To form a polar question: Move the auxiliary to the front. The answer to a polar question is the declarative sentence that results from undoing this movement.[4]

The next step is to complicate matters a tad and ask how well (3) generalizes to other cases, say like those in (4):

(4)  John might say that Bill is leaving

The answer is “not that well.” Why? The pesky ‘the’ in (3). In (4), there is a pair of potentially moveable Auxs and so (3) is inoperative as written. The following fix is then considered:

            (3’) Move the Aux closest to the front to the front.

This serves to disambiguate which Aux to target in (4) and we can go on. As you all no doubt know, the next question is where the fun begins: what does “closest” mean? How do we measure distance? It can have a linear interpretation: the “leftmost” Aux and, with a little bit of grammatical analysis, we see that it can have a hierarchical interpretation: the “highest” Aux. And now the illustration of the POS logic begins: the data in (1), (2) and (4) cannot choose between these options. If this is representative of what there is in the PLD relevant to AC, then the data accessible to the child cannot choose between (3’) where ‘closest’ means ‘leftmost’ and (3’) where ‘closest’ means ‘highest.’ And this, of course, raises the question of whether there is any fact of the matter here. There is, as the data in (5) shows:

(5)  a. The man who is sleeping is happy
b. Is the man who is sleeping happy
c. *Is the man who sleeping is happy

The fact is that we cannot form a polar question like (5c) to which (5a) is the answer and we can form one like (5b) to which (5a) is the answer. This argues for ‘closest’ meaning ‘highest.’ And so, the rule of AC in English is “structure” dependent (as opposed to “linear” dependent) in the simple sense of ‘closest’ being stated in hierarchical, rather than linear, terms.

Furthermore, choice of the hierarchical conception of (3’) is not and cannot be based on the evidence if the examples above are characteristic of the PLD. More specifically, unless examples like (5) are part of the PLD it is unclear how we might distinguish the two options, and we have every reason to think (e.g. based on Childes searches) that sentences like (5b,c) are not part of the PLD. And, if this is all correct, then we have reason for thinking that: (i) that a rule like AC exists in English and whose properties are in part a product of the PLD we find in English (as opposed to Brazilian Portuguese, say) (ii) that AC in English is structure dependent, (iii) that English PLD includes examples like (1), (2) and maybe (4) (though not if we are a degree-0 learners) but not (5) and so we conclude (iv) if AC is structure dependent, then the fact that it is structure dependent is not itself a fact derivable from inspecting the PLD. That’s the simple POS argument.

Now some observations: First, the argument above supports the claim that the right rule is structure dependent. It does not strongly support the conclusion that the right rule is (3’) with ‘closest’ read as ‘highest.’ This is one structure dependent rule among many possible alternatives. All we did above is compare one structure dependent rule and one non-structure dependent rule and argue that the former is better than the latter given these PLD.  However, to repeat, there are many structure dependent alternatives.[5] For example, here’s another that bright undergrads often come up with:

            (3’’) Move the Aux that is next to the matrix subject to the front

There are many others. Here’s the one that I suspect is closest to the truth:

            (3’’) Move Aux

(3’’) moves the correct Aux to the right place using the very simple rule (3’’) in conjunction with general FL constraints. These constraints (e.g. minimality, the Complex NP constraint (viz. bounding/phase theory)) themselves exploit hierarchical rather than linear structural relations and so the broad structure dependence conclusion of the simple argument follows as a very special case.[6] Note, that if this is so, then AC effects are just a special case of Island and Minimality effects. But, if this is correct, it completely changes what an empiricist learning theory alternative to the standard rationalist story needs to “learn.” Specifically, the problem is now one of getting the ML to derive cyclicity and the minimality condition from the PLD, not just partition the class of acceptable and unacceptable AC outputs (i.e. distinguish (5b) from (5c)). I return to a little more discussion of this soon, but first one more observation.

Second, the simple case above uses data like (5) to make the case that the ‘leftmost’ aux cannot be the one that moves. Note that the application of (3’)-‘leftmost’ here yields the unacceptable string (5c). This makes it easy to judge that (3’)-‘leftmost’ cannot be right for the resulting string is clearly unacceptable regardless of what it is intended to mean. However, using this sort of data is just a convenience for we could have reached the exact same conclusion by considering sentences like (6):

(6)  a. Eagles that can fly swim
b. Eagles that fly can swim
c. Can eagles that fly swim

(6c) can be answered using (6b) not (6a). The relevant judgment here is not a simple one concerning a string property (i.e. it sounds funny) as it is with (5c). It is rather unacceptability under an interpretation (i.e. this can’t mean that, or, it sounds funny with this meaning). This does not change the logic of the example in any important way, it just uses different data, (viz. the kind of judgment relevant to reaching the conclusions is different).

Berwick, Pietroski, Yankama and Chomsky (BPYC) emphasize that data like (6), what they dub constrained homophony, best describes the kind of data linguists typically use and have exploited since, as Chomsky likes to say, “the earliest days of generative grammar.” Think: flying planes can be dangerous, or I saw the woman with the binoculars, and their disambiguating flying planes is/are dangerous and which binoculars did you see the woman with.  At any rate, this implies that the more general version of the AC phenomena is really independent of string acceptability and so any derivation of the phenomenon in learning terms should not obsess over cases like (5c). They are just not that interesting for the POS problem arises in the exact same form even in cases where string acceptability is not a factor.

Let’s return briefly to the first point and then wrap up. The simple discussion concerning how to interpret (3’) is good for illustrating the logic of POS. However, we know that there is something misleading about this way of framing the question. How do we know this? Well, because, the pattern of the data in (5) and (6) is not unique to AC movement. Analogous dependencies (i.e. where some X outside of the relative clause subject relates to some Y inside it) are banned quite generally. Indeed, the basic fact, one, moreover that we all have known about for a very long time, is that nothing can move out of a relative clause subject. For example: BPYC discuss sentences like (7):

(7)  Instinctively, eagles that fly swim

(7) is unambiguous, with instinctively necessarily modifying fly rather than swim. This is the same restriction illustrated in (6) with fronted can restricted in its interpretation to the matrix clause. The same facts carry over to examples like (8) and (9) involving Wh questions:

(8)  a. Eagles that like to eat like to eat fish
b. Eagles that like to eat fish like to eat
c. What do eagles that like to eat like to eat

(9)  a. Eagles that like to eat when they are hungry like to eat
b. Eagles that like to eat like to eat when they are hungry
c. When do eagles that like to eat like to eat

(8a) and (9a)  are appropriate answers to (8c) and (9c) but (8b) and (9b) are not. Once again this is the same restriction as in (7) and (6) and (5), though in a slightly different guise. If this is so, then the right answer as to why AC is structure dependent has nothing to do with the rule of AC per se (and so, plausibly, nothing to do with the pattern of AC data). It is part of a far more general motif, the AC data exemplifying a small sliver of a larger generalization. Thus, any account that narrowly concentrates on AC phenomena is simply looking at the wrong thing! To be within the ballpark of the plausible (more pointedly, to be worthy of serious consideration at all), a proffered account must extend to these other cases of as well. That’s the problem in a nutshell.[7]

Why is this important? Because criticisms of the POS have exclusively focused on the toy example that Chomsky originally put forward to illustrate the logic of POS.  As noted, Chomsky’s original simple discussion more than suffices to motivate the conclusion that G rules are structure dependent and that this structure dependence is very unlikely to be a fact traceable to patterns in the PLD. But the proposal put forward was not intended to be an analysis of ACs, but a demonstration of the logic of the POS using ACs as an accessible database. It’s very clear that the pattern attested in polar questions extends to many other constructions and a real account of what is going on in ACs needs to explain these other data as well. Suffice it to say, most critiques of the original Chomsky discussion completely miss this. Consequently, they are of almost no interest.

Let me state this more baldly: even were some proposed ML able to learn to distinguish (5c) from other sentences like it (which, btw, seems currently not to be the case), the problem is not just with (5c) but sentences very much like it that are string kosher (like (6)). And even were they able to accommodate (6) (which so far as I know, they currently cannot) there is still the far larger problem of generalizing to cases like (7)-(9). Structure dependence is pervasive, AC being just one illustration. What we want is clearly an account where these phenomena swing together; AC, Adjunct WH movement, Argument Wh Movement, Adverb fronting, and much much more.[8] Given this, the standard empiricist learning proposals for AC are trying (and failing) to solve the wrong problem, and this is why they are a waste of time. What’s the right problem? Here’s one: show how to “learn” the minimality principle or Subjacency/Barriers/Phase theory from PLD alone. Now, were that possible, that would be interesting. Good luck.

Many will find my conclusion (and tone) harsh and overheated. After all isn’t it worth trying to see if some ML account can learn to distinguish good from bad polar questions using string input? IMO, no. Or more precisely, even were this done, it would not shed any light on how humans acquire AC. The critics have simply misunderstood the problem; the relevant data, the general structure of the phenomenon and the kind of learning account that is required. If I were in a charitable mood, I might blame this on Chomsky. But really, it’s not his fault. Who would have thought that a simple illustrative example aimed at a general audience should have so captured the imagination of his professional critics! The most I am willing to say is that maybe Fodor is right and that Chomsky should never have given a simple illustration of the POS at all. Maybe he should in fact be banned from addressing the uninitiated altogether or only if proper warning labels are placed on his popular works.

So, to end: why am I not impressed by empiricist discussions of AC? Because I see no reason to think that this work has yielded or ever will yield any interesting insights to the problems that Chomsky’s original informal POS discussion was intended to highlight.[9] The empiricist efforts have focused on the wrong data to solve the wrong problem.  I have a general methodological principle, which I believe I have mentioned before: those things not worth doing are not worth doing well. What POS’s empiricist critics have done up to this point is not worth doing. Hence, I am, when in a good mood, not impressed. You shouldn’t be either.






[1] One point before getting down and dirty: what follows is not at all original with me (though feel free to credit me exclusively). I am repeating in a less polite way many of the things that have been said before. For my money, the best current careful discussion of these issues is in Berwick, Pietroski, Yankama and Chomsky (see link to this below). For an excellent sketch on the history of the debate with some discussion of some recent purported problems with the POS arguments, see this handout by Howard Lasnik and Juan Uriagereka.
[2] I believe (actually I know, thx Howard) that the case is first discussed in detail in Language and Mind (L&M) (1968:61-63). The argument form is briefly discussed in Aspects (55-56), but without attendant examples. The first discussion with some relevant examples is L&M. The argument gets further elaborated in Reflections on Language (RL) and Rules and Representations (RR) with the good and bad examples standardly discussed making their way prominently into view. I think that it is fair to say that the Chomsky “analysis” (btw, these are scare quotes) that has formed the basis of all of the subsequent technical discussion and criticism is first mooted in L&M and then elaborated in his other books aimed at popular audiences. Though the stuff in these popular books is wonderful, it is not LGB, Aspects, the Black Book, On Wh movement, or Conditions on transformations. The arguments presented in L&M, RL and RR are intended as sketches to elucidate central ideas. They are not fully developed analyses, nor, I believe, were they intended to be. Keep this in mind as we proceed.
[3] Of course, not sentences, but utterances thereof, but I abstract from this nicety here.
[4] Those who have gone through this know that the notion ‘Aux’ does not come tripping off the tongue of the uninitiated. Maybe ‘helping verb,’ but often not even this.  Also, ‘move’ can be replaced with ‘put’ ‘reorder’ etc.  If one has an inquisitive group, some smart ass will ask about sentences like ‘Did Bill eat lunch’ and ask questions about where the ‘did’ came from. At this point, you usually say (with an interior smile), to be patient and that all will be revealed anon.
[5] And many non-structure dependent alternatives, though I leave these aside here.
[6] Minimality suffices to block (4) where the embedded Aux moves to the matrix C. The CNPC suffices to block (5c). See below for much more discussion.
[7] BTW, none of this is original with me here. This is part of BPYC’s general critique.
[8] Indeed, every case of A’-movement will swing the same way. For example: in It’s fresh fish that eagles that like to eat like to eat, the focused fresh fish is complement of the matrix eat not the one inside the RC.
[9] Let me add one caveat: I am inclined to think that ML might be useful in studying language acquisition combined with a theory of FL/UG. Chomsky’s discussion in Chapter 1 of Aspects still looks to me very much like what a modern Bayesian theory with rich priors and a delimited hypothesis space might look like. Matching Gs to PLD even given this, does not look to me like a trivial task (and work by those like Yang, Fodor, Berwick) strike me as trying to address this problem. This, however, is very different from the kind of work criticized here, where the aim has been to bury UG not to use it. This has been a both a failure and, IMO, a waste of time.

Tuesday, September 24, 2013

When UG?

MP makes the working assumption that whatever happened to allow FL to emerge in its current state happened recently in evo time. This, in turn relies on assuming that precursors of us were without our UG (though they may have had quite a bit of other stuff going on between the ears, in fact, they MUST have had quite a bit of stuff going on there). This assumption was recently challenged by Dediu and Levinson (D&L). Here's an evaluation of their paper by Berwick, Hauser and Tattersall (BHT) (here). BHT argue that there is no there there, a feature, it appears, of much of Levinson's current oeuvre (see here). They observe that the evidence for the quick time frame is sorta/kinda supported by the archeological record, but that such evidence can hardly be dispositive as it is not fine grained enough to address the properties of "the core linguistic competence" of our predecessors as this "does not fossilize." However, such that exists does appear to (weakly) support the envisaged timeframe proposed (roughly 100,000 years). Indeed, as BHT note, D&L misrepresent an important source (Somel et. al) which, concludes, contrary to D&L that: "There is accumulating evidence that human brain development was fundamentally reshaped through several genetic events within the short time space between the human-Neandertahl split and the emergence of modern humans."

So take a look. The D&L paper got a lot of play, but if BHT are right (that's where my money is) then it's pretty much a time sink with little to add to the discussion. You surprised? I'm not. But read away.

Wednesday, January 16, 2013

More on Darwin's Problem


Berwick, Friederici, Chomsky and Bolhuis (BFCB) have a newpaper that discusses Darwin’s Problem and Broca’s Problem (i.e. how brains embody language). The paper is a good short review of some of the relevant issues.  I found parts very provocative and timely. Parts confusing. Here are some (personal) highlights with commentary.

1. BFCB review two facts that set boundary conditions on any evolutionary speculations; (i) “[h]uman language appears to be a recent evolutionary development” (roughly in the last 100,000 years citing Tattersall) and “the capacity for language has not evolved in any significant way since human ancestors left Africa” (roughly 50-80,000 years ago). In sum “that the human language faculty emerged suddenly in evolutionary time and has not evolved since. (p.1)”  These two features suggest two conclusions.

First that UG emerged more or less fully formed and that whatever precipitated its emergence was something pretty simple. It was simple in two ways. It’s design structure was not the result of a lot of selective massaging and whatever triggered the change must have been pretty minimal, e.g. one mutation. I use ‘precipitate’ deliberately. The suggested picture is of a chemical reaction where the small addition of a single novel element results in a drastic qualitative change. For language the idea is that some small addition to the pre-existing cognitive apparatus results in the distillation of FL/UG.

I like this picture a lot. It is the one that Chomsky has presented several times in outlining target of minimalist speculation. If one assumes (as I do) that GB (or its very near cousins, viz. LFG, GPSG, HPSG, RG etc.), for example, roughly describes FL/UG then the project is to try to understand how something of this apparent complexity is actually quite simple.  This will involve two separate but related projects: (i) eliminating the internal modularity of FL/UG as described by GB and (ii) showing that many of the operational constraints are actually reflections of more general cognitive/computational features of mammal minds.

I have discussed (ii) in various other posts (see here, here and here). As regards (i), Chomsky’s unification of Ross’s Islands via Subjacency and ‘Move alpha’ (see ‘On Wh Movement’), offers a good model of what to look for, though the minimalist unification envisioned here is far more ambitious as it involves unifying domains of grammar that Generative Grammar (GG) has taken to be very different from day one. For example, since the get-go GG has distinguished phrase structure rules from movement rules and both from construal rules. Unificationist ambitions (aka: theoretical hubris?) motivate trying to reduce these apparently distinct kinds of rules to a common core.  You gentle readers will no doubt know of certain current suggestions of how to unify Phrase Structure and Movement rules as species of Merge (E and I respectively). There has also been a small (in my humble opinion, much too  small!) industry aiming to unify movement and control (yours truly among others, efforts reviewed in Boeckx, Hornstein andNunes) and movement and binding (starting with Chomsky’s adoption of Lebeaux’s suggestion regarding reflexives in Knowledge of Language). From my seat in the peanut gallery, these attempts have been very suggestive and largely persuasive (I would think that wouldn’t I?), though there are still some puzzles to be tamed before victory is declared.  At any rate, aside from unification being a general scientific virtue, the project gains further empirical motivation in the context of Darwin’s problem given the boundary conditions adumbrated in BFCB.

The second consequence is that the evolution of FL/UG has little to do with natural selection (NS). Why? If FL emerged 100,000 years ago and humanity started dispersing 80,000 years ago then this leaves a very short time for NS to work its (generally assumed) gradual magic. Note whatever took place must have happened entirely before the move out of Africa for otherwise we would expect group variation in FL/UG. If NS was the prime factor in the evolution of FL/UG why did it stop after a mere 20,000 years. Did NS only need 20,000 years to squeeze out all the possible variation?  If so, there couldn’t have been much to begin with (i.e. the system that emerged was more or less fully formed). If not, then why do we see no variation in FL/UG across different groups of humans. Over the last 40,000 years we have encountered many isolated groups of people, with very distinctive customs living in very diverse and remote environments. Despite these manifest differences all humans share a common FL, as attested to by the fact that kids from any of these groups can learn the language of any other in essentially the same way (even the Piraha!). The absence of any perceptible group differences in FL/UG suggests that NS did not drive the change from pre-grammatical to grammatical or if it did so then there was very little variation to begin with.

2. BFCB provide an argument that communication is “ancillary to language design.” This relates to a previous post (here) where I discussed two competing evolutionary scenarios, one driven by communication, the other by enhanced cognition.  As even a non-careful reader will have surmised, I am sympathetic to the second scenario. However, truth be told, I don’t understand the argument BFCB provide for the claim that communicative efficacy is only (at best) a secondary consideration for grammar design. The paper notes “the deletion of copies” which “make[s] sentence production easier renders sentence perception harder.” They conclude from this that deletion “follows the computational dictates of factor (iii)” (i.e. third factor concerns) over the “principle of communicative efficiency.”  This, they continue, supports the conclusion “that externalization (a fortiori communication) is ancillary to language design.(4)”

Here’s what I don’t get: why is communicative efficiency measured by easing the burden on the receiver rather than on the sender? Why is ease of interpretation diagnostic of a communicative end but ease of expression is not and is taken instead to reflect third factor concerns? Inquiring minds would love an answer as this presumed asymmetry appears to license the conclusion that deletion is a third factor consequence and that mapping to AP is a late accretion. I don’t see the logic here.

Moreover, doesn’t this further imply that there is no deletion on the way to CI? And don’t we regularly assume that there are “LF” deletions (e.g. see Chomsky’s 1993 paper that launched the Minimalist Program).  Why should there be “deletion” of copies at LF if deletion is simply a way of reducing the computational burden arising from having to express phonological material. I don’t get it. Help!

3. The paper has an important discussion of human lexicalization and what it means for Darwin’s problem. Human language has two distinctive features.

The first is the nature of the computational system, viz. it’s hierarchical recursion. I’ve discussed this elsewhere (here and here) so I will spare you more of the same.

The second concerns computational atoms, i.e. words.  There are at least two amazing things about them. First, we have soooo many and they are learned soooo quickly! Words are “learned with amazing rapididty, one per waking hour at the peak period of language acquisition. (5)” I’ve discussed some of this before and Lila has chimed in with various correctives.  However, as syntacticians like me tend to focus on grammar, rather than words, the large size and rapid speed of vocabulary acquisition bears constant repeating. Just like no other animal has anything like human grammatical structure, no other animal has anything quite like our lexicon, either quantitatively or qualitatively.

Let’s spend a second on these qualitative features.  As BFCB note human lexical items “appear to be radically different from anything found in animal communication. (4)” In discussing work by Laura Petitto (one of Nym Chimpsky’s original handlers), BFCB highlight her conclusion that “chimps do not really have “names for things” at all. They only have a hodge-podge of loose associations,” in contrast with even the youngest children whose earliest words are “used in a kind-concept constrained way” (5). 

In fact these lexical constraints are remarkably complex.  Chomsky has repeatedly noted (e.g. starting with Reflections on Language and in almost every subsequent philo book since) that “[e]ven the simplest elements of the lexicon do not pick out (‘denote’) mind independent entities. Rather their regular use relies crucially on the complex ways in which humans interpret the world: in terms of such properties as psychic continuity, intention and goal, design and function, presumed cause and effect, Gestalt properties and so on” (5). This raises Plato’s problem in the domain of lexical acquisition and, given the vast noted difference between human lexical concepts and animal “words,” a strong version of Darwin’s problem as well.

It would be nice if we could tie the two distinctive features of human language (viz. unbounded hierarchical structure and vast and intricate vocabulary) together somehow.  Boy would it be nice. I have a rough rule of scientific thumb: keep miracles to a minimum! We already need (at least) one for grammar, now it looks like we need a second for the human lexicon. Can these be related? Please?!

Here’s some idle speculation with the hope that wishing might make it so.  First consider the complexity of lexicalized concepts. In the previous post on Darwin's Problem, I noted the H-VSK hypothesis that what grammar adds is the capacity to combine predicates from otherwise encapsulated modules together into single representations. I suggested that this fits well with the autonomy of syntax thesis, which is basically another way of describing the fact that grammatical operations, unlike module internal operations, are free to apply to predicates independently of their module specific “meanings.”  Autonomy, in effect, allows distinct predicates to be brought together.  The power to conjoin properties from different modules smells similar to what we find lexical items in human language doing, viz. they combine disparate features from various modules (natural physics, persons, natural biology etc.) to construct complex predicates.  If so, the emergence of an abstract syntax may be a pre-condition for the formation of human lexical concepts, which are complex in that they combine features from various informationally encapsulated modules (Note: this conjecture has roots in earlier speculations of the Generative Semanticists). 

Let’s now address the size of the lexicon and its speed of acquisition. Lila noted in her posts (see here and her paper here) that syntactic bootstrapping lies behind the explosion of lexical acquisition that we witness in kids. Until the syntax kicks in, lexical items are acquired slowly and laboriously. After it kicks in, we get an explosion in the growth of the lexicon.  So, following H-VSK, grammar matters in forming complex predicates (aka lexical items) as it allows words to combine cross module features and, following Gleitman and colleagues, grammar underpins the explosive growth of the lexicon.  If this is correct, then maybe the two miracles are connected, but please don’t ask me for the details. As I said, this is VERY speculative.

4. Broca’s problem gets a lot of airtime in BFCB.  I am no expert in the matters discussed, but I confess to having been more skeptical than I expected about the reported results.  Friends more knowledgeable than I am in these matters tell me that the reported results are extremely contentious and that very reasonable people strongly disagree with the specific reported claims. Here is a paper by Rogalsky and Hickok that reviews the evidence that Broca’s area shows specific sensitivity to syntactic structure. Their conclusion does not fit well with that in BFCB: “…our review leads us to conclude that there is no compelling evidence that there are sentence specific processing regions within Broca’s area” (p. 1664). Oh well.

To end: IMO, the most valuable part of BFCB, is how it frames Darwin’s problem in the domain of language.  It correctly stresses that before addressing the evolution question we need to know what it is that we think evolved; the basic design of the system.  FL has two parts: a hierarchical recursive syntax and a conceptually distinctive and very large lexicon. FL seems to have sprung up very rapidly. The Minimalist Program asks how this could have happened and has started to engage in the unification necessary to offer a reasonable conjecture. It’s a sign of a fecund research program that it renews itself by adding new questions and combining them with old results to make new conjectures and launch new research projects. As BFCB show, by this measure, Generative Grammar is alive and well. 

Monday, November 26, 2012

Merging Birds


In the last several years I have become a really big fan of singing mice.  It seems that unbeknownst to us, these white little fur balls have been plunging from aria to aria while gorging on food pellets and simultaneously training their ever-vigilant grad student minders to react appropriately whenever they pressed a bar.  Their songs sound birdish though at a higher pitch. Now it seems that many kinds of mice sing, not only those complaining of incarceration. I was delighted and amazed (though as my daughter pointed out, we’ve known since the first Feival film that mice are great singers). 

I don’t know how extensively rodent operettas have been studied, but recently there has been a lot of research on the structure of bird song and interesting speculation about what it may tell us about the species specificity of the kind of hierarchical recursion we find in natural language (NL). Berwick, Beckers, Okanoya and Bolhuis (BBOB; hmm, kind of a stuttering version of Berwick’s first name) provide an extensive linguist friendly review of the relevant literature which I recommend to the ornithophile with interests in UG. 

BBOB’s review is especially relevant to anyone interested in the evolution of the faculty of language (FL) (ahem, I’m talking to all you minimalists out there!). They note “many striking parallels between speech and vocal production and learning in birds and humans” but also note qualitative differences “when one compares language syntax and birdsong more generally (5/1).” The value of the review, however, is not in these broad conclusions but in the detailed comparisons between phonological vs syntactic vs birdsong structure that it outlines. In particular, both birdsong and the human sound system display precedence based dependencies (1st order markov), adjacency-based dependencies, some (limited) non-adjacent dependencies, and the grouping of elements into “chunks” (“phrases,” “syllables”).  In effect, birdsongs seem restricted to linear precedence relations alone, just what Heinz and Idsardi propose suffices to represent the essentials of the human sound system. Importantly, there is no evidence that birdsong allows for the kind of hierarchical recursion that is typical of syntactic structures:

Birdsong does not admit such extended self-nested structures, even in the nightingale song chunks are not contained within other song chunks, or song packets within other song packets or contexts within contexts (5/6) (my emphasis).

Nor do they provide any evidence for unbounded dependencies, unboundedly hierarchical asymmetric “phrases,” or displacement relations (aka movement), all characteristic features of NLs.

The BBOB paper also contains an interesting comparison of songbird and human brains remarking on various possible shared vocalization homologies in human and bird brain architecture. Even FoxP2, (that ubiquitous rascal) makes a cameo appearance, with BBOB briefly reviewing the current speculations concerning how “this system may be part of a “molecular toolkit that is essential for sensory-guided motor learning” in the relevant regions of songbirds and humans (5/9).”

All in all then I found this a very useful guide to the current state of the art, especially for those with minimalist interests.

Why minimalists in particular? Because it has possible bearing on a currently active speculation regarding the species specificity and domain specificity of Merge.  Merge, recall, is the minimalist replacement for phrase structure rules (and movement). It’s the operation responsible both for unbounded hierarchical embedding and displacement.  So if birdsong displays context free patterns one source for this could be the presence of Merge as a basic operation in the songbird brain. BBOB carefully review the evidence that birdsong patterns exceed the descriptive power of finite transition networks and demand the resources of context free grammars. They conclude that there is currently “no compelling evidence” that they do (5/14). Furthermore, BBOB note that there is no evidence for displacement-like operations in birdsong, the second product of a merge-like operation. Thus, at this time, NLs alone provide clear evidence of context free and displacement structures. So, if Merge is the operation that generates such structures, there is currently no evidence that Merge has arisen in any species other than humans or in any domain other than syntax.

Why is this important for minimalists? The minimalist Genesis story goes as follows: Some “miracle” occurred in the last 100,000 years that allowed for NLs to arise in humans. Following Chomsky, let’s call this miracle “Merge.” By hypothesis, Merge is a very “simple” addition to the cognitive repertoire. Conceptually, there are (at least) two ways it might have been added: (i) Merge is a linguistically specific miracle or (ii) it is a more general cognitive one. If (ii), then we might expect Merge to have arisen before in other species and to be expressed in other cognitive domains, e.g. birdsong.  This is where BBOB’s conclusions are important for they indicate that there is currently no evidence in birdsong for the kind of structures (i.e. ones displaying unbounded nested dependencies and displacement) Merge would generate. Thus, at present, the only cognitive products of Merge we have found occur in species that have NLs, i.e. us.

Moreover, as BBOB emphasize the impact of Merge is only visible in a subpart of our linguistic products. It is a property of syntactic structures not phonological ones.  Indeed, as BBOB show, human sound systems and birdsong systems look very similar.  This suggests that Miracle Merge is quite a picky operation, exercising its powers in just a restricted part of FL (widely construed).  So not only is Merge not cognitively general, it’s not even linguistically general. Its signature properties are restricted to syntactic structures.

If this is correct, then it suggests (to me at least) that Merge is a linguistically local miracle and so proprietary to FL and so part of UG. This, I believe, comports more with Chomsky’s earlier conception of Merge, than his current one.  The former sees the capacity to build bigger and bigger hierarchically embedded structures (and movement) as resting on being able to spread “edge features” (EF) from lexical items to the complexes of lexical items that Merge forms.  So given two lexical items (LI) (each with an inherent EF), a complex inherits an EF (presumably from one of its participants) and this inherited EF is what licenses the further merging of the created complex with other EF bearing elements (LIs and earlier products of Merge). Inherited EFs then are essentially the products of labeling (full disclosure: I confess to liking this idea as I outlined/adopted a version of it here (Btw, it makes a wonderful stocking stuffer so buy early buy often!) and labeling is the miracle primarily responsible for the e(/I)mergence (like that?) of both phrase structure and displacement.

Chomsky’s more current view seems to be that labeling (and so EFs) are dispensable and that Merge alone is the source of phrase structure and movement. There is no need for EFs as Merge is defined as being able to apply to any cognitive objects at all, primitive or constructed.  In particular, both lexical items and complexes of lexical items formed by prior applications of Merge are in the domain of Merge. EFs are unnecessary and so, with a hat tip to Ockham, should be dispensed with. 

And this brings us back to birds, their songs and their brains.  It would have been a powerful piece of evidence in favor of this latter conception were a signature of merge attested in the cognitive products of some other species for it would have been evidence that the operation isn’t FL/UG peculiar.  Birdsong was a plausible place to look and it appears that it isn’t there.  BBOB’s review locates the effects of Merge exclusively to the syntax of NL.  Were Merge more domain general and less species specific we might have expected other dogs to bark (or sing more complex songs). And though absence of evidence should not be mistaken for evidence of absence, at least right now, it looks like Merge is very domain specific, something more compatible with Chomsky’s first version of Merge than his second.