Comments

Showing posts sorted by relevance for query does anyone learn anything. Sort by date Show all posts
Showing posts sorted by relevance for query does anyone learn anything. Sort by date Show all posts

Monday, April 8, 2013

Give Nim, Give Me



(Photo credit: Herb Terrace)

Einstein was a very late talker. “The soup is too hot”, as the legend has it, were his first words at the very ripe age of three. The boy genius hadn't seen anything worth commenting on.

The credulity of such tales aside, they do contain a kernel of truth: a child doesn’t have to say something, anything, just because he can. This poses a challenge for the study of child language, since naturalistic production is often the only, and certainly the most accessible, data at hand. A child’s linguistic knowledge may not be fully reflected in their speech, which we have known since Lila’s deconstruction of the telegraphic stage. Some expressions may not show up because we haven’t waited long enough, while others—an extraction violation, for instance—will never be said for they are unsayable.

In recent years, what-you-say-is-what-you-know appears to be gaining popularity, as interests in usage based theories of language are on the rise. Here is a warmup. The expression "give me" is proposed as a frozen phrase (Lieven et al. 1992, Tomasello 2010), rather than syntactically composed, spawning cottage industries such as "formulaic languages", which some regard as a transient stage in language evolution (Wray 1998). True, "give" and "me" make a good tag team: "give me freedom", "give me cheese", "give me now" … “gimme coffee” (an old favorite of mine), and they dwarf other combinations.  Take the speech of Adam, Eve and Sarah from Roger Brown's classic study: the frequencies of "give me", "give him", and "give her" are:

95 (93 give me, 2 gimme): 15 (give him): 12 (give her), or 7.91 : 1.23 : 1

So “give me” does seem especially formulaic ... right? Well, not if you check the frequencies of "me", "him", and "her" from the same three kids:

2870 (me) : 466 (him) : 364 (her), or 7.88 : 1.28 : 1

Nothing much can be concluded from these six numbers but there seems to be pretty good support for the null hypothesis that “give” and pronouns combine completely independently. The Brown data has been around for forty years; it's just nobody had bothered to check. (Use the grep, Luke.)

Nowadays everyone does statistics but we still need reasonable hypotheses to test for and against. Usage based theories have plenty of p-values: one can easily show that the frequencies of "give me/him/her" are statistically significantly different from "chance"--but what is "chance"? If we know anything about the statistics of language, it is that language is not "chance" (Zipf 1949). To make the argument against grammar, one would need to show, at the minimum, that the observed distribution in child language is statistically inconsistent with the predicted distribution of a grammar.  Judiciously chosen null hypotheses are needed, not gut feelings: so long to the "gimme" myth.

A few years ago, Virginia Valian came to Penn to give a talk. It concerned the distribution of determiner-noun combinations in child English. Virginia was the first to show that English children’s determiners are virtually error free (1986), thereby providing evidence for an abstract grammar. Not so quick, the usage-based folks say, because the absence of errors could be the result of children memorizing specific word combinations from adult speech, which would also be error free. (I fully endorse such skepticism.) We need some other statistical benchmarks to show the presence of grammar. 

Diversity is a popular measurement. Suppose there is a rule DP→DN, where D is either a or the, and N stands for a singular noun, yielding "a/the car", "the/a pizza", etc. Shouldn't the interchangeability of "a" and "the", per grammar, be reflected in the diversity of nouns that appear with both of them?  Young children's determiner use, however, only shows 20-40% of diversity (Pine & Lieven 1996); perhaps they just memorize determiner-noun combinations from the adult input (Tomasello 2000, Cognition). 

Along with Stephanie Solt and John Stewart, Virginia showed that mothers' speech contains comparable, and comparably low, diversity measures as their toddlers’ (2009, J. Child Language). After her talk, I pulled out some numbers from the Brown Corpus: not Roger Brown, but the collection of English print materials at Brown University, the grandmother of all modern linguistic corpora. Only 25% of singular nouns that combine with either "a" and "the" combine with both. That's lower than some child samples from Pine & Lieven (1996), so two year olds have a better command of English grammar than professional writers. Now that is absurd. 

One reaction would be to abandon the premise that syntactic diversity is a direct reflection of grammatical complexity. Not a bad idea, and much of the purported evidence for usage based theory vanishes. Another reaction would be to go for the extra credit, by characterizing the statistical profile of syntactic diversity that can be expected from a grammar. If the child used 100 distinct nouns, and paired them with either "a" or "the" 500 times, how many of the 100 will be paired with both, assuming the rule DP→DN is at work? Virginia's work was inspirational. I was also knee deep in Zipfian waters, thanks to the work of Erwin Chan, Constantine Lignos and my colleague Mitch Marcus. They showed that pretty much everywhere you look--words, lemmas, morphological inflections, syntactic rules--language follows Zipf-like distributions, which can be exploited for fun and benefit.

If a sample contains 100 nouns (types), then a good many of them must occur only once since they will inevitably fall on Zipf's long and flat tail: these fellows will never get to meet both determiners.  Even for those that do show up multiple times, they may still be monogamous, just as when you toss a fair coin 3 times, it may land on heads 3 times in a roll. And grammar is no fair coin. Nouns tend to have a favored determiner, even though both combinations are possible. For instance, "the bathroom" is more commonly used than "a bathroom" but we say "a bath" a lot more often than "the bath". These imbalances are probably not a matter of grammar, which presumably does not encode the frequency of bodily needs, but they will conspire to produce low syntactic diversity, and thus the impression of grammatical absence.

After a bit of probability theory exercise [1], we can use a formula to calculate the expected diversity from the sample and vocabulary size (e.g., 500 and 100). The key here is multiplication--the statistical hallmark of independence--of the noun probabilities with determiner-noun combination probabilities, both of which can be well approximated by Zipf’s law. I was surprised to see how well it worked, and in fact had to learn new statistics just to be sure. We mostly use statistics to show one set of values and another (e.g., experimental results vs. "chance") are statistically different, but being different is not the same as being the same. Lin's concordance correlation coefficient (there is an R package, of course), first invented in biostatistics to verify drug effectiveness across trials, confirmed the observation.   In other words, children's syntactic diversity appears exactly what one might expect from a grammar rule, once the general statistical properties of language are taken into account. [2]

Someday we may have a bunch of these statistical profilers, like what evolutionary geneticists use to detect natural selection at the molecular level. Let me make very clear what this work does and does not show. It does show that at least one part of child language makes use of an abstract rule of grammar but it does not mean that all parts of child language do. It does show that children can merge but it does not tell us how they learn what to merge with. It does show the presence of grammatical ability in very young children, but it does not say how that ability got there in the first place, ontogetically or phylogenetically.

Which brings me to Nim Chimpsky and the evolution of language. The continuity between primate language and early child language is believed to hold “the most promising guide to what happened in language evolution” (Hurford 2011, p590), presumably on the apparent formulaic similarities between them. If the numbers worked out for children, who seem to have a grammar after all, perhaps Nim is due for a similar upgrade?  Whatever one thinks of Project Nim--I had to fight back tears--it produced the only publicly available corpus of primate language. Nim acquired about 125 signs of ASL, and produced thousands of multiple sign combinations, the vast majority of which being two sign combinations (Terrace 1979, Nim).  These have been described as rule-like constructions, each consisting of two closed class functors such as “give” and “more,” along with open class items such as “apple,” “Nim,” or “eat.”  Signs do not combine with uniform frequency either, with “eat”, “banana”, “me”, “Nim" etc. among the predictable favorites. What’s Nim’s syntactic diversity if he combined signs under a rule? Run the numbers: the poor guy didn’t seem to have a grammar, just as his trainers concluded (Terrace et al. 1979).

Syntactic diversity in human language, usage based learning and Nim Chimpsky.



Moral for the day: the null hypothesis, once properly formulated, may come back to bite your statistical hand.  All very exciting. When I explained this work to some of my non-linguist friends (I do have a few!), their reaction was one of surprise, though not the kind I had in mind. “Why would anyone think kids learn language by copying us?  Just this morning, Maggie said ___”, to be filled by one of the darndest things kids say.  They do wonder about vocabulary, boys vs. girls, and bilingualism, but no one is remotely concerned about the combinatorics of grammar that are, literally, screaming in their faces. Perhaps linguists do worry too much. 

[1] Thanks to Ruochuan Liu and Qiuye Zhao for spotting an error early on.
[2] Could a usage cum memory-retrieval model account for the same finding? I don’t think one knows for sure,  since it has been difficult to pin down the mechanics of usage based learning so it’s unclear what quantitative predictions it makes. I won’t dwell on the matter here but refer you to the paper under discussion, where a concrete proposal (Tamales 2000, Cognitive Linguistics, p77) is tested but came up short. 

Wednesday, December 12, 2012

Does Anyone Ever Learn Anything?


Let’s follow Jimmy and Judy from birth to about five.  At birth they say precious little. At five you can’t shut them up. What happened in these five years? Answer: they learned their native language, English say. Obvious no?  Yes. But is it right? Did Judy and Jimmy learn English?  Well, to paraphrase a recent political celebrity, it all depends on what you mean by learn and English.  It seems undeniable that Judy and Jimmy developed a capacity absent (or at least invisible[1]) at birth and this capacity can be exercised to converse with some natives, the English speaking ones, but not others, the Mandarin speaking ones. However, does this imply that they learned English? 

Linguists have long understood that labels like English, French, Swahili, Mandarin, etc. are more convenience terms than terms of art (here’s where linguists mention the Weinreich quip that a language is just a dialect with an army and a navy; real cognoscenti adding sotto voce that dialects are just idiolects with epaulettes). However, recent research suggests that we have been far too cavalier about the first half of this doublet.  We can all agree that Judy and Jimmy acquired (competence in) English, but did they learn English? Recent (and, as we shall see, not so recent) research into word learning suggest that here we need to look before we leap, something, it appears, that kids do not do, at least when it comes to early word acquisition. I’ve already discussed some of the research by MSTG (here) which argues (very convincingly in my view) that the early stages of lexical acquisition do not involve the careful statistical weighing of competing alternatives but involves jumping to a conclusion mentally clutched with fierce determination and, with time, forgotten if incorrect, only to set up another ill supported jump into the lexical abyss (boy was that fun to write!). MSTG support this conclusion by considering learning situations less factitious than the contrived set-ups near and dear to the psych lab.  When kid and adults are asked to consider more natural (and hence visually busy) filmed vignettes in which the targets of lexical labeling are not clearly segregated and identified on pristine picture cards, they acquire word meanings all at once or not at all. This result is important for several reasons.

First and foremost, it provides a concrete illustration of why we should not equate acquisition with learning.  MSTG provides diagnostics of learning (multiple hypotheses, statistical weighing of these alternatives, gradual convergence on the right result) and argues that learning so understood fails to hold in more realistic contexts of lexical acquisition. Specific conclusion; in at least one demonstrable case, acquisition does not equal learning.  More general conclusion; it is an empirical question whether in any given acquisition context it is true that learning (now understood to be one mechanism among others for the acquisition of knowledge) is taking place.  In other words, Jimmy and Judy certainly acquired English but whether they learned it is entirely open for empirical grabs. Chomsky’s repeated suggestion that we understand language acquisition as a species of growth rather than learning makes an analogous point.  MSTG makes the case more crisply, I believe, by providing clear diagnostics of learning and showing that there are core cases of “learning” where these signature properties of learning are demonstrably absent.

Second, MSTG provide a rationale for why learning doesn’t hold in their examined cases. Commenting on this (here) I observed that the MSTG results suggest that in such busy contexts the prerequisites for “cross situational learning” do not exist and this is why an alternative acquisition strategy is employed. Following MSTG’s lead, I even suggested that learning requires the kind of structured hypothesis space provided by the contrived set ups that MSTG’s more realistic vignettes argue against. This constituted a kind of compromise position; learning applies where acquirers have well structured hypothesis spaces and leaping to conclusions holds where this fails to hold.[2]  However, (and learn from this you soft-hearted open-minded intellectual compromisers out there) once again Norbert’s natural generosity of spirit and desire for intellectual group-hug kumbaya moments led him astray. It seems that even this compromise position concedes too much to learning mechanisms. In a companion paper, Trueswell, Medina, Hafri and Gleitman (TMHG) extend the MSTG results to include the more stylized learning environments in which options are clearly marked and lexical targets (aka referents) are crisply identified.[3]  Even in the artificial setting of the experimental psych lab kids and adults do the darndest things!

More specifically, like MSTG, TMGH identify the quiddity of “cross situational learning” with the following mechanism:

…keeping track of multiple hypotheses about a word’s meaning across successive learning instances, and gradually converg[ing] on the correct meaning via an intersective statistical process (128).

What they demonstrate is that even in simple stylized psych lab contexts when acquirers “are placed in the initial novice state of identifying word-to-referent mappings across learning instances, evidence for such a multiple hypothesis tracking procedure is strikingly absent (128),” and learners don’t “track the cross-trial statistics in order to build the correct mappings between words and referents (129).”  What acquirers do is “make a single conjecture upon hearing the word used in context and carry that conjecture forward to be evaluated for consistency with the next observed context (129).”  If confirmed, the guess is retained, if disconfirmed, acquirers guess again as if de novo.

TMGH show this (same with MSTG) by considering the dynamics of knowledge fixation: “how learning accuracy unfolded across learning instances (130).” This is very interesting stuff and I strongly recommend that you look at the details. However, just as interesting is the little bit of history that TMGH review. TMGH acknowledge tapping into a long history of criticism of this traditional gradual/comparative conception of learning.  In the late 1950s and early 1960s first Irvin Rock (1957) and then William Estes (1960) used very analogous kinds of arguments to demonstrate that verbal learning was not learning at all but was one-trial guessing. They did not fare so well. As Roediger and Arnold (RA) (2012) put it “the verbal learning establishment rose up to smite down these new ideas (2).”  It is instructive to read the RA paper for it shows the weakness of the counter-arguments used against Rock and Estes that nonetheless carried the day.  Just like today, learning was less a hypothesized mechanism for acquisition than a definitional truth about it.[4] 

The TMHG paper ends with an interesting paradox that I would like to briefly discuss as well. Initial word learning is slow and laborious (35-140 words at 14 months) but unbelievably rapid thereafter (12,000 words by age 6, i.e. roughly 7 words per day for a little under 5 years).  Why the change? TMHG note other work (Lila’s syntactic bootstrapping hypothesis) which proposes that “the acquisition of syntax and other linguistic knowledge by toddlers and young preschoolers during this time period provides a rich database of additional constraints that permit the learning of many additional words.” It is conceivable that with this knowledge in place, cross situational learning might finally become operative, though as TMHG correctly observe, this is decidedly an empirical question and it is “plausible that a propose-but-verify word learning procedure is at work all along the course of word learning throughout most of the life cycle (150).”  It would be interesting in the extreme, in my view were either conclusion correct.

If the latter conclusion proved true (propose and verify all the way down), then language acquisition might have nothing to do with learning in the technical sense. Of course, this is consistent with grammar acquisition being a case of learning, i.e. perhaps we don’t learn words but we do learn parameters. Maybe, but I’d be very skeptical. If even word acquisition isn’t an instance of learning then it seems to me that the burden of proof that any area of language acquisition involves learning would be pretty high.

If, on the other hand, the former option is correct, (viz. that learning only kicks in when grammar is there to buttress it) then its role in accounting for language acquisition would, in my opinion, be quite modest.  Yes, learning plays a role, but really most of the action lies with the constraints that grammars place on the process. This does not mean that we should not study the extras that learning might be adding (though remember the possibility mooted in the prior paragraph), but I doubt that these cognitive titivations will generate much excitement if they only operate in restricted hypothesis spaces. Learning is interesting when there are lots of options that need sifting, not so much when the range of possible end states is highly restricted. 

MSTG and TMHG show us that terminology matters. If we call acquisition “learning” we’ve loaded the research dice. If MSTG and TMHG are right (and I’d bet quite a bit that they are: any suckers out there?) it looks like we’ve repeatedly loaded them to come up snake eyes. It’s time we cut our losses and open our minds to the possibility that ‘language learning’ like ‘Justice Roberts,’ ‘military intelligence,’ and ‘western civilization’ is an oxymoron.



[1] I include this hedge for every week another (usually French speaking) psychologist shows that the youngest kids seem to have the most prodigious knowledge. It seems that we have nothing to teach those little know-it-alls. 
[2] A kind of learning as first resort, jumping to conclusions a last resort, method. C.f. TMHG p. 130.
[3] It is doubtful that the meaning of a lexical item is simply its referent for reasons that Chomsky has belabored (sadly, quite unsuccessfully) over the years. There is far more structure to lexical items than what they refer to. However, for current purposes, this only further dramatizes the inadequacy of learning as a mechanism for lexical acquisition. 
[4] TMHG also cite another paper by Gallistel and friends that I have not yet read but will try to get hold of and blog about when I do. It argues (cited in TMHG 151), that “in most subjects, in most paradigms, the transition from a low level of responding to an asymptotic level is abrupt.” Oh my. I’ll keep you posted.

Monday, February 11, 2013

What's Chomsky Thinking Now


Here’s a short post linking to what I think is a very accessible summary of Chomsky’s current views about language and its biological basis. The discussion is not extended and the cognoscenti will not likely learn anything new (though I did).  Here are some points he touches on:

1.     Distinction between generalizations about language and UG, which is “the genetic basis for language.
2.     That there are no real group differences (vs individual differences) between humans as regards FL/UG, implying that there has been no significant change in FL for a very long time.
3.     Early theories of UG allowed for a huge amount of variation between languages. Over the last 40 years, theory has narrowed the range of this difference.
4.     A simple statement of the goals of a useful theory: “A plausible theory has to account for the variety of languages and the detail that you see in the surface study of languages – and, at the same time, be simple enough to explain how language could have emerged quickly, through some small mutation of the brain, or something like that.”
5.     Analogy between UG and Jacob’s Universal Genome hypothesis.
6.     Argument for species specific native capacity for language starts from the simple observation that humans alone can “pick out anything that is relevant to language” from the “great blooming, buzzing confusion” that is the stimulus input.
7.     Fast mapping of words to meanings indicating taht Quine’s “museum myth” is in fact reality.
8.     Pattern recognition insufficient for language acquisition.
9.     Theory of Mind orthogonal to language acquisition problem.
10.  Linguistic interests are different from Languistic ones.
11.  Is the idea that culture influences language meaningful?

The distinction in (1) is particularly important to reiterate as there has been rampant confusion on this point. Chomsky’s views are not Greenberg’s and a lot of the criticism of Universals has come from running the two together.  I also liked the observations concerning the implications of research on autism for what look like Tomasello-like views about language development. The remarks are blunt, but raise a relevant point.

I also found the observations concerning how little UG has apparently changed over the last 30,000 years important to stress (2). Recall the Papuans (and the Piraha?) who were very isolated until recently are all capable of learning the same languages in the same way as anyone else. In fact, I know of no group of people whose kids suitably located cannot learn any language, on contrast to any other animal. Why? because humans have essentially the same UG and other animal's don't have one. This is sufficient to raise the generative research question: what do we have that they don't and how did it get there? 

Similarly culture changes have left UG pretty much intact (11), so far as we can tell. Older varieties of English, Icelandic, Japanese look from a UG vantage point pretty much like their contemporary counterparts, indicating that the same UG operated then as does now. Chomsky develops these points more elaborately elsewhere, but if you are like me and people ask what you do then this is something short and readable to give them, if, of course, these are the questions you are interested in.

Saturday, July 30, 2016

The sludge at the bottom of the barrel

It is somewhat surprising that Harper’s felt the need to run a hit piece by Tom Wolfe on Chomsky in its August issue (here). True, such stuff sells well. But given that there are more than enough engaging antics to focus on in Cleveland and Philadelphia one might have thought that they would save the Chomsky bashing for a slow news period. It is a testimony to Chomsky’s stature that there is a publisher of a mainstream magazine who concludes that even two national conventions featuring two of the most unpopular people ever to run for the presidency won’t attract more eyeballs than yet another takedown of Noam Chomsky and Generative Grammar (GG).

Not surprisingly, content wise there is nothing new here. It is a version of the old litany. Its only distinction is the over the top nuttiness of the writing (which, to be honest, has a certain charm in its deep dishonesty and nastiness) and its complete disregard for intellectual integrity. And, a whiff of something truly disgusting that I will get to at the very end. I have gone over the “serious” issues that the piece broaches before in discussions of analogous hit jobs in the New Yorker, the Chronicle of Higher Education, and Aeon (see here and here for example). Indeed, this blog was started as a response to what this piece is a perfect example of: the failure of people who criticize Chomsky and GG to understand even the basics of the views they are purportedly criticizing.

Here’s the nub of my earlier observations: Critics like Everett (among others, though he is the new paladin for the discontented and features prominently in this Wolfe piece too) are not engaged in a real debate for the simple reason that they are not addressing positions that anyone holds or has ever held. This point has been made repeatedly (incuding by me), but clearly to no avail. The present piece by Wolfe continues in this grand tradition. Here's what I've concluded: pointing out that neither Chomsky nor GG has ever held the positions being “refuted” is considered impolite. The view seems to be that Chomsky has been rude, sneaky even, for articulating views against which the deadly criticisms are logically refractory. Indeed, the critics refusal to address Chomsky’s actual views suggests that they think that discussing his stated positions would only encourage him in his naughty ways. If Chomsky does not hold the positions being criticized then he is clearly to blame for these are the positions that his critics want him to hold so that they can pummel him for holding them. Thus, it is plain sneaky of him to not hold them and in failing to hold them Chomsky clearly shows what a shifty, sneaky, albeit clever, SOB he really is because any moderately polite person would hold the views that Chomsky’s critics can demonstrate to be false! Given this, it is clearly best to ignore what Chomsky actually says for this would simply encourage him in articulating the views he in fact holds, and nobody would want that. For concreteness, let’s once again review what the Chomsky/GG position actually is regarding recursion and Universal Grammar (UG).

The Wolfe piece in Harper’s is based on Everett’s critique of Chomsky’s view that recursion is a central feature of natural language. As you are all aware, Everett believes that he has discovered a language (Piraha) whose G does not recurse (in particular, that forbids clauses to be embedded within clauses). Everett takes the putative absence of recursion within Piraha to rebut Chomsky’s view that recursion is a central feature of human natural language precisely because he believes that it is absent from Piraha Gs. Everett further takes this purported absence as evidence against the GG conception of UG and the idea that humans come with a native born linguistic facility to acquire Gs.  For Everett human linguistic facility is due to culture, not biology (though why he thinks that these are opposed to one another is quite unclear). All of these Everett tropes are repeated in the Wolf piece, and if repetition were capable of improving the logical relevance of non-sequiturs, then the Wolfe piece would have been a valuable addition to the discussion.

How does the Everett/Wolfe “critique” miss the mark? Well, the Chomsky-GG view of recursion as a feature of UG does not imply that every human G is recursive. And thinking that it does is to confuse Chomsky Universals (CU) with Greenberg Universals (GU). I have discussed this before in many many posts (type in ‘Chomsky Universals’ or ‘Greenberg Universals’ in the search box and read the hits). The main point is that for Chomsky/GG a universal is a design feature of the Faculty of Language (FL) while for Greenberg it is a feature of particular Gs.[1] The claim that recursion is a CU is to say that humans endowed with an FL construct recursive Gs when presented with the appropriate PLD. It makes no claim as to whether particular Gs of particular native speakers will allow sentences to licitly embed within sentences. If this is so, then Everett’s putative claim that Piraha Gs do not allow sentential recursion has no immediate bearing on the Chomsky-GG claims about recursion being a design feature of FL. That FL must be able to construct Gs with recursive rules does not imply that every G embodies recursive rules. Assuming otherwise is to reason fallaciously, not that such logical niceties have deterred Everett and friends.

Btw: I use ‘putative claim’ and ‘purported absence’ to highlight an important fact. Everett’s empirical claims are strongly contested. Nevins, Pesetsky and Rodrigues (NPR) have provided a very detailed rebuttal of Everett’s claims that Piraha Gs are recursiveless.[2] If I were a betting man, my money would be in NPR. But for the larger issue it doesn’t matter if Everett is right and NPR are wrong. Thus, even were Everett right about the facts (which, I would bet that he isn’t) it would be irrelevant to his conclusion regarding the implications of Piraha for the Chomsky/GG claims concerning UG and recursion.

So what would be relevant evidence against the Chomsky/GG claim about the universality of recursion? Recall that the UG claim concerns the structure of FL, a cognitive faculty that humans come biologically endowed with. So, if the absence of recursion in Piraha Gs resulted from the absence of a recursive capacity in Piraha speakers’ FLs then this would argue that recursion was not a UG property of human FLs. In other words, if Piraha speakers could not acquire recursive Gs then we would have direct evidence that human FLs are not built to acquire recursive Gs. However, we know that this conditional is FALSE. Piraha kids have no trouble acquiring Brazilian Portuguese (BP), a language that everyone agrees is the product of a recursive G (e.g. BP Gs allow sentences to be repeatedly embedded within sentences).[3]  Thus, Piraha speakers’ FLs are no less recursively capable than BP speakers’ FLs or English speakers’ FLs or Swahili speakers’ FLs or... We can thus conclude that Piraha FLs are just human FLs and have as a universal feature the capacity to acquire recursive Gs.

All of this is old hat and has been repeated endlessly over the last several years in rebuttal to Everett’s ever more inflated claims. Note that if this is right, then there is no (as in none, nada, zippo, bubkis, gornisht) interesting “debate” between Everett and Chomsky concerning recursion. And this is so for one very simple reason. Equivocation obviates the possibility of debate. And if the above is right (and it is, it really is) then Everett’s entire case rests on confusing CUs and GUs. Moreover, as Wolfe’s piece is nothing more than warmed over Everett plus invective, its actual critical power is zero as it rests on the very same confusion.[4]

But things are really much worse than this. Given how often the CU/GU confusion has been pointed out, the only rational conclusion is that Everett and his friends are deliberately running these two very different notions together. In other words, the confusion is actually a strategy. Why do they adopt it? There are two explanations that come to mind. First, Everett and friends endorse a novel mode of reasoning. Let’s call it modus non sequitur, which has the abstract form “if P why not Q.”  It is a very powerful method of reasoning sure to get you where you want to go. Second possibility: Everett and Wolfe are subject to Sinclair’s Law, viz. It is difficult to get a man to understand something when his salary depends upon his not understanding it. If we understand ‘salary’ broadly to include the benefits of exposure in the high brow press, then … All of which brings us to Wolfe’s Harper’s piece.

Happily for the Sinclair inclined, the absence of possible debate does not preclude the possibility of considerable controversy. It simply implies that the controversy will be intellectually barren. And this has consequences for any coverage of the putative debate.  Articles reprising the issues will focus on personalities rather than substance, because, as noted, there is no substance (though, thank goodness, there can be heroes engaging in the tireless (remunerative) pursuit of truth). Further, if such coverage appears in a venue aspiring to cater to the intellectual pretensions of its elite readers (e.g. The New Yorker, the Chronicle and, alas, now Harper’s) then the coverage will require obscuring the pun at the heart of the matter. Why? Because identifying the pun (aka equivocation) will expose the discussion as, at best, titillating gossip for the highbrow, at middling, a form of amusing silliness (e.g. perfect subject matter for Emily Litella) and, at worst, a form of celebrity pornography in the service of character assassination. Wolfe’s Harper’s piece is the dictionary definition of the third option.

Why do I judge Wolfe’s article so harshly? Because he quotes Chomsky’s observation that Everett’s claims even if correct are logically irrelevant. Here’s the full quote (39-40):
“It”—Everett’s opinion; he does not refer to Everett by name—“amounts to absolutely nothing, which is why linguists pay no attention to it. He claims, probably incorrectly, it doesn’t matter whether the facts are right or not. I mean, even accepting his claims about the language in question—Pirahã—tells us nothing about these topics. The speakers of this language, Pirahã speakers, easily learn Portuguese, which has all the properties of normal languages, and they learn it just as easily as any other child does, which means they have the same language capacity as anyone else does.”
A serious person might have been interested in finding out why Chomsky thought Everett’s claims “tell us nothing these topics.” Not Wolfe. Why try to understand issues that might detract from a storyline? No, Wolfe quotes Chomsky without asking what he might mean. Wolfe ignores Chomsky's identification of the equivocation as soon as he notes it. Why? Because this is a hit piece and identifying the equivocation at the heart of Everett’s criticism would immediately puncture Wolfe’s central conceit (i.e. heroic little guy slaying the Chomsky monster).

Wolfe clearly hates Chomsky. My reading of his piece is that he particularly hates Chomsky’s politics and the article aims to discredit the political ideas by savaging the man. Doing this requires demonstrating that Chomsky, who, as Wolfe notes is one of the most influential intellectuals of all time, is really a charlatan whose touted intellectual contributions have been discredited. This is an instance of the well know strategy of polluting the source. If Chomsky’s (revolutionary) linguistics is bunk then so are his politics. A well-known fallacy this, but not less effective for being so. Dishonest and creepy? Yes. Ineffective? Sadly no.

So there we have it. Another piece of junk, but this time in the style of the New Journalism. Before ending however, I want to offer you some quotes that highlight just how daft the whole piece is. There was a time that I thought that Wolfe was engaging in Sokal level provocation, but I concluded that he just had no idea what he was talking about and thought that stringing technical words together would add authority to his story. Take a look at this one, my favorite (p. 39):

After all, he [i.e. Chomsky, NH] was very firm in his insistence that it [i.e. UG, NH] was a physical structure. Somewhere in the brain the language organ was actually pumping the UG through the deep structure so that the LAD, the language acquisition device, could make language, speech, audible, visible, the absolutely real product of Homo sapiens’s central nervous system. [Wolfe’s emphasis, NH].

Is this great, or what! FL pumping UG through the deep structure. What the hell could this mean? Move over “colorless green ideas sleep furiously” we have a new standard for syntactically well-formed gibberish. Thank you Mr Wolfe for once again confirming the autonomy of syntax.

Or this encomium to cargo cult science (37):

It [Everett’s book, NH] was dead serious in an academic sense. He loaded it with scholarly linguistic and anthropological reports of his findings in the Amazon. He left academics blinking . . . and nonacademics with eyes wide open, staring.

Yup, “loaded” with anthro and ling stuff that blinds professionals and leaves neophytes agog. Talk of scholarship. Who could ask for more? Not me. Great stuff.

Here’s one more, where Wolfe contrasts Chomsky and Everett (31):

Look at him! Everett was everything Chomsky wasn’t: a rugged outdoorsman, a hard rider with a thatchy reddish beard and a head of thick thatchy reddish hair. He could have passed for a ranch hand or a West Virginia gas driller.

Methodist son of a cowboy rather than the son of Russian Askenazic Jews infatuated with political “ideas long since dried up and irrelevant,” products “perhaps” of a shtetl mentality (29). Chomsky is an indoor linguist “relieved not to go into the not-so-great outdoors,” desk bound “looking at learned journals with cramped type” (27) and who never left the computer, much less the building” (31). Chomsky is someone “very high, in an armchair, in an air conditioned office, spic and span” (36), one of those intellectuals with “radiation-bluish computer screen pallors and faux-manly open shirts” (31) never deigning to muddy himself with the “muck of life down below” (36). His linguistic “hegemony” (37) is “so supreme” that other linguists are “reduced to filling in gaps and supplying footnotes” (27).

Wowser. It may not have escaped your notice that this colorful contrast has an unsavory smell. I doubt that its dog whistle overtones were inaudible to Wolfe. The scholarly blue-pallored desk bound bookish high and mighty (Ashkenazi) Chomsky versus the outdoorsy (Methodist) man of the people and the soil and the wilderness Everett. The old world shtetl mentality brought down by a (lapsed) evangelical Methodist (32). Trump’s influence seems to extend to Harper’s. Disgusting.

That’s it for me. Harper’s should be ashamed of itself. This is not just junk. It is garbage. The stuff I quoted is just a sampling of the piece’s color. It is deeply ignorant and very nasty, with a nastiness that borders on the obscene. Your friends will read this and ask you about it. Be prepared.




[1] Actually, Greenberg’s own Universals were properties of languages not Gs. More exactly, they describe surface properties of strings within languages. As recursion is in the first instance a property of systems of rules and only secondarily a property of strings in a language, I am here extending the notion Greenberg Universal to apply to properties all Gs share rather than all languages (i.e. surface products of Gs) share.
[2] Incidentally, Wolfe does not address these counterarguments. Instead he suggests that NPR are Chomsky’s pawns who blindly attack anyone who exposes Chomsky’s fallacies (see p.35).  However, reading Wolfe’s piece indicates that the real reason he does not deal with NPT’s substantive criticisms is that he cannot. He doesn’t know anything so he must ignore the substantive issues and engage in ad hominem attacks. Wolfe has not written a piece of popular science or even intellectual history for the simple reason that he does not appear to have the competence required to do so.
[3] It is worth pointing out that sentence recursion is just one example of recursion. So, Gs that repeatedly embed DPs within DPs or VPs within VPs are just as resursive as those that embed clauses within clauses.
[4] See Wolfe’s discussion of the “law” of recursion on 30-31. It is worth noting that Wolfe seems to think that “discovering” recursion was a big deal. But if it was Chomsky was not its discoverer, as his discussion of Cartesian precursors demonstrates. Recursion follows trivially from the fact of linguistic creativity. The implications of the fact that humans can and do acquire recursive Gs are significant. The fact itself is a pretty trivial observation.