Comments

Showing posts with label Michael Tomasello. Show all posts
Showing posts with label Michael Tomasello. Show all posts

Wednesday, May 2, 2018

Mendivil-Giro cleans the Augean stables

I am delighted to be writing this very short post advertising a very nice paper. It has appeared in the Journal of Linguisticsbut is available on lingbuzz (here). The paper is aptly entitled Is Universal Grammar ready for retirement? A short review of a longstanding misinterpretation. The author is Jose-Luis Mendivil-Giro (MG) (put in appropriate diacritics on the vowels). Here is the abstract:

In this paper I consider recent studies that deny the existence of Universal Grammar (UG), and I show how the concept of UG that is attacked in these works is quite different from Chomsky’s, and thus that such criticisms are not valid. My principal focus is on the notions of “linguistic specificity” and of “innateness”, and I conclude that, since the controversy about UG is based on misinterpretations, it is rendered sterile and thus does unnecessary harm to linguistic science. I also address the underlying reasons for these misunderstandings and suggest that, once they have been clarified, there is much scope for complementary approaches that embrace different research traditions within current theoretical linguistics.

The paper reads quickly and is surprisingly judicious and generous without being conciliatory.  Readers will note that I have made similar points far less charitably in FoL. MG surmises that the reason for the multiple confusions he identifies lies in the perfectly reasonable fact that different people are (or can be) interested in different issues relating to the wide ranging concept of ‘language.’ Perhaps. There are indeed different people interested in different things and given the complexity of the phenomena we categorize under the term ‘language.’ Further, MG is right to think that these different approaches are complementary rather than incompatible. FoL has made exactly this point several times. However, I believe that MG is being far too generous with GGs critics. I doubt that MG has correctly identified the source of the confused discussion in the literature. And one reason I believe this is that MG’s point has been made repeatedly over the last 60 years to absolutely no avail. GGers have generally bent over backwards conceding that there is room for non-GG style work in investigating the myriad properties that language knowledge and use have. What GG has insisted upon is that it’s own style of work addresses real questions and provides legitimate answers to these questions. Critics have repeatedly rejected this, as MG’s own excellent review of the literature amply demonstrates. So, if there is a confusion (or “misinterpretations”), it is rabid, and not traceable to mere differences to tastes in scientific questions. It has deeper roots. 

Ok, let me say it: the difference really lies in two incompatible conceptions of what science consists in, especially as regards the mental/behavioral sciences. The Empiricism/Rationalism (E/R) divide is the one that I have in mind, but as I have discussed it endlessly on FoL I will not go over it again here. Suffice it to say, that ifone is an Eist then GG is basically muddleheaded confusion. It cannotbe right and so its results need notbe considered. Consequently, if GG’s critics were largely Eish, it would explain the depth of their misunderstanding and their congenital inability to resist confusion/misinterpretation. 

Here’s what I mean. The tenor of many (most?) of the critiques as MG notes hardly ever go into any detail concerning specific GG proposals. As MG notes this results in critiques that are overwhelmingly dumb. The sheer ignorance of the critical discussion is wondrous to behold. The critics that MG cites and discusses really appear to know nothing at all and many (most?) completely ignore everything that GG has discovered over 60 years of research. MG notes this, and seems a bit disoriented by the fact that the main culprits seem so blithely uninformed. And it is not just one or two. They are alllike this, from Chater and Christiensen to Tomasello, Everettt, Levinson etc. etc. etc. Their critiques are really useless (and many times based on simple equivocation (I am talking to you Everett!), even if they contain a grain of truth or two (though color me very skeptical, I have been told that Tomasello’s stuff has someinteresting points) that are worth preserving given a reasonable conception of the enterprise. These kinds of “misunderstanding” are best explained methodologically. The critics don’t go into the details because they don’t believe the problem is one of detail. It is one of principle. The GG enterprise is faulty because its Rish presuppositions are untenable. If you believe this (and these people do, really!), then it is no wonder that they don’t do a deep dive into the details and confront what GGers take to be their most significant contributions.

 In other words, for the critics, the problem is the GG belief that a reasonable view of language would root the research program in an Rish vision of science in general and the mental/behavioral sciences in particular. The critics, being Eish, reject this, and as the divide between E and R conceptions is wide, we can identify its basic unbridgeability as the underlying source of the shockingly shoddy criticisms that MG so ably surveys. Given this, I am far less hopeful than MG is that “there is a glimmer of hope” (p. 23) that these disagreements will be resolved in a rational manner.[1]They cannot be for the very idea of what is the right form of “rational” inquiry is what is being debated.

I have other quibbles with the paper. For example, I found the discussion of reduction and emergence in section 3 somewhat confusing in that it mixes up two different questions: how do linguistic claims get cashed out in wetware? and are linguistic primitives reducible to those of other cognitive domains? These are different questions (as I am sure MG knows) but the paper seems to run them together. The question of FL’s linguistic “specificity” relates more to the second than the first. Of course, if we assume that cognition supervenes on brains and brains are made up of regular biological material then linguistic objects, dependencies and principles even if very linguistically sui generiswill live in biological tissue of these brains. Where else?[2]

However, that is not, nor has it ever been the relevant issue. The question has always been whether the FoL is cognitively independent. To put this crudely in “program” talk: is the FoL program just cobbled together from routines extant in other domains of animal cognition or does it require its own specific features (primitives, subroutines, addressing mechanisms etc.). One might imagine that FL is a kind of Rube Goldberg device assembled from bits and pieces of other available cognitive faculties. This is a possibility. However, I personally doubt it, and the Merge Hypothesis (i.e. that Merge is the linguistically specific sauce that one needs to add to general cognitive and computational powers to yield FL) does as well, though it limits the specificity to this one small operation. 

Honesty compels us (me!) to admit that, to date, no minimalist account has managed to eliminate all operations rather than Merge in accounting for well established features of FL. So, to date, there is reason to think that there is more to the UG parts of FL than just Merge.[3]So whereas the Minimalist Program’s ambitions are alive and well, to date, there is still quite a bit of air between the hopes and the results. And to date, there is good reason to think that FL has quite a bit more UG in it than the standard advertising supposes. This is not a serious problem for the program, but it is worth keeping in mind when we advertise the ambitions given that the program is not exactly in its infancy anymore (it’s a robust 25 years old).

I have other quibbles as well, but enough really. MG has written a terrific paper which makes some very useful points (e.g. I love the discussion in section 4 a lot and his discussion of Tomasello, Everett and Chater and Christiansen are excellent). The paper should be widely read and I hope that it helps change the discussion to a more reasonable one. It shoulddo this. But even if it fails to blunt the overwhelming stupidity of the common critiques, it is a very good paper for insidersto read. I suspect that nowadays many GGers do not really care for the larger cognitive biological issues that once animated the field. This makes it hard to properly rebut the many claims that GG is dead that abound in the popular press. MG’s paper is a good starting point for those interested in reclaiming the cognitive/biological roots of the GG enterprise.

That said I am going to end on a pessimistic note. Despite MG’s excellent discussion, I doubt it will much change the discussion for the reasons outlined above. We are entering a new age of Eism (Deep Learning and Big Data being conspicuous signs of this), and not just in otherareas of cognition.  Its allure is alive in linguistics as well. The idea that FL exists and has special features and that it is a proper object of linguistic study is, IMO, actually taken to be rather quaint within linguisticcircles. GGers with a cognitive bent should not only worry about the barbarians at the gate, the horse has been dragged within the city limits. Let’s hope that MG’s reasonable discussion can redirect this tide, but I am not counting on it.

Last point, I was delighted to see that a major journal published MG’s paper. I could not imagine this appearing in today’s Cognitionor LIor NLLT. Kudos to the Journal of Linguistics.  


[1]Of course, that said, one should always be ready to integrate useful findings from those one disagrees with, even deeply. Those grains of truth are (perhaps) worthwhile.
[2]Though who knows, maybe there really is mind stuff. The belief that there is isn’t is largely a matter of faith.
[3]Indeed, an interesting paradox, IMO, of much contemporary Minimalist work is that it is not Merge that does most of the Grammatical heavy lifting. Rather the prime grammatical operation is AGREE and the long distance feature checking that accompanies it. I-Merge is a very secondary feature of most contemporary accounts and nobody had bothered to consider how linguistically specific the properties of AGREE are. To the degree that they are not, this is a problem for the idea that onlyMerge is linguistically proprietarty. Ditto with the features of the basic lexical atoms. Their idiosyncrasies have been well discussed by Chomsky. To the degree that they remain, there is more to UG than Merge.

Friday, December 5, 2014

The verdict is in regarding Evans' book

Ok, the verdict is in. remember when I wondered (here) if Alun Anderson accurately reported the content of Vyvyan Evans’ book (The Language Myth))? Well he did! It seems that Evans actually made the arguments Anderson attributes to him. How do I know? Norbert must have read the book, you are thinking. Nope. Easier. Evans outlines his views here and it is actually as confused and uninformed as Anderson’s review suggests. In fact, it may even be worse than Anderson’s review suggests.  This paper has no clue. Not even a scintilla of one. The piece is less than unenlightened, it is aphotic. And it is precisely for this reason that I cannot recommend Evans' short paper highly enough. It is a pedagogue's dream. How so?

Well first it is very short. In fact, it’s hard to believe that someone could cram so much misunderstanding into so short a format. Clearly Evans does have talent. It is so replete with the standard confusions that one need send an eager student no further than to this short piece for an example of the kinds of conceptual errors critics of GG seem drawn to (here a nice metaphor of moths and flames seems called for, but I will show strength of literary character and resist). This short piece exhibits every possible mistake. It’s a godsend. Moreover, because it is both short and so replete with misinformation, there is really no need to ever buy the book (so don't!). It is inconceivable that the book could possibly provide more illustrative examples of miscomprehension than is packed into this dense eight pages. So both cheap and serviceable. Perfect.

BTW, those thinking of using this material for teaching purposes (and I am not joking, you really should consider making criticism of these views part of the graduate degree process (even undergrads would benefit)) might also consider using Anderson’s excellent review. It too, as I indicated previously, is thoroughly misguided and misinformed, bit sometimes it’s nice to have two versions of the same song.

As I’ve already (pre)reviewed the Evans' material in the earlier post on Anderson’s reviewed, I will only add a few cursory remarks here:

1.     The confusion between Greenberg and Chomsky Universals is evident throughout. It’s amazing how hard this simple distinction has been for critics to grasp. But this piece is not only conceptually myopic, it is also intellectually lazy. Here’s a piece of critical etiquette: If one wants to sink the GG ship one needs to go after the trophy cases. There are many of these in GG and they are not hard to find. However the observation, yet again, that languages really really look different is not one of these. We know. We’ve always known. We even have GGers who have worked on (gasp!) Salish. So when making a case, discuss the serious work. Failure to do so is a sure fire indicator that nothing of interest has been discussed.

2.     There is a difference between two oft used metaphors for describing innate structure: (i) UG is a blueprint, (ii) UG is an operating system. The difference between the two caused quite a bit of confusion within biology almost as much as it is doing in linguistics (or so I remember Gould telling us in his discussion of the homunculus problem). Blueprints invite the search for Greenberg Universals (every house has 2 bathrooms, 2 bedrooms, a kitchen and a foyer: find them!). Operating systems not so much. Windows is not OS, though they can do many of the same things, they are different systems with different formal requirements. In particular, there are few surface indications of what operating system you are using. In contrast, it is actually quite easy to tell whether some set of blueprints are those of the apartment/house/building you are looking at (I speak from some personal experience here). The GG conception of UG is more like operating system than blueprint. But really we don’t need these metaphors anymore (or at least we should not rest content with them): UG is a function that given PLD yields Gs. Your job, Mr Phelps, should you agree to take it, is to describe the fine structure of this function. In so doing GGers have discovered some very subtle features of this function and critics, to be serious, must address them in detail. Though I am a big fan of hand-waving (especially on sultry summer days) it generally fails to advance the topics of interest. Evans piece is an excellent illustration of this truism.

3.     All the same non-Chomskys make their appearance. Everett is cited (yet again) as showing that Gs need not be recursive and so disconfirming UG and Tomasello is trotted out to represent the idea that a penchant for co-operative behavior is all we need to get human language.  Starlings are schlepped out to argue that one finds recursion in non-humans and Neanderthals make an appearance to testify language facility pre-dating homo sapiens. All of this asserted with nary a hint of skepticism (or apparent knowledge that each one of these claims has been contested and all are likely to be false). Wouldn’t it be nice sometime if someone (e.g. Tomasello, Evans, Anderson) explained to anyone how it is that a penchant for co-operation yielded island effects or the ECP or Principle C, or …you get it. I understand that many other species display quite a bit of co-operative behavior without appearing to be grammatically competent. There is a simple reason why accounts purporting to explain the intricacies of language based on co-operation are never forthcoming, and a moments reflection tells us what this is: there is no route from the former to the latter that does not need to engage with the exactly the questions that Evans begs. And this is evident throughout: just like Anderson, faithful describer of his views, Evans prefers to abstract away from facts and arguments enjoying instead the broad oh-my-god-how-nutty-this-all-looks summary conclusions.

4.     Did I say ‘competent’ above? Here’s another great feature of the Evans’ piece. He seems not to understand the competence-performance distinction as he keeps talking about “speech” rather than linguistic competence. Evidence for the former is not in and of itself evidence for the latter. Birds “speak” but are not syntactically competent. Not even starlings.

5.     Oh yes, the Neo-Darwinian synthesis makes its gradualist appearance here to argue against the emergence of an FL about 100 kya (where is Gould now that we really need him) as does the assertion that being non-localized in the brain (there is no “spot specialized just for language” indicates that the idea that there is a language module is over the top.  I assume that similar reasoning could show us that there is no electrical system in a car. After all it is everywhere.

There is a lot more in this slender cornucopia of misapprehension and I could go on, but you should enjoy this yourself.  I recommend this short prĂ©cis of Evans work to all looking for a paper to hand out to a class that concisely embodies all the mistakes that it is possible to make about the Chomsky program in GG. Though appalling, the piece is very useful.


One last exhortation: linguists should make it their business to loudly criticize this junk at every opportunity. It is getting wide distribution. I got it from Aeon and it was picked up in The Browser. It has just the ingredients to make it big: another one of those Chomsky-has-been-proven-wrong memes that seem so popular. So, criticize this in all venues, especially where non-linguists gather. Consider it part of your linguistic public service.

Monday, April 8, 2013

Give Nim, Give Me



(Photo credit: Herb Terrace)

Einstein was a very late talker. “The soup is too hot”, as the legend has it, were his first words at the very ripe age of three. The boy genius hadn't seen anything worth commenting on.

The credulity of such tales aside, they do contain a kernel of truth: a child doesn’t have to say something, anything, just because he can. This poses a challenge for the study of child language, since naturalistic production is often the only, and certainly the most accessible, data at hand. A child’s linguistic knowledge may not be fully reflected in their speech, which we have known since Lila’s deconstruction of the telegraphic stage. Some expressions may not show up because we haven’t waited long enough, while others—an extraction violation, for instance—will never be said for they are unsayable.

In recent years, what-you-say-is-what-you-know appears to be gaining popularity, as interests in usage based theories of language are on the rise. Here is a warmup. The expression "give me" is proposed as a frozen phrase (Lieven et al. 1992, Tomasello 2010), rather than syntactically composed, spawning cottage industries such as "formulaic languages", which some regard as a transient stage in language evolution (Wray 1998). True, "give" and "me" make a good tag team: "give me freedom", "give me cheese", "give me now" … “gimme coffee” (an old favorite of mine), and they dwarf other combinations.  Take the speech of Adam, Eve and Sarah from Roger Brown's classic study: the frequencies of "give me", "give him", and "give her" are:

95 (93 give me, 2 gimme): 15 (give him): 12 (give her), or 7.91 : 1.23 : 1

So “give me” does seem especially formulaic ... right? Well, not if you check the frequencies of "me", "him", and "her" from the same three kids:

2870 (me) : 466 (him) : 364 (her), or 7.88 : 1.28 : 1

Nothing much can be concluded from these six numbers but there seems to be pretty good support for the null hypothesis that “give” and pronouns combine completely independently. The Brown data has been around for forty years; it's just nobody had bothered to check. (Use the grep, Luke.)

Nowadays everyone does statistics but we still need reasonable hypotheses to test for and against. Usage based theories have plenty of p-values: one can easily show that the frequencies of "give me/him/her" are statistically significantly different from "chance"--but what is "chance"? If we know anything about the statistics of language, it is that language is not "chance" (Zipf 1949). To make the argument against grammar, one would need to show, at the minimum, that the observed distribution in child language is statistically inconsistent with the predicted distribution of a grammar.  Judiciously chosen null hypotheses are needed, not gut feelings: so long to the "gimme" myth.

A few years ago, Virginia Valian came to Penn to give a talk. It concerned the distribution of determiner-noun combinations in child English. Virginia was the first to show that English children’s determiners are virtually error free (1986), thereby providing evidence for an abstract grammar. Not so quick, the usage-based folks say, because the absence of errors could be the result of children memorizing specific word combinations from adult speech, which would also be error free. (I fully endorse such skepticism.) We need some other statistical benchmarks to show the presence of grammar. 

Diversity is a popular measurement. Suppose there is a rule DP→DN, where D is either a or the, and N stands for a singular noun, yielding "a/the car", "the/a pizza", etc. Shouldn't the interchangeability of "a" and "the", per grammar, be reflected in the diversity of nouns that appear with both of them?  Young children's determiner use, however, only shows 20-40% of diversity (Pine & Lieven 1996); perhaps they just memorize determiner-noun combinations from the adult input (Tomasello 2000, Cognition). 

Along with Stephanie Solt and John Stewart, Virginia showed that mothers' speech contains comparable, and comparably low, diversity measures as their toddlers’ (2009, J. Child Language). After her talk, I pulled out some numbers from the Brown Corpus: not Roger Brown, but the collection of English print materials at Brown University, the grandmother of all modern linguistic corpora. Only 25% of singular nouns that combine with either "a" and "the" combine with both. That's lower than some child samples from Pine & Lieven (1996), so two year olds have a better command of English grammar than professional writers. Now that is absurd. 

One reaction would be to abandon the premise that syntactic diversity is a direct reflection of grammatical complexity. Not a bad idea, and much of the purported evidence for usage based theory vanishes. Another reaction would be to go for the extra credit, by characterizing the statistical profile of syntactic diversity that can be expected from a grammar. If the child used 100 distinct nouns, and paired them with either "a" or "the" 500 times, how many of the 100 will be paired with both, assuming the rule DP→DN is at work? Virginia's work was inspirational. I was also knee deep in Zipfian waters, thanks to the work of Erwin Chan, Constantine Lignos and my colleague Mitch Marcus. They showed that pretty much everywhere you look--words, lemmas, morphological inflections, syntactic rules--language follows Zipf-like distributions, which can be exploited for fun and benefit.

If a sample contains 100 nouns (types), then a good many of them must occur only once since they will inevitably fall on Zipf's long and flat tail: these fellows will never get to meet both determiners.  Even for those that do show up multiple times, they may still be monogamous, just as when you toss a fair coin 3 times, it may land on heads 3 times in a roll. And grammar is no fair coin. Nouns tend to have a favored determiner, even though both combinations are possible. For instance, "the bathroom" is more commonly used than "a bathroom" but we say "a bath" a lot more often than "the bath". These imbalances are probably not a matter of grammar, which presumably does not encode the frequency of bodily needs, but they will conspire to produce low syntactic diversity, and thus the impression of grammatical absence.

After a bit of probability theory exercise [1], we can use a formula to calculate the expected diversity from the sample and vocabulary size (e.g., 500 and 100). The key here is multiplication--the statistical hallmark of independence--of the noun probabilities with determiner-noun combination probabilities, both of which can be well approximated by Zipf’s law. I was surprised to see how well it worked, and in fact had to learn new statistics just to be sure. We mostly use statistics to show one set of values and another (e.g., experimental results vs. "chance") are statistically different, but being different is not the same as being the same. Lin's concordance correlation coefficient (there is an R package, of course), first invented in biostatistics to verify drug effectiveness across trials, confirmed the observation.   In other words, children's syntactic diversity appears exactly what one might expect from a grammar rule, once the general statistical properties of language are taken into account. [2]

Someday we may have a bunch of these statistical profilers, like what evolutionary geneticists use to detect natural selection at the molecular level. Let me make very clear what this work does and does not show. It does show that at least one part of child language makes use of an abstract rule of grammar but it does not mean that all parts of child language do. It does show that children can merge but it does not tell us how they learn what to merge with. It does show the presence of grammatical ability in very young children, but it does not say how that ability got there in the first place, ontogetically or phylogenetically.

Which brings me to Nim Chimpsky and the evolution of language. The continuity between primate language and early child language is believed to hold “the most promising guide to what happened in language evolution” (Hurford 2011, p590), presumably on the apparent formulaic similarities between them. If the numbers worked out for children, who seem to have a grammar after all, perhaps Nim is due for a similar upgrade?  Whatever one thinks of Project Nim--I had to fight back tears--it produced the only publicly available corpus of primate language. Nim acquired about 125 signs of ASL, and produced thousands of multiple sign combinations, the vast majority of which being two sign combinations (Terrace 1979, Nim) These have been described as rule-like constructions, each consisting of two closed class functors such as “give” and “more,” along with open class items such as “apple,” “Nim,” or “eat.”  Signs do not combine with uniform frequency either, with “eat”, “banana”, “me”, “Nim" etc. among the predictable favorites. What’s Nim’s syntactic diversity if he combined signs under a rule? Run the numbers: the poor guy didn’t seem to have a grammar, just as his trainers concluded (Terrace et al. 1979).

Syntactic diversity in human language, usage based learning and Nim Chimpsky.



Moral for the day: the null hypothesis, once properly formulated, may come back to bite your statistical hand.  All very exciting. When I explained this work to some of my non-linguist friends (I do have a few!), their reaction was one of surprise, though not the kind I had in mind. “Why would anyone think kids learn language by copying us?  Just this morning, Maggie said ___”, to be filled by one of the darndest things kids say.  They do wonder about vocabulary, boys vs. girls, and bilingualism, but no one is remotely concerned about the combinatorics of grammar that are, literally, screaming in their faces. Perhaps linguists do worry too much. 

[1] Thanks to Ruochuan Liu and Qiuye Zhao for spotting an error early on.
[2] Could a usage cum memory-retrieval model account for the same finding? I don’t think one knows for sure,  since it has been difficult to pin down the mechanics of usage based learning so it’s unclear what quantitative predictions it makes. I won’t dwell on the matter here but refer you to the paper under discussion, where a concrete proposal (Tamales 2000, Cognitive Linguistics, p77) is tested but came up short.