Comments

Showing posts with label Lila Gleitman. Show all posts
Showing posts with label Lila Gleitman. Show all posts

Tuesday, November 18, 2014

Why Morphology? (part deux)

As Chomsky has repeatedly emphasized, natural language (NL) has two distinctive features; it’s hierarchically recursive and it contains a whole bunch of lexical items (LI) with rather distinctive properties (when compared to what one finds in animal communication systems (e.g. they are not stimulus bound and they are very conceptually labile)). A natural question that arises is whether these two features are related to one another; does the fact that NLs are hierarchically recursive (actually generated by Gs that are recursive and produce hierarchical linguistic objects, or the corresponding I-language is such, but let’s forget the niceities for now) have any causal bearing on the fact that NL lexicons are massive or vice versa?[1] The only proposal that I know of that links these two facts together causally is Lila Gleitman’s hypothesis that vocab acquisition leverages syntactic knowledge. The syntactic bootstrapping thesis is that Gs facilitate (and hence accelerate) vocab acquisition in LADs. By “vocab” I mean “open” class contentful items like ‘shoe’ and ‘rutabaga’ rather than “closed” class items like ‘in’ or ‘the’ or the plural ‘-s’ morpheme. NLs have a surprisingly large number of open class “words” when compared to other animal systems (at least 3 orders of magnitude larger) and it is reasonable to ask why this is so.[2]

Gleitman’s syntactic bootstrapping thesis provides a possible answer: without syntactic leverage, acquiring words is a very very arduous process (see here, here, and here for discussion about how first words are acquired), and as only humans have syntax, only they will have large vocabs. Oh, btw, by acquisition I intend something pretty trivial; tagging a concept with a label (generally phonological, but not exclusively so, think ASL). I don’t think that this is all there is to vocab acquisition,[3] but it turns out that even this is surprisingly difficult to accomplish, (contrary to intimation to the contrary in the philo of language literature, see Quine on the ‘museum myth’) and it requires lots of machinery to pull off (e.g. a way of generating labels, namely, something like a phonology and a way of identifying the things that need tagging).

I mention all of this because I just recently heard a lecture by Anne Christophe (see slides here) that goes into some details about the possible mechanics behind this leveraging process that bears on an earlier question I have been wondering about for a long time: why do NLs have so much phonologically overt morphology?  For any English speaker, morphology seems like a terrible idea (it’s a mess and a pain to learn, you should hear my German). However, Christophe argues that morphology and closed class items serve to facilitate the labeling process that drives open class vocab acquisition in humans. In other words, her work (and here I mean that of the work from her lab as there are many contributors here as the slides make clear) sketches the following picture: closed class items (and I including morphological stuff here) provide excellent environments for the identification and tagging of open class expressions. And as closed class items are in fact very closely tied to (correlated with) syntactic structure (think Jabberwocky!!), morphology, broadly construed, is what enables kids to leverage syntax to build vast lexicons. If correct, this forges a close causal connection between the two distinctive properties NLs display; large vocabs and syntactic structure. Let’s call this the Gleitman-Christophe thesis (GCT).

What’s the evidence? Effectively, morphology provides relatively stable landmarks between which content words sit thus allowing for easy identification and tagging. In other words, morphology allows the LAD to zero in on content words by locating them in relatively fixed morpho-syntactic frames. And very young LADs can use this information for they are known to be very good at acquiring this fixed syntactic scaffolding using non-linguistic distributional (and statistical) methods (see slide 18).[4] Closed class items are acquired early (they are ubiquitous, frequent and short) and it has been shown that kids can and do exploit them to “find” content words. Christophe reports on new work that shows how useful this can be for categorizing words into initial semantic classes (e.g. distinguishing Ns (canonically things) from Vs (canonically eventish)).  She then goes on to describe how acquiring a few words can further enhance the process of acquiring yet more vocab. The first few words act as “seeds” that once acquired serve to further leverage the acquisition process. In sum, Christophe describes a process in which prosody (which Christophe discusses in the first part of the talk), morphology and syntax greatly facilitate word acquisition, which in turn enables yet more word acquisition. We thus get a virtuous circle anchored in morpho-syntax, which is in turn anchored in epistemologically prior pre-linguistic capacities to find fixed phono-morpho-syntactic landmarks that allow LADs to quickly fix on new words.[5] This all provides good evidence in favor of the GCT. It also allows one to begin to formulate one answer to my earlier question: Why Morphology?

So what’s morphology for? (This is a dumb question really, for it could be for lots of things. However, thinking functionally is natural for ‘why’ questions. See below for a just so story). Well among other things It is there to support large vocabs and this, I would suggest is a very big deal.  I once suggested the following thought experiment: I/you am/are going to Hungary (I/you speak NO Hungarian) and am/are made the following offer: I/you can have a vocab of 50k words and no syntax whatsoever or a perfect syntax and a vocab of 10 words. Which would I/you prefer? I/(you?) would take door number 1. One can get a long way on 50k words. Nor is this only for purposes of “communication” (though were there an indirect tie between communicative enhancement and morpho-syntax I would be ok with that). Phenomenologically speaking, tagging a concept has an effect not unlike making an implicit assumption explicit. And explicitness is a very good way to enhance thought (indeed, it feels like it allows one to entertain novel thoughts). Having a word for something allows it to be conceptually accessible and salient in a way that having the concept inchoately does not. In fact, I am tempted to say, that having a word for something can change the way you think (i.e. it can affect cognitive competence, not just enhance performance).[6] So, tagging concepts really helps and tagging a lot of concepts really really helps. Thus, if you are thinking of putting in an Amazon order for a syntax, I would suggest asking for one that also supports large scale vocab acquisition (tagging concepts) and the GCT argues that such a syntax would come with phonologically overt morphology and phonologically overt closed class items that our pre-linguistic proclivities can exploit to build large lexicons quickly. In other words, if you want an NL that is good for thinking (and maybe also for communication) get one that has the kinds of relations between morphology and syntax that we actually find.

Note that this view of morphology leaves (completely?) open the question of how morphology functions inside UG. It is consistent with this view that G operations are driven by feature checking requirements, some of which become realized overtly in the phonology (this is characteristic of early Minimalist proposals). It is also consistent with the view that they are not (e.g. that they are mere by-products of grammatical operations rather than drivers thereof (this is what we find in later EPP based conceptions of grammatical operations in later minimalism). It is consistent with the idea that morphology exists to fit syntax to phonology (readjustment rules), or that it’s not (i.e. it’s a functionally useless ornamentation).  All the GCT requires is that there be reliable correlations between overt markers and the syntax so as to allow the LAD to leverage the syntax in order to acquire a content rich lexicon.

If this is right, then it might serve as the start of an account of why there is so much overt morphology and/or closed class items (very frequent items that function to limn syntactic structure) in NLs.  In fact, it suggests that there should be no NL that eschews both, though what the mix needs to be can be quite open. So Chinese and English don’t have much morphology, but they have classifiers (Chinese) or Determiners and Verbal morphology (English) and this can serve to do the lexicon building job (Anne C tells me that all the stuff done on French discussed in the slides replicates in Mandarin).

As an aside, IMO, one of the puzzles of morphology is why some languages seem to have so much (Georgian) and some seem to have so little (English, Chinese). If morphology is that important in the grammar either functionally (e.g. for processing) or grammatically (e.g. it drives operations) then why should some NLs express so much of it overtly and some have almost none at all that is visible. The GCT offers a possible way out: morphology per se is not what we should be looking at. Rather it is morphology plus closed class items; items that give you fixed frames for triangulating on content “words.” There needs to be a sufficiency of these to undergird vocab acquisition, but the mix need not be crucial (in fact it is not even clear how much of both is sufficient or if there may be a cost in having too much. Or even if these queries make any sense).

Let me end here. NLs are stuffed with what appears to be “useless” (and as an English speaker, cumbersome) morphology. And useless it may be from a purely grammatical point of view (note I say may, leaving the question open). But GCT suggests that overt grammar dependent fixed points can be very useful for building lexicons given our pre-linguistic capacities. And given the virtues of a good sized lexicon for thinking and communicating, a syntax that can support this should have advantages over one that doesn’t (that’s the just so story, btw). If correct, and the data collected so far is non-trivial, this is nice to know for it serves to possibly bridge two big facts about NLs, facts that to date seem (or more accurately, seemed to me) to be entirely independent of one another.





[1] Nothing I say touches on the fact that NL lexical items have very distinctive properties, at least when compared to symbols in animal systems. In other words, the lexicon presents two puzzles: (i) why is it so big? (ii) why do human lexical items function so differently from non-human ones. What follows tries to say something about the first, but leaves the second untouched. Chomsky has discussed some of these distinctive features and Paul Pietroski has a forthcoming book that discusses the topic of lexicalization. 
[2] So far as I know, no animal communication system has any closed class expressions, which, if they have no syntax, would not be a surprise given Gleitman’s conjecture.
[3] Paul Pietroski has a terrific forthcoming book on what lexicalization consists in “semantically” and I intend my views on the matter to closely track his. However, for present purposes, we can ignore the semantic side of the word acquisition process and concentrate solely on the process of phonetically tagging LIs.
[4] That kids are really good at identifying morphology has always surprised me. This is far less the case in second language acquisition if my experience is anything to go on. At any rate, it seems that kids rarely make “errors” of commission morphologically. If they screw up, which is surprisingly infrequently, it manifests as errors of omission. Karin Stromswold has work from a while ago documenting this.
[5] This might also provide a model for how to think of the thematic hierarchy.  Is this part of UG or not? Well, one reason for thinking not is that it is very hard to define theta roles so that they apply across a broad class of verbs. Dowty showed how hard it is to define ‘agent’ and ‘patient’ etc. and Grimshaw noted that it is largely irrelevant for syntactic concerns anyhow. When does it matter? Real, theta roles matter for UTAH. We need to know where a DP starts its derivational life.  If this is its sole role, then what one needs theta roles for is not semantic interpretation in general but for priming the syntax.  In this case, all one needs are a few thematically well behaved verbs (like ‘eat,’ ‘hug,’ ‘hit,’ ‘kiss,’) and to get the syntax off the ground. Once flying, we don’t need thematic information any more for there arise other ways, some of them being morpho-syntactic, to figure our where a DP came from (think case, or fixed word order position or agreement patterns). At any rate, like that morphology case, the thematic hierarchy need not be part of UG to be linguistically very important to the LAD.
[6] Think of this as a very weak (almost truistic) version of the Sapir-Whorf hypothesis.

Monday, March 11, 2013

Learning Fast and Slow I: How Children Learn Words



My inaugural post is the first of a three part sequence on language acquisition, which has become something of a signature issue for this blog. The title, and indeed the theme, are taken from Kahneman's recent book (Thinking, Fast and Slow). The analogy is appropriate. Anyone who has looked at child language--especially when quantitatively and cross linguistically--can see that some parts are learned fast while others proceed at a much more gradual pace. Depending on the department one works in, the fast and the slow side tend to get exaggerated to the point of mutual exclusion; I will name names, I promise.

These posts will be based on some strands of my own work--sorry but I can only talk about things I know. In the later parts of this sequence, I will deal with aspects of child language that decidedly call for either fast or slow learning, but I will kick things off with a case that requires a bit of both: How Children Learn Words.

As Noam and Lila reminded us over the years, word meanings are very complicated and associationist learning is hopeless.  A recent anecdote. My daughter goes to a wonderful Montessori school and one of the things they do is called "Word of the Week", when a grownup word is illustrated with examples from pictures, stories, and group activities. So for a long time, our three year old thought "loyalty" meant hugging, "drizzle" meant watering seeds for them to grow, and "coincidence" meant two (specific) girls wearing identical polkadot leggings. Eventually she learned these words; I just have no  idea how. Perhaps Jerry Fodor is right after all but I digress. 

The task at hand is much simpler. In fact, "How Children Learn Words" is false advertising because research in this business is really "How children learn words that go with medium sized bright objects". But we are aiming lower still: "How children learn which words go with which medium sized bright objects." 

Not very exciting for those who worry about type shifting and generalized quantifiers. Yet a sizable cognitive science literature deals is devoted to this very topic, and identifying what goes with what is clearly a necessary component in any theory of word learning. The revival of associationist learning, I suppose, comes from the realization that statistical learning is "more powerful than previously thought". Surely not all words, or every instance of them, will be neatly aligned with their "meanings"--really, and very very loosely, “referents”. But if the target meaning is associated with a word sufficiently frequently, and more frequently than its competitors, then the learner may be able to detect and learn from such statistical correlations. Again, scientists now know we are really good at statistics. The research on cross-situational learning is to see if the learner can tabulate word-meaning co-occurrence statistics over multiple learning instances to figure out which words pick out which objects. 

Apparently they can.  Consider a typical cross-situational learning study adapted from Smith & Yu Smith (2008) where the subject hears the word "ball" accompanied by two learning instances:

The cross situational learner may construct a mental table that lists the frequencies with which the objects (BALL, BAT, DOG, following Jerry’s CARBURATOR convention) are paired with "ball". It is clear that BALL will be the winner since it co-occurs with “ball” more consistently than others.  Several computational models, starting with Jeff Siskind's MIT Thesis in the 1990s, have implemented variants of this idea: some are quite straightforward implementations of cross tabulation (Yu & Ballard 2007, Fazley et al. 2010) while others situate learning in a Bayesian framework that enumerates and evaluates all possible lexicons (Frank et al. 2009). 

A few years ago, Jon Stevens, a graduate student at Penn, implemented the variational learning model (Yang 2002; a subspecies of Reinforcement Learning to be featured in later posts) and tested it on some manually annotated video data from CHILDES. Each utterance is treated as two sets, one for a bag of words (i.e., no linguistic relation whatsoever) and the other for a bunch of reasonably observable/identifiable "things" (i.e., medium sized objects) For each word, the learner establishes a probabilistic distribution overs its candidate meanings. When a word is heard, the learner chooses a candidate meaning according to its probability and checks if it's present in the set of things. If so, the associated probability goes up (reward); otherwise, it goes down (penalty). This model worked well enough but was quite a bit worse than the far more complex Bayesian model (Frank et al. 2009). The computational linguist in me knew that any model, however simple, would not get in the game as long as some other model, however complex, was producing higher F-scores. 

The variational model for word learning turns out to be too slow and indecisive, hedging its bets not unlike a cross situational learner. In a series of clever and illuminating experiments, my colleagues Lila and John have uncovered some limits of cross situational learning.  Here is a briefest summary since the topic has come up before. Not only are subjects incapable of tracking cross situational statistics, they appear to follow a strategy dubbed Propose-but-Verify. If a candidate meaning is confirmed, the learner keeps it; if it is disconfirmed, the learner replaces it with a new candidate meaning. Crucially, the learner appears to assess only one candidate meaning at a time: if a meaning is confirmed, the learner ignores everything else in the scene completely. 

Learnability aficionados will recognize that Propose-but-Verify is the triggering model (Gibson & Wexler 1994) in disguise: for "a candidate meaning", read "a candidate grammar". If a grammar (e.g., parameter setting) works well for an input sentence, it is kept; otherwise the learner tries a new setting and moves on. (This is hypothesis testing par excellence, and I never understood why folks would characterize the triggering algorithm as an instance of domain specific learning.) But triggering is too fast for its own good. A tried and true parameter setting can be overturned on the basis of one counter example, possibly a speech error, which brings about serious learnability problems (Berwick & Niyogi 1996).  

Lila and John’s experiments have led to a new model of word learning, perhaps the simplest model possible. It combines fast and single minded hypothesis testing with slow and conservative probabilistic learning from the variational model. We call it Pursuit. The learner always tests, or pursues, the most probable word meaning, hence only one hypothesis at any time: the variational model, by contrast, samples among the candidates probabilistically--and that's the key difference. Just as before, if a meaning (the most probable) is confirmed, its probability goes up and if it is disconfirmed, its probability goes down and the learner adds another candidate from the current scene (with some initial probability). Next time when the word is heard, the learner still goes for the most probable candidate, which may well be a failed candidate from before but has not slipped down the pecking order.  In other words, failed candidates get to hang around, depending on how dominant they were before failing, rather than being booted out altogether.

The Pursuit model is decidedly non-optimal. Take a look at the BALL-BAT-DOG example. The cross situational models always learn BALL. The Pursuit model only gets it 75% of the time. If you guess BALL in the 1st instance, it will be pursued and succeeds in the 2nd; all good. But if you guess BAT in the 1st instance, which will be disconfirmed in the 2nd, you only have a 50% chance of guessing BALL.  Yet surprisingly--no B.S. Norbert, we really weresurprised--the Pursuit model came out Best in Show, in a matter of seconds rather than hours as the Bayesian model requires.  Obligatory performance table below; for details, see The Pursuit of Word Meanings.


Model
Precision
        Recall                
F-score
cross-situational                 
0.28
        0.21
0.24
Propose-but-Verify
0.04
        0.31
0.08
Bayesian
0.50   
        0.29
0.37
Pursuit
0.44
        0.38
0.41


Here is why. While mommy may not “show them the thing whereof they would have them have the idea;  and then repeat to them the name that stands  for it, as “white”, “sweet”, “milk”, “sugar”, “cat”, “dog””, as John Locke speculated, they never set up cross situational learning conditions where meanings are uniformly spread out across words and learning instances. If one looks at word learning data children receive, one cannot fail to note that words are indeed embedded in multiply ambiguous settings. Yet every now and then, mommy does point to “cake” and say “cake”, and there is a ball around, and bright orange bouncy one at that, when the child hears “ball!”. The Pursuit model latches on to these low ambiguity learning instances, even though they are few and far between, and pursues them with abandon.  These relatively salient cues are diluted under cross situational learning, which averages everything including many (more) highly noisy ones. The artificial world of cross situational learning does favor the cross situational model but bears little relevance to reality, which is indeed one of the main points in Lila and John’s papers.

"Shoe."

A very small step toward word learning, to be sure. But even a small step needs to be taken deliberately, neither too fast nor too slow.




Wednesday, December 12, 2012

Does Anyone Ever Learn Anything?


Let’s follow Jimmy and Judy from birth to about five.  At birth they say precious little. At five you can’t shut them up. What happened in these five years? Answer: they learned their native language, English say. Obvious no?  Yes. But is it right? Did Judy and Jimmy learn English?  Well, to paraphrase a recent political celebrity, it all depends on what you mean by learn and English.  It seems undeniable that Judy and Jimmy developed a capacity absent (or at least invisible[1]) at birth and this capacity can be exercised to converse with some natives, the English speaking ones, but not others, the Mandarin speaking ones. However, does this imply that they learned English? 

Linguists have long understood that labels like English, French, Swahili, Mandarin, etc. are more convenience terms than terms of art (here’s where linguists mention the Weinreich quip that a language is just a dialect with an army and a navy; real cognoscenti adding sotto voce that dialects are just idiolects with epaulettes). However, recent research suggests that we have been far too cavalier about the first half of this doublet.  We can all agree that Judy and Jimmy acquired (competence in) English, but did they learn English? Recent (and, as we shall see, not so recent) research into word learning suggest that here we need to look before we leap, something, it appears, that kids do not do, at least when it comes to early word acquisition. I’ve already discussed some of the research by MSTG (here) which argues (very convincingly in my view) that the early stages of lexical acquisition do not involve the careful statistical weighing of competing alternatives but involves jumping to a conclusion mentally clutched with fierce determination and, with time, forgotten if incorrect, only to set up another ill supported jump into the lexical abyss (boy was that fun to write!). MSTG support this conclusion by considering learning situations less factitious than the contrived set-ups near and dear to the psych lab.  When kid and adults are asked to consider more natural (and hence visually busy) filmed vignettes in which the targets of lexical labeling are not clearly segregated and identified on pristine picture cards, they acquire word meanings all at once or not at all. This result is important for several reasons.

First and foremost, it provides a concrete illustration of why we should not equate acquisition with learning.  MSTG provides diagnostics of learning (multiple hypotheses, statistical weighing of these alternatives, gradual convergence on the right result) and argues that learning so understood fails to hold in more realistic contexts of lexical acquisition. Specific conclusion; in at least one demonstrable case, acquisition does not equal learning.  More general conclusion; it is an empirical question whether in any given acquisition context it is true that learning (now understood to be one mechanism among others for the acquisition of knowledge) is taking place.  In other words, Jimmy and Judy certainly acquired English but whether they learned it is entirely open for empirical grabs. Chomsky’s repeated suggestion that we understand language acquisition as a species of growth rather than learning makes an analogous point.  MSTG makes the case more crisply, I believe, by providing clear diagnostics of learning and showing that there are core cases of “learning” where these signature properties of learning are demonstrably absent.

Second, MSTG provide a rationale for why learning doesn’t hold in their examined cases. Commenting on this (here) I observed that the MSTG results suggest that in such busy contexts the prerequisites for “cross situational learning” do not exist and this is why an alternative acquisition strategy is employed. Following MSTG’s lead, I even suggested that learning requires the kind of structured hypothesis space provided by the contrived set ups that MSTG’s more realistic vignettes argue against. This constituted a kind of compromise position; learning applies where acquirers have well structured hypothesis spaces and leaping to conclusions holds where this fails to hold.[2]  However, (and learn from this you soft-hearted open-minded intellectual compromisers out there) once again Norbert’s natural generosity of spirit and desire for intellectual group-hug kumbaya moments led him astray. It seems that even this compromise position concedes too much to learning mechanisms. In a companion paper, Trueswell, Medina, Hafri and Gleitman (TMHG) extend the MSTG results to include the more stylized learning environments in which options are clearly marked and lexical targets (aka referents) are crisply identified.[3]  Even in the artificial setting of the experimental psych lab kids and adults do the darndest things!

More specifically, like MSTG, TMGH identify the quiddity of “cross situational learning” with the following mechanism:

…keeping track of multiple hypotheses about a word’s meaning across successive learning instances, and gradually converg[ing] on the correct meaning via an intersective statistical process (128).

What they demonstrate is that even in simple stylized psych lab contexts when acquirers “are placed in the initial novice state of identifying word-to-referent mappings across learning instances, evidence for such a multiple hypothesis tracking procedure is strikingly absent (128),” and learners don’t “track the cross-trial statistics in order to build the correct mappings between words and referents (129).”  What acquirers do is “make a single conjecture upon hearing the word used in context and carry that conjecture forward to be evaluated for consistency with the next observed context (129).”  If confirmed, the guess is retained, if disconfirmed, acquirers guess again as if de novo.

TMGH show this (same with MSTG) by considering the dynamics of knowledge fixation: “how learning accuracy unfolded across learning instances (130).” This is very interesting stuff and I strongly recommend that you look at the details. However, just as interesting is the little bit of history that TMGH review. TMGH acknowledge tapping into a long history of criticism of this traditional gradual/comparative conception of learning.  In the late 1950s and early 1960s first Irvin Rock (1957) and then William Estes (1960) used very analogous kinds of arguments to demonstrate that verbal learning was not learning at all but was one-trial guessing. They did not fare so well. As Roediger and Arnold (RA) (2012) put it “the verbal learning establishment rose up to smite down these new ideas (2).”  It is instructive to read the RA paper for it shows the weakness of the counter-arguments used against Rock and Estes that nonetheless carried the day.  Just like today, learning was less a hypothesized mechanism for acquisition than a definitional truth about it.[4] 

The TMHG paper ends with an interesting paradox that I would like to briefly discuss as well. Initial word learning is slow and laborious (35-140 words at 14 months) but unbelievably rapid thereafter (12,000 words by age 6, i.e. roughly 7 words per day for a little under 5 years).  Why the change? TMHG note other work (Lila’s syntactic bootstrapping hypothesis) which proposes that “the acquisition of syntax and other linguistic knowledge by toddlers and young preschoolers during this time period provides a rich database of additional constraints that permit the learning of many additional words.” It is conceivable that with this knowledge in place, cross situational learning might finally become operative, though as TMHG correctly observe, this is decidedly an empirical question and it is “plausible that a propose-but-verify word learning procedure is at work all along the course of word learning throughout most of the life cycle (150).”  It would be interesting in the extreme, in my view were either conclusion correct.

If the latter conclusion proved true (propose and verify all the way down), then language acquisition might have nothing to do with learning in the technical sense. Of course, this is consistent with grammar acquisition being a case of learning, i.e. perhaps we don’t learn words but we do learn parameters. Maybe, but I’d be very skeptical. If even word acquisition isn’t an instance of learning then it seems to me that the burden of proof that any area of language acquisition involves learning would be pretty high.

If, on the other hand, the former option is correct, (viz. that learning only kicks in when grammar is there to buttress it) then its role in accounting for language acquisition would, in my opinion, be quite modest.  Yes, learning plays a role, but really most of the action lies with the constraints that grammars place on the process. This does not mean that we should not study the extras that learning might be adding (though remember the possibility mooted in the prior paragraph), but I doubt that these cognitive titivations will generate much excitement if they only operate in restricted hypothesis spaces. Learning is interesting when there are lots of options that need sifting, not so much when the range of possible end states is highly restricted. 

MSTG and TMHG show us that terminology matters. If we call acquisition “learning” we’ve loaded the research dice. If MSTG and TMHG are right (and I’d bet quite a bit that they are: any suckers out there?) it looks like we’ve repeatedly loaded them to come up snake eyes. It’s time we cut our losses and open our minds to the possibility that ‘language learning’ like ‘Justice Roberts,’ ‘military intelligence,’ and ‘western civilization’ is an oxymoron.



[1] I include this hedge for every week another (usually French speaking) psychologist shows that the youngest kids seem to have the most prodigious knowledge. It seems that we have nothing to teach those little know-it-alls. 
[2] A kind of learning as first resort, jumping to conclusions a last resort, method. C.f. TMHG p. 130.
[3] It is doubtful that the meaning of a lexical item is simply its referent for reasons that Chomsky has belabored (sadly, quite unsuccessfully) over the years. There is far more structure to lexical items than what they refer to. However, for current purposes, this only further dramatizes the inadequacy of learning as a mechanism for lexical acquisition. 
[4] TMHG also cite another paper by Gallistel and friends that I have not yet read but will try to get hold of and blog about when I do. It argues (cited in TMHG 151), that “in most subjects, in most paradigms, the transition from a low level of responding to an asymptotic level is abrupt.” Oh my. I’ll keep you posted.