Comments

Showing posts with label syntactic bootstrapping. Show all posts
Showing posts with label syntactic bootstrapping. Show all posts

Tuesday, December 18, 2012

Lila on Learning (and Norbert gets chastised)


Dear Norbert,

I hesitate to say ANYTHING in response to your last post because its positive tone leaves me glowing, wonderful if everyone thinks the same about our findings and our new work.   But there are some foundational issues where you and I diverge and they're worth a good vetting, I think.   I really was totally surprised at your equating slow/probabilistic/associationistic learning with "learning in general," i.e. that if some change in an organism's internal state of knowledge wasn't attributable to the action of this very mechanism, then it isn't learning at all, but "something else."   My own view of what makes something learning is much more prosaic and excludes the identification of the mechanism:  if you come to know something via information acquired from the outside, it's learning.   So acquiring knowledge of Mickey Mantle's batting average in 1950 or finding out that the pronunciation of the concept 'dog' in English is "dog" both count as learning, to me.   At the opposite extreme of course are changes due entirely, so far as we know, to internal forces, e.g., growing hair on your face or replacing your down with feathers (though let's not quibble, some minimal external conditions have to be met, even there).   The in-between cases are the ones to which Plato drew our attention, and maybe Noam is most responsible in the modern era for reviving:

“But if the knowledge which we acquired before birth was lost by us at birth, and afterwards by the use of the senses we recovered that which we previously knew, will not that which we call learning be a process of recovering our knowledge, and may not this be rightly termed recollection?”
(Plato Phaedo [ca. 412 BCE])

I take UG to be a case, something (postulated) as an internal state, probably a logically necessary one, that is implicit in the organism before knowledge (say, information about English or Urdu) is obtained from the outside, and which guides and helps organize acquisition of that language-specific knowledge.  I have always believed that most language acquisition is consequent on and derivative from this structured, preprogrammed, basis, yet something crucially comes from the outside and yields knowledge of English, above and beyond the framework principles and functions of UG.   Syntactic bootstrapping, for example, is meant to be a theory describing how knowledge of word meaning is acquired within (and "because of") certain pre-existing representational principles, for example that -- globally speaking -- "things" surface as NP's (further divided into the animate and inanimate) and "events/states" as clauses, so the structure NP gorps that S would be licensed for "gorp" iff its semantics is to express a relation between a sentient being and an event, e.g., "knowing" but not "jumping."  

The problems we have in mind to address in the recent work you've been discussing, to my delight, are: how do you ever discover where the NP's etc. are, in English?  That its subjects precede its verbs, roughly speaking.   This knowledge comes from outside (it is not true of all languages) and has to be acquired to concretize the domain-specific procedure in which you learn word meanings, in part at least, by examining the structures for which they're licensed. I have argued that, at earliest stages, you can't make contact with this preexisting knowledge just because you don't know, e.g. where the subject of the sentence is.   To find out, you have to learn a few "seed words" via a domain-general procedure (available across all the species of animals we know perhaps barring the paramecia) .  That procedure has almost always (since Hume, anyhow) been conceived as association (in its technical sense).   As I keep mentioning, success in this procedure is vanishingly rare, it is horrible, because of the complexity and variability of the external world, though reined in to some extent by some (domain-specific) perceptual-conceptual biases (see Markman and Wachtel, inter alia).  Apparently, you can only acquire a pitiful set of whole-object concrete concept labels by this technique.   Though it is so restrictive, we take it as crucial:  it is the only possibility that keeps "syntactic bootstrapping" from being absolutely circular, it provides the enabling data for SB, enough nouns to hint as to where the subject is, hence given this knowledge of "man" and "flower" and the observation of a man watering a flower, you learn not only the verb "water" but the fact that English is SVO.  

So back to the point:  I think word learning starts with a domain-general procedure that acquires "dog" by its cooccurrence with dog-sightings, given innate knowledge of the concept ‘dog,’ and one learns (yes) that English is SVO as a second step, and as above.  This early procedure gives you concretized PS representations, domain-specific, language-specific ones, that allow you to infer "a verb with mental content" from appearance in certain clausal environments.   That's my story.  

What my argument with you, now, is about, is the contrapositive to Plato, I am asking:   "If you have to have information from the outside, information as to the licensing conditions for "dog" (and, ultimately, "bark"), is acquiring that information not rightly termed learning?"    I think it is.   But the mechanism turns out (if we're right) to be more like brute-force triggering than by probabilistic compare-and-contrast across instances.    The exciting thing, as you mention, is that others through the last century (e.g., Rock, Bower, several others) insisted that learning in quite different areas might work that way too.   Most exciting I think is the work starting in the 1940's and continued in the exquisite experimental work of Randy Gallistel, showing that it is probably true of the wasps, the birds and the bees, the rats, as well, even learning stupid things in the laboratory (well, not Mickey Mantle's batting average, but, at least, where the food is hidden in the maze).

Wednesday, December 12, 2012

Does Anyone Ever Learn Anything?


Let’s follow Jimmy and Judy from birth to about five.  At birth they say precious little. At five you can’t shut them up. What happened in these five years? Answer: they learned their native language, English say. Obvious no?  Yes. But is it right? Did Judy and Jimmy learn English?  Well, to paraphrase a recent political celebrity, it all depends on what you mean by learn and English.  It seems undeniable that Judy and Jimmy developed a capacity absent (or at least invisible[1]) at birth and this capacity can be exercised to converse with some natives, the English speaking ones, but not others, the Mandarin speaking ones. However, does this imply that they learned English? 

Linguists have long understood that labels like English, French, Swahili, Mandarin, etc. are more convenience terms than terms of art (here’s where linguists mention the Weinreich quip that a language is just a dialect with an army and a navy; real cognoscenti adding sotto voce that dialects are just idiolects with epaulettes). However, recent research suggests that we have been far too cavalier about the first half of this doublet.  We can all agree that Judy and Jimmy acquired (competence in) English, but did they learn English? Recent (and, as we shall see, not so recent) research into word learning suggest that here we need to look before we leap, something, it appears, that kids do not do, at least when it comes to early word acquisition. I’ve already discussed some of the research by MSTG (here) which argues (very convincingly in my view) that the early stages of lexical acquisition do not involve the careful statistical weighing of competing alternatives but involves jumping to a conclusion mentally clutched with fierce determination and, with time, forgotten if incorrect, only to set up another ill supported jump into the lexical abyss (boy was that fun to write!). MSTG support this conclusion by considering learning situations less factitious than the contrived set-ups near and dear to the psych lab.  When kid and adults are asked to consider more natural (and hence visually busy) filmed vignettes in which the targets of lexical labeling are not clearly segregated and identified on pristine picture cards, they acquire word meanings all at once or not at all. This result is important for several reasons.

First and foremost, it provides a concrete illustration of why we should not equate acquisition with learning.  MSTG provides diagnostics of learning (multiple hypotheses, statistical weighing of these alternatives, gradual convergence on the right result) and argues that learning so understood fails to hold in more realistic contexts of lexical acquisition. Specific conclusion; in at least one demonstrable case, acquisition does not equal learning.  More general conclusion; it is an empirical question whether in any given acquisition context it is true that learning (now understood to be one mechanism among others for the acquisition of knowledge) is taking place.  In other words, Jimmy and Judy certainly acquired English but whether they learned it is entirely open for empirical grabs. Chomsky’s repeated suggestion that we understand language acquisition as a species of growth rather than learning makes an analogous point.  MSTG makes the case more crisply, I believe, by providing clear diagnostics of learning and showing that there are core cases of “learning” where these signature properties of learning are demonstrably absent.

Second, MSTG provide a rationale for why learning doesn’t hold in their examined cases. Commenting on this (here) I observed that the MSTG results suggest that in such busy contexts the prerequisites for “cross situational learning” do not exist and this is why an alternative acquisition strategy is employed. Following MSTG’s lead, I even suggested that learning requires the kind of structured hypothesis space provided by the contrived set ups that MSTG’s more realistic vignettes argue against. This constituted a kind of compromise position; learning applies where acquirers have well structured hypothesis spaces and leaping to conclusions holds where this fails to hold.[2]  However, (and learn from this you soft-hearted open-minded intellectual compromisers out there) once again Norbert’s natural generosity of spirit and desire for intellectual group-hug kumbaya moments led him astray. It seems that even this compromise position concedes too much to learning mechanisms. In a companion paper, Trueswell, Medina, Hafri and Gleitman (TMHG) extend the MSTG results to include the more stylized learning environments in which options are clearly marked and lexical targets (aka referents) are crisply identified.[3]  Even in the artificial setting of the experimental psych lab kids and adults do the darndest things!

More specifically, like MSTG, TMGH identify the quiddity of “cross situational learning” with the following mechanism:

…keeping track of multiple hypotheses about a word’s meaning across successive learning instances, and gradually converg[ing] on the correct meaning via an intersective statistical process (128).

What they demonstrate is that even in simple stylized psych lab contexts when acquirers “are placed in the initial novice state of identifying word-to-referent mappings across learning instances, evidence for such a multiple hypothesis tracking procedure is strikingly absent (128),” and learners don’t “track the cross-trial statistics in order to build the correct mappings between words and referents (129).”  What acquirers do is “make a single conjecture upon hearing the word used in context and carry that conjecture forward to be evaluated for consistency with the next observed context (129).”  If confirmed, the guess is retained, if disconfirmed, acquirers guess again as if de novo.

TMGH show this (same with MSTG) by considering the dynamics of knowledge fixation: “how learning accuracy unfolded across learning instances (130).” This is very interesting stuff and I strongly recommend that you look at the details. However, just as interesting is the little bit of history that TMGH review. TMGH acknowledge tapping into a long history of criticism of this traditional gradual/comparative conception of learning.  In the late 1950s and early 1960s first Irvin Rock (1957) and then William Estes (1960) used very analogous kinds of arguments to demonstrate that verbal learning was not learning at all but was one-trial guessing. They did not fare so well. As Roediger and Arnold (RA) (2012) put it “the verbal learning establishment rose up to smite down these new ideas (2).”  It is instructive to read the RA paper for it shows the weakness of the counter-arguments used against Rock and Estes that nonetheless carried the day.  Just like today, learning was less a hypothesized mechanism for acquisition than a definitional truth about it.[4] 

The TMHG paper ends with an interesting paradox that I would like to briefly discuss as well. Initial word learning is slow and laborious (35-140 words at 14 months) but unbelievably rapid thereafter (12,000 words by age 6, i.e. roughly 7 words per day for a little under 5 years).  Why the change? TMHG note other work (Lila’s syntactic bootstrapping hypothesis) which proposes that “the acquisition of syntax and other linguistic knowledge by toddlers and young preschoolers during this time period provides a rich database of additional constraints that permit the learning of many additional words.” It is conceivable that with this knowledge in place, cross situational learning might finally become operative, though as TMHG correctly observe, this is decidedly an empirical question and it is “plausible that a propose-but-verify word learning procedure is at work all along the course of word learning throughout most of the life cycle (150).”  It would be interesting in the extreme, in my view were either conclusion correct.

If the latter conclusion proved true (propose and verify all the way down), then language acquisition might have nothing to do with learning in the technical sense. Of course, this is consistent with grammar acquisition being a case of learning, i.e. perhaps we don’t learn words but we do learn parameters. Maybe, but I’d be very skeptical. If even word acquisition isn’t an instance of learning then it seems to me that the burden of proof that any area of language acquisition involves learning would be pretty high.

If, on the other hand, the former option is correct, (viz. that learning only kicks in when grammar is there to buttress it) then its role in accounting for language acquisition would, in my opinion, be quite modest.  Yes, learning plays a role, but really most of the action lies with the constraints that grammars place on the process. This does not mean that we should not study the extras that learning might be adding (though remember the possibility mooted in the prior paragraph), but I doubt that these cognitive titivations will generate much excitement if they only operate in restricted hypothesis spaces. Learning is interesting when there are lots of options that need sifting, not so much when the range of possible end states is highly restricted. 

MSTG and TMHG show us that terminology matters. If we call acquisition “learning” we’ve loaded the research dice. If MSTG and TMHG are right (and I’d bet quite a bit that they are: any suckers out there?) it looks like we’ve repeatedly loaded them to come up snake eyes. It’s time we cut our losses and open our minds to the possibility that ‘language learning’ like ‘Justice Roberts,’ ‘military intelligence,’ and ‘western civilization’ is an oxymoron.



[1] I include this hedge for every week another (usually French speaking) psychologist shows that the youngest kids seem to have the most prodigious knowledge. It seems that we have nothing to teach those little know-it-alls. 
[2] A kind of learning as first resort, jumping to conclusions a last resort, method. C.f. TMHG p. 130.
[3] It is doubtful that the meaning of a lexical item is simply its referent for reasons that Chomsky has belabored (sadly, quite unsuccessfully) over the years. There is far more structure to lexical items than what they refer to. However, for current purposes, this only further dramatizes the inadequacy of learning as a mechanism for lexical acquisition. 
[4] TMHG also cite another paper by Gallistel and friends that I have not yet read but will try to get hold of and blog about when I do. It argues (cited in TMHG 151), that “in most subjects, in most paradigms, the transition from a low level of responding to an asymptotic level is abrupt.” Oh my. I’ll keep you posted.

Monday, December 3, 2012

I'm posting, with her permission, a comment to me from Lila Gleitman

I am posting this here, rather than as comment for there are papers she cites that I have added links to.


Hi Norbert, I read your blog and there is much to say.   Re the header above (i.e. False Truisms NH), I only fear that you, like the rest of the world, might be taking these findings as an excuse to write off the domain-specific structure-dependent scheme for the lexicon that I have devoted myself to, last several decades!   As we say in this new paper, but without room for discussion at all, is that only a minute, tiny, microscopically small set of the words are acquired "by" observation (and even this leaves aside that no one has a clue about how seeing a dog could teach you the "meaning" of dog, though it could spotlight, pun, the intended referent in best cases).   We did not sample the words.   We deliberately chose from the teeniest set of whole-object basic-level nouns those very few that our subjects could guess with any accuracy at all (roughly, 1/2 the time, with all other words guessed correctly a laughable maximum of 7% of the time), and only from specially chosen "good" exemplars.   All the same, "syntactic bootstrapping" skirts with circularity unless there is another -- non-syntactic -- procedure to learn a few "seed" words, those that, once and painfully acquired, can help you build a representation of the input (something like a rudimentary clause, or at least, to decide on the structural position of the Subject NP); it is that improved input that next makes a snap of learning all the words.   So if we now say (and we do) that there is a procedure for asyntactic, domain-general, learning of first words, this alone can't do for all the rest (try learning probably, think etc from watching scenarios in which these words are uttered -- or try even learning jump, tail, animal from the evidence of the senses, as Plato rightly noted).   So please though I'm pleased at your response to this new work, don't forget that its role is minute over the lexical stock, though crucial as the starting point.

Second, it turns out (I believe) that one trial learning is the rule rather than an exception that god makes for word learning in particular.  Despite my fondness for the small-set-of-options story we told in that first paper (don't you love it! something right about it) it turns out that subjects behave the same way if you put them in the icon-to-sound experimental condition studied by our opponents.  And I have attached the paper showing this!   The paper you read examines the natural case (using video of real contexts) and this paper examines the fake case, but in so doing achieves a level of experimental control we couldn't attain originally.  Again it is one-trial learning with no savings of the choices not made.  I think you'll like this, because it really exposes the logic.   And most important, it turns out (we at least mention this past-and-present literature on one-trial learning, particularly Gallistel, who speaks for the ants and the wasps) that learning in general (until, as I say, it becomes structured) across tasks and species has this character.   There has always been a small subset of the psychologists (Rock, Guthrie, Gallistel...) who denied, on their data + logic that associationist, gradualist, learning was the key, but they have always been overwhelmed in number and influence by the associationists.  As you point out, even you (even Chomsky, if you want to go back to his early writings) think/thought about phoneme/morpheme learning this way.   A marvelous new paper from Roediger traverses this literature and ably makes the case that learning in the general case is more determinative and less statistical than you thought.

We have some new work, I think important, showing the temporal and situational conditions that actually support the primitive first procedure for word learning.