Comments

Showing posts with label Gleitman. Show all posts
Showing posts with label Gleitman. Show all posts

Friday, October 10, 2014

Two kinds of Poverty of Stimulus arguments

There are two kinds of questions linguists would like to address: (1) Why do we see some kinds of Gs and never see others and (2) Why do kids acquire the particular Gs that they do. GG takes it that the answer to (2) is usefully informed by an answer to (1). One reason for thinking this is that both questions have a similar structure. Kids are exposed to products of a G and on the basis of these products they must infer the structure of the G that produces it. In other words, from a finite set of examples, a Language Acquisition Device (LAD) must infer the correct underlying function, G, that generates these examples. What does ‘correct’ mean? That G is correct which not only covers the finite set of given examples, but also correctly predicts the properties of the unbounded number of linguistic objects that might be encountered. In other words, the “right” G is one that correctly projects all possible unseen data from exposure to the limited input data.[1] GG calls the input examples the ‘primary linguistic data’ (PLD), and contrasts this with ‘linguistic data’ (LD), which comprises the full range of possible linguistic expressions of a given language L (e.g. ‘Who did John see’ is an example of PLD, ‘*Who did John see a man who likes’ is an example of LD). The correct G is that G which covers the PLD and also covers all the non-observed LD. As LD is in effect infinite, and PLD is necessarily finite, there’s a lot of unseen stuff that G needs to cover.[2] 

The very general characterization, let’s call it the Projection Problem (PrP), can cover both (1) and (2) above. Indeed, the standard PoS argument is based on a specific characterization of PrP. How so?

First, a standard PoS argument gives the following characterization of the PLD. It consists of well-formed, “simple,” sound/meaning (SM) pairs generated from a single G. In other words, the data used to infer the right G is “perfect” (i.e. no noise to speak of) but circumscribed (i.e. only “simple” data (see here for some discussion)).[3] Second, it assumes that the data is abundant. Indeed, it is counterfactually presumed that the PLD is presented “all at once,” rather than in smaller incremental chunks.[4] In short, the PoS makes two important assumptions about the PLD: (i) it is restricted to “simple” data, (ii) it is noiseless, homogeneous, and abundant (i.e. there is no room for variance as there would be were the data presented incrementally in smaller bits). Last, the LAD is also assumed to be “perfect” in having no problem in accurately coding the information the PLD contains and no problems computing its structure and relating it to the G that generated it. This idealization eliminates another source of potential noise. Thus, the quality of the data wrt input and intake, is assumed to be flawless.

Given these (clearly idealized) assumptions the PoS question is how does the LAD go from PLD/LAD so described to a G able to generate the full range of data (i.e. both simple and complex)? The idealization isolates the core of the PoS argument: getting from PLD to the “correct” G is massively underdetermined by the PLD even if we assume that the PLD is of immaculate quality. The standard PoS conclusion is that the only way to explain why some kinds of Gs are unattested is to assume that some (logically possible) inductions from PLD to G are formally illict. That’s the Projection Problem as it relates to (1). UG (i.e. formal restrictions on the set of admissible Gs) is the proposed answer.

Next step: assume now that we have a fully developed theory of UG. In other words, let’s assume that we have completely limned the borders of possible Gs. We are still left with question (2). How does the LAD acquire the specific G that it does? How does the LAD use the PLD to select one among the many possible Gs? Note that it appears (at least at first blush) that restricting our attention to selecting the specific G compatible with the given PLD from among the grammatically possible Gs (rather than from all the logically possible Gs) simplifies the problem. There are a whole lot of Gs that LAD need never consider precisely because they are grammatically impossible. And it is conceivable that finding the right G among the grammatically admissible ones requires little more than matching PLD to Gs. So, one possible interpretation of the original Chomsky program is that once UG is fixed, acquisition reduces to simple learning (e.g. once the UG principles are specified, acquisition is little more than standard matching of data to Gs). On this view, UG so restricts the class of accessible Gs that using PLD to search for the right G is relatively trivial. 

There is another possibility, however. Even with the invariant principles fixed (i.e. even once we specified the impossible (kinds of) Gs), the PLD is still too insubstantial to select the right G given PLD (i.e. the PLD still underdetermines choice of the right G). On this second scenario, additional machinery (perhaps some of it domain specific) is required to navigate the remaining space of possible grammatical options. Or another way of putting this: fixing the invariant principles of UG does not suffice to uniquely select a G given PLD? 

There is reason to think Chomsky, at least in Aspects took door number 2 above.[5] In other words, “since the earliest days of generative grammar” (as Chomsky likes to say), it has been assumed that a usable acquisition model will likely need both a way of eliminating the impossible Gs and another (perhaps related, perhaps not) set of principles to guide the LAD to its actual G.[6] So, in addition to invariant principles of UG, GG also deployed markedness principles (i.e. “priors”) to play a hefty explanatory role. So, for example, say the principles of UG delimit the borders of the hypothesis space, Gs within the borders being possible. Acquisition theory (most likely) still requires that the Gs within the borders have some kind of preferential ordering, with some Gs better than others.

To repeat, this is roughly the Aspects view of the world and it is one that fits well with the Bayes conception where in addition to a specification of the hypotheses entertained, some are endowed with higher priors than others. P&P models endorse a similar conception as some parameters, the unmarked ones, are treated as more equal than others. Thus, while the invariant principles and open parameters delimit the space of G options, markedness theory (or the evaluation metric) is responsible for getting an LAD to specific parameter values on the basis of the available PLD.

This division of labor seems reasonable, but is not apodictic.  There is a trading relation between specifying high priors and delimiting the hypothesis space. Indeed, saying that some option is impossible amounts to setting the prior for this option to 0 and saying that it is necessary amounts to setting the prior to 1.  Moreover, given our current state of knowledge, it is unclear what the difference is between assuming that something is impossible given PLD versus saying that it is very improbable. However, it is not unreasonable, IMO, to divide the problem up as above as several kinds of things really do seem unattested while other things though possible are not required.

With this as background, I want to now turn to a kind of PoS argument that builds on (steals from?) a terrific paper that I’ve recently read by Gigerenzer and Brighton (G&B) (here) and that I have been recommending to all and any in my general vicinity in the last week.

G&B discuss the role of biases in inductive learning. The discussion is under the rubric of heuristics. They note that biases/heuristics have commonly been motivated on grounds of reducing computational complexity. As noted several times before in other posts (e.g. here), many inductive theories are computationally intensive if implemented directly. In fact, so intensive as to be intractable.  I’ve mentioned this wrt Bayesian models and several commentators noted (here) that there are reasons to hope that these problems can be finessed using various well-known (in the sense of well-known to those in the know, i.e. not to me) statistical sampling methods/algorithms. These methods can be used to approximate the kinds of solutions the computationally intractable direct Bayesian methods would produce were they tractable. Let’s call these methods “heuristics.” If correct, this constitutes one good cognitive argument for heuristics; they reduce the computational complexity of a problem making its solution tractable.  As G&B note, on this conception, heuristics (and the biases they incorporate) are the price one has to pay for tractability. Or; though it would be best to do the obvious calculation, such calculations are sadly intractable and so we use heuristics to get the calculations done even though this sacrifices (or might sacrifice) some accuracy for tractability. They call this the accuracy-effort tradeoff (AET). As G&B put it:

If you invest less effort the cost is lower accuracy. Effort refers to searching for more information, performing more computation, or taking more time; in fact these typically go together. Heuristics allow for fast and frugal decisions; thus, it is commonly assumed that they are second best approximations of more complex “optimal” computations and serve the purpose of trading off accuracy for effort. If information were free and humans had eternal time, so the argument goes, more information and computation would always be better (109).

G&B note that this is the common attitude towards heuristics/biases.[7]  They exist to make the job doable. And though G&B agree that this might be one reason for them, they think that it is not the most important helpful feature that heuristics/biases have. So what is?  G&B highlight a second feature of biases/heuristics; what they call the “bias-variance dilemma” (BVD).  They describe it as follows:[8]

… achieving a good fit to observations does not necessarily mean we have found a good model, and choosing a model with the best fit is likely to result in poor predictions…(118).

Why? Because

…bias is only one source of error impacting on the accuracy of model predictions. The second source is variance, which occurs when making inferences from finite samples of noisy data. (119).

In other words, a potentially very serious problem is “overfitting,” a problem that flexible models standardly enjoy. In G&B’s words:

The more flexible the model, the more likely it is to capture not only the underlying pattern but unsystematic patterns such as noise…[V]ariance reflects the sensitivity of the induction algorithm to the specific contents of samples, which means that for different samples of the environment, potentially very different models are being induced.  [In such circumstances NH] a biased model can lead to more accurate predictions than an unbiased model. (119)

Hence the dilemma: To best cover the input data set, “model must accommodate a rich class of patterns in order to insure low bias.” But “[t]he price is an increase in variance, as the model will have greater flexibility, this will enable it to accommodate not only systematic patterns but also accidental patterns such as noise” (119-120). Thus a btter fit to the input may have deleterious effects on predicting future data. Hence the BVD:

Combating high bias requires using a rich class of models, while combating high variance requires placing restrictions on this class of models. We cannot remain agnostic and do both unless we are willing to make a bet on what patterns will occur. This is why “general purpose” models tend to be poor predictors of the future when data are sparse (120).

And the moral G&B draw?

The bias-variance dilemma shows formally why a mind can be better off with an adaptive toolbox of biased specialized heuristics. A single, general-purpose tool with many adjustable parameters is likely to be unstable and incur greater prediction error as a result of high variance. (120)

What consequences might the BVD have for work on language? Well, note first of all that it provides the template for an additional kind of PoS argument. In contrast to the standard one reviewed above, this one holds when we relax the standard idealizations reviewed above; in particular, the assumption that the PLD is noise free and that it is provided all-at-once.  We know that these assumptions are false, what the BVD suggests is that when these are relaxed we potentially encounter another kind of inductive problem in which biases can be empirically very useful. I say “suggests” rather than “shows” because as G&B demonstrate quite nicely, whether the problem is a real one, depends on how sparse and noisy the relevant PLD is. 

The severity of the BVD problem in linguistics will likely depend on the particular linguistic case being studied. So for example, work by Gleitman, Trueswell and friends (discussed here, here, here) suggests that at least early word learning occurs in very noisy data sparse environments. This is just the kind that G&B point to as favor shallow non-intensive data analysis. The procedure that Gleitman, Trueswell and friends argue for seems to fit well into this picture.

I’m no expert in the language acquisition literature, but from what I’ve seen, the scenarios that G&B argue promote BVDs are rife in the wild. I sure looks like many people converge to (very close to) the same G despite plausibly having very different individual inputs (isn’t this the basis for the overwhelming temptation to reify languages?  Believe me my Polish parent English PLD was quite a bit different from that of my Montreal peers and we ended up sounding and speaking very much the same). If so, the kinds of biased systems that GG is very comfortable with will be just what G&B ordered. However, whether this always holds or even whether it ever holds is really an empirical question.[9]

G&B contrasts heuristic systems with more standard models, including Bayesian models, exemplar models, multiple regression models etc. that embody Carnap’s “principle of total evidence” (110). From what G&B say (and I have sort of confirmed by doing econometrician on the campus interviews), it appears that most of the current favored approaches to rational decision making embody this principle, at least as an ideal. As a favorite conceit is to assume that cognitively speaking, humans are very rational, indeed optimal decision makers, Carnap’s principle is embodied in most of the common approaches (indeed Bayesians love to highlight this). Theories that embody Carnap’s principle understand “rational decision making as the process of weighing and adding all information” up to computational tractability.  The phenomena that G&B isolates (what the paper dubs “less is more” effects) challenge this vision. These effects, G&B argues, illustrate that it’s just false that more is always better even in the absence of computational constraints. Rather, in some circumstances, the ones that G&B identifies, shallow and blinkered is the way to go. And if this is correct, then the empirical questions will have to be settled on a case by case basis, sometimes favoring total evidence based models and sometimes not. Further, if this is correct (and if Bayesian models are species of total evidence models) then whether a Bayesian approach is apposite in a given cognitive context becomes an empirical question, the answer depending on how well behaved the data samples are.

Third, it would not be surprising (at least to me) were there two (or at least two) kinds of native FL biases, corresponding to the two kinds of PoS arguments discussed above.  It is possible that the biases motivated via the classical PoS argument (the invariances that circumscribe the class of possible Gs) alone suffice to lead the LAD to its specific G. However, this clearly need not be so.  Nor is it obvious (again at least to me) that the principles that operate within the circumscribed class of grammatically possible grammars would operate as well within the wider class of logically possible ones.  Indeed, when specific examples are considered (e.g. ECP effects, island effects, binding effects) the case for the two-prong attack on the PoS problem seems reasonable. In short, there are two different kinds of PoS problems invoking different kinds of mechanisms.

G&B ends with a description of two epistemological scenarios and the worlds where they make sense.[10] Let me recap them, comment very briefly and end.

The first option has a mind with no biases “with an infinitely flexible system of abstract representations.” This massive malleability allows the mind to “reproduce perfectly” “whatever structure the world has.” This mind works best with “large samples of observations” drawn from world that is “relatively stable.” Because such a mind “must choose from an infinite space of representations, it is likely to require resource intensive cognitive processing.” G&B believes that exemplar models and neural networks are excellent models for this sort of mind. (136)

The second mind makes inferences “quickly from a few observations.” The world it lives in changes in unforeseen ways and the data it has access to is sparse and noisy. To overcome this it uses different specialized biases that can “help to reduce the estimation error.” This mind need not have “knowledge of all relevant options, consequences and probabilities both now and in the future” and it “relies on several inference tools rather than a single universal tool.” Last, in this second scenario intensive processing is not required nor favored.  Rather minds come packed with specialized heuristics able to offset the problems that small noisy data brings with it.

You probably know where I am about to go. The first kind of mind seems more than just a tad familiar from the Empiricist literature. “Infinitely flexible” minds that “reproduce perfectly” “whatever structure the world has” sound like the perfect wax tablets waiting to faithfully receive the contours that the world via “large sample of observations” is ready to structure it with. The second with its biases and specialized heuristics has a definite Rationalist flavor. Such minds contain domain specific operations. Sound familiar? 

What G&B adds to the standard Empiricism-Rationalism discussion is not these two conceptions of different minds, but the kinds of advantages we can expect from each given the nature of the input and the “worlds’ that produce it. When a world is well behaved, G&B observes, minds can be lightly structured and wait for the environment to do its work. When it is a blooming buzzing confusion bias really helps. 

There is a lot more in the G&B paper. I found it one of the more stimulating and thought provoking things I’ve read in the last several years. If G&B is correct, the BVD is rich in consequences for language acquisition models that begin to loosen the idealizations characteristic of Plato’s Problem ruminations. Most interestingly, at least to me, coarsening the idealization adds new reasons for assuming that biological systems come packed with rich innately structured minds. In the right circumstances, they don’t only relieve computational burdens, they allow for good inference, indeed better inference than a mind that more carefully tracks the world and intensively computes the consequences of this careful tracking. Interesting, very interesting. Take a look.




[1] This problem goes back to the very beginning GG, see Stanley Peters’ paper “The Projection Problem: How is a grammar to be selected” in Goals of Linguistic Theory. As he noted in his paper, this problem is closely tied to the question of Explanatory Adequacy. The logic outlined above is very clearly articulated in Peters’ paper. He describes the projection problem  as the “problem of providing a general scheme which specifies the grammar (or grammars) tht can be provided by a human upon exposure to a possible set of basic data” (172).
[2] Note that the projection problem can hold for finite sets as well. The issue is how to select the function that covers the unobserved on the basis of the observed (i.e. how to generalize from a small sample to a larger one). How does a system “project” to the unobserved data based on the observed sample. The infinity assumption allows for a clear example of the logic of projection. It is not a necessary feature.
[3] Peters also zeros in on the idea that PLD is “simple.” As he puts it: “as has often been remarked, one rarely hears a fully grammatical sentence of any complexity…One strategy open to him [the LAD, NH] is to put the greatest confidence in short utterances, which are likely to be less complex than longer ones and thus more likely to be grammatical” (175).
[4] As noted here this assumption quite explicit in Aspects is known to be a radical idealization. However, this does not indicate that it has baleful consequences. It does seem that kids in the same linguistic environment come to acquire very similar competences (no doubt the source of our view that languages exist). This despite the reasonable conjecture that they are not exposed (or intake) exactly the same (kinds of) sentences in the same order. This suggests that order of presentation is not that critical and this is follows from the all-at-once idealization. That said, I return to this assumption below. For some useful discussion see Peters where the idealization is defended (p.175).
[5] Again, see Peters for illuminating discussion.
[6] This overstates the case. The evaluation measure did no tell the LAD how to construct a G given PLD. Rather it specified how to order Gs as better or worse given PLD.  In other words, it specifies how to rank two given Gs. Specifying how to actually build these was considered (and probably still is) too ambitious a goal.
[7] I suspect that the general disdain for priors in Bayesian accounts is the belief that they do not fundamentally alter the acquisition scenario. What I mean by this is that though they may accelerate or impede the rate at which one gets to the best result, over enough time the data will overwhelm the priors so that even if one starts, as it were, in the wrong place in the hypothesis space, the optimal solution will be attained. So priors may affect computation and the rate of convergence to the optimum but it cannot fundamentally alter the destination.
[8] By “fit” here G&B mean fit with the input data sets.
[9] So Jeff Lidz noted that perhaps all LADs enjoy a good number of rich learning encounters where sufficient amounts of the same good data is used.  In other words, though the data overall might stink, there are reliable instances where the data is robust and there are where the acquisition action takes place.  This is indeed possible, it seems to me, and this is what makes the BVD problem an empirical, rather than a conceptual, one.
[10] There are actually three, but I ignore the first as it has little real interest.

Monday, July 29, 2013

More on Word Acquisition


In some earlier posts (e.g. here, here), I discussed a theory of word acquisition developed by Medina, Snedeker, Trueswell and Gleitman (MSTG) that I took to question whether learning in the classical sense ever takes place.  MSTG propose a theory they dub “Propose-but-Verify” (PbV) that postulates that word learning in kids is (i) essentially a one trial process where everything but the first encounter with a word is essentially irrelevant, (ii) at any given time only one hypothesis is being entertained (i.e. there is no hypothesis testing;/comparison going on) and (iii) that updating only occurs if the first guess is disconfirmed, and then it occurs pretty rapidly.  MSTG’s theory has two important features. First, it proceeds without much counting of any sort, and second, the hypothesis space is very restricted (viz. it includes exactly one hypothesis at any given time). These two properties leave relatively little for stats to do as there is no serious comparison of alternatives going on (as there’s only one candidate at a time and it gets abandoned when falsified).

This story was always a little too good to be true. After all, it seems quite counterintuitive to believe that single instances of disconfirmation would lead word acquirers (WA) to abandon a hypothesis.  And not surprisingly, as is often the case, those things too good to be true might not be. However, a later reconsideration of the same kind of data by a distinguished foursome (partially overlapping, partially different) argues that the earlier MSTG model is  “almost” true, if not exactly spot on.

In a new paper (here) Stevens, Yang, Trueswell and Gleitman (SYTG) adopt (i)-(iii) but modify it to add a more incremental response to relevant data. The new model, like that older MSTG one, rejects “cross situational learning” which SYTG take to involve “the tabulation of multiple, possibly all, word-meaning associations across learning instances” (p.3) but adds a more gradient probabilistic data evaluation procedure. The process works as follows. It has two parts.

First, for “familiar” words, this account, dubbed “Pursuit with abandon” (p. 3) (“Pursuit” (P) for short), selects the single most highly valued option (just one!) and rewards it incrementally if consistent with the input and if not it decreases its score a bit while also randomly selecting a single new meaning from “the available meanings in that utterance” (p. 2) and rewarding it a bit. This take-a-little, give-a-little is the stats part. In contrast to PbV, P does not completely dump a disconfirmed meaning, but only lowers its overall score somewhat.  Thus, “a disconfirmed meaning may still remain the most probable hypothesis and will be selected for verification the next time the word is presented in the learning data” (p. 3). SYTG note that replacing MSTG’s one strike you’re out “counting” procedure, with a more gradient probabilistic evaluation measure adds a good deal of “robustness” to the learning procedure.

Second, for novel words, P encodes “a probabilistic form of the Mutual Exclusivity Constraint…[viz.] when encountering novel words, children favor mapping to novel rather than familiar meanings” (p. 4).  Here too the procedure is myopic, selecting one option among many and sticking with it until it fails enough to be replaced via step one above.

Thus, the P model, from what I can tell, is effectively the old PbV model but with a probabilistic procedure for, initially, deciding on which is the “least probable” candidate (i.e. to guide an initial pick) and for (dis)confirming a given candidate (i.e. to up/downgrade a previously encountered entry).  Like the PbV, P is very myopic. Both reject cross situational learning and concentrate on one candidate at a time, ignoring other options if all goes well and choosing at random if things go awry.

This is the P model. Using simulations based on Childes data, the paper goes on to show that this system is very good when compared both with PbV and, more interestingly, with more comprehensive theories that keep many hypothesis in play throughout the acquisition process. To my mind, the most interesting comparison is with Bayesian approaches. I encourage you to take a look at the discussion of the simulations (section 3 in the paper).  The bottom line is that the P model bested the three others on overall score, including the Bayesian alternative.  Moreover, SYTG was able to identify the main reason for the success: non-myopic comprehensive procedures fail to sufficiently value “high informative cues” provided early in the acquisition process.  Why? Because comprehensive comparison among a wide range of alternatives serves to “dilute the probability space” for correct hits, thereby “making the correct meaning less likely to be added to the lexicon” (P. 6-7).  It seems that in the acquisition settings found in CHILDES (and in MSTGs more realistic visual settings), this dilution prevents WAs from more rapidly building up their lexicons. As SYTG put it:

The advantage of the Pursuit model over cross-situational models derives from its apparent sub-optimal design. The pursuit of the most favored hypothesis limits the range of competing meanings. But at the same time, it obviates the dilution of cues, especially the highly saline first scene…which is weakened by averaging with more ambiguous leaning instances…[which] are precisely the types of highly salient instances that the learner takes advantage of…(p. 7).

There is a second advantage of the P model as compared to a more sophisticated and comprehensive Bayesian approach.  SYTG just touch on this, but I think it is worth mentioning. The Bayesian model is computationally very costly. In fact, SYTG notes that full simulations proved impractical as “each simulation can take several hours to run” (p. 8).  Scaling up is a well-known problem for Bayesian accounts (see here), which is probably why Bayesian proposals are often presented as Marrian level 1 theories rather than actual algorithmic procedures. At any rate, it seems that the computational cost stems from precisely the feature that makes Bayesian models so popular: their comprehensiveness. The usual procedure is to make the hypothesis space as wide as possible and then allow the “data” to find the optimal one. However, it is precisely this feature that makes the obvious algorithm built on this procedure intractable. 

In effect, SYTG show the potential value of myopia, i.e. in very narrow hypothesis spaces. Part of the value lies in computational tractability. Why? The narrower the hypothesis space, the less work is required of Bayesian procedures to effectively navigate the space of alternatives to find the best candidate.  In other words, if the alternatives are few in number, the bulk of explaining why we see what we get will lie not with fancy evaluation procedures, but with the small set of options that are being evaluated. How to count may be important, but it is less important the fewer things there are to count among. In the limit, sophisticated methods of counting may be unnecessary, if not downright unproductive.

The theme that comprehensiveness may not actually be “optimal” is one that SYTG emphasize at the end of their paper. Let me end this little advertisement by quoting them again:

Our model pursues the [i.e. unique NH] highly valued, and thus probabilistically defined, word meaning at the expense of other meaning candidates.  By contrast, cross-situational models do not favor any one particular meaning, but rather tabulate statistics across learning instances to look for consistent co-occurrences. While the cross-situational approach seems optimally designed [my emph, NH], its advantage seems outweighed by its dilution effects that distract the learner away from clear unambiguous learning instances…It is notable that the apparently sub-optimal Pursuit model produces superior results over the more powerful models with richer statistical information about words and their associated meanings: word learning is hard, but trying to hard may not help.

I would put this slightly differently: it seems that what you choose to compare may be as (more?) important than how you choose to compare them. SYTG reinforces MSTG’s earlier warning about the perils of open-mindedness. Nothing like a well designed narrow hypothesis space to aid acquisition. I leave the rationalist/empiricist overtones of this as an exercise for the reader. 

Sunday, December 2, 2012

A False Truism?


It’s my strong impression that everyone working on language, be they generativist partisans or enemies, takes it for granted that linguistic performance (if not competence) involves a heavy dose of stats. So, for example, the standard view of language learning is basically Bayesian, i.e. structured hypothesis space with grammars ordered by some sort of simplicity metric the “winner” being the simplest one consistent with the incoming data, the procedure involving a comparison of alternatives ranked by simplicity and conformity with the data.  This is indistinguishable from the set up in Aspects chapter 1.
Same with language processing, where alternative hypotheses about the structure of the incoming sounds/words/sentences are compared and assessed; the simplest one best fitting the data carrying the parsing day.[1] Thus, wisdom has it that language use requires the careful, gradual and methodical assessment of alternatives, which involves iteratively trading off some sophisticated measure of goodness of fit against some measure of simplicity to eventually get to a measure of believability. Consequently, nobody really wonders anymore whether statistical estimation/calculation is a central feature of our cognitive lives but which particular versions are correct.  Call this the “Stats Truism.”  Jeff Lidz recently sent me a very interesting paper that suggests that this unargued for presupposition, one incidentally that I have shared (do share?), should be moved from the truism column to the an-assumption-that-needs-argument-and-justification pile. Let me explain.

Kids acquire grammars very quickly. By about age 5 a kid’s grammar is largely set.  What’s equally amazing is that by age six, kids typically have a vocabulary of 6,000-8,000 words, which translates into an acquisition rate of roughly 6-8 words per day from the day they were born. This is a hell of a rate! (remember, for the first several years kids sleep half the day (if their parents are lucky) and cry for the other half (do I remember that!)). Psychologists have investigated how they do this and the current wisdom is that they employ some kind of fast mapping of sounds to concepts (so much for Quine’s derision of the museum myth!). Medina, Snedeker, Trueswell and Gleitman (MSTG) study this fast mapping process and what they discover is a serious challenge for the Stats Truism. Here’s why. They find that kids learn words more or less as follows: they quickly jump to a conclusion about what a word means and don’t let go!  They don’t consider alternative possibilities, they don’t revise prior estimates of their initial guess, and they don’t much care about the data after their initial guess (i.e. they don’t guess again when given disconfirming evidence for their initial hyporthesis, at least for a while).[2] Or, in MSTG’s own words:

1.     “Learners hypothesize a single meaning based on their first encounter with a word (3/6).” Thus, learning is essentially a one trial process where everything but the first encounter is irrelevant.
2.     “Learners neither weight nor even store back-up alternative meanings (3/5).” Thus, there is no (fancy (i.e. Bayesian) or crude (i.e. simple counting)) hypothesis comparison/testing going on as in word learning kids only ever entertain a single hypothesis.
3.     “On later encounters, learners attempt to retrieve their single hypothesis from memory and test it against a new context, updating only if it is disconfirmed. Thus they do not accrue a “bests final hypothesis by comparing multiple …semantic hypotheses (3/6).” Indeed, MSTG note that “a false hypothesis once formed blocks the formation of new ones (4/6).”

In sum, kids are super impulsive, narrow-minded, pig-headed learners, at least for words (though I suspect parents may not find this discovery this so surprising).

MSTG note that if they are correct (and the experiments are cool (and pretty convincing) so look them over) this constitutes a serious challenge to standard statistical models in which “each word-meaning hypothesis [is] based on the properties of all past learning instances regardless of the order in which they are encountered [and] numerous hypotheses [are held] in mind (with changing weights) until some learning threshold is reached (3/6).” SMTG’s point: at least in this domain the Stats Truism isn’t.

A further interesting feature of this paper is that it suggests why the Truism doesn’t hold.  MSTG argue that the standard experimental materials for investigating word learning in the lab don’t scale up. In other words, when materials that more adequately reflect real world situations are used the signature properties of statistical learning disappear.

The main difference between real world situations and the lab context is well summed up in an aphorism I once heard from Lila (the G in MSTG): “A picture is worth a thousand words, and that’s the problem.”  MSTG observes that in real life situations learners cannot match “recurrent speech events to recurrent aspects of the observed world because “the world of words and their contexts is enormously [i.e. too?-NH] complex (1/6).” As MSTG note, most word learning investigations abstract away from this complexity, either by assuming stylized learning contexts or assuming that a kid’s attention is directed in word learning settings so that the noise is cancelled out and the relevant stimulus is made strongly salient. MSTG provide pretty good evidence that these assumptions are unrealistic and that the domain in which words are learned is very busy and very noisy. As they note:

The world of words and their contexts is enormously complex. Few words are taught systematically…[I]n most instances, the situations of word use arise adventitiously as adults interact socially with novices. Words are heard buried inside multiword utterances and in situations that vary in almost endless ways…so that usually a listener could not be warranted in selecting a unique interpretation for a new item.

The important finding is that in these more realistic contexts, the signature properties of statistical learning disappear (not mitigated, not reduced, disappear).

MSTG suggest two reasons for this. First, within any given context too many things are plausibly relevant so picking out exactly what’s important is very hard to do.  Second, across contexts what’s relevant can change radically, so determining which features to consider is very difficult. In effect, there is no relevance algorithm either within or across contexts that the child can use to direct its attention and memory resources. Together these factors overwhelm those capacities that we successfully deploy in “stripped down laboratory demonstrations (1/6).”[3] Said less coyly: the real world is too much of a "blooming buzzing confusion" for statistical methods to be useful.  This suggestion sounds very counterintuitive, so let’s consider it for a moment.

It is generally believed that the virtue of statistical models is that they are not categorical and for precisely this reason they are better able to deal with the gradient richness of the real world (see here for example and here for discussion).  What MSTG’s results suggest is that crude rules of thumb (e.g. the first is best) are better suited to the real world than are sophisticated statistical models. Interestingly, others have made similar suggestions in other domains. For example, Andrew Haldane (here) invidiously compares bank supervision policies that depend on large numbers of weighted variables that are traded off against one another in statistically sophisticated ways with policies that use a single simple standard such as the leverage ratio. As he shows, the simple single standard blunt models outperform the fancy statistical ones most of the time. Gerd Gigerenzer (here) provides many examples where sophisticated statistical models are bested by rather ham fisted heuristic algorithms that rely on one-good-reason decision procedures rather than decision rules that weigh, compare and evaluate many reasons. Robert Axelrod observes how very simple rules of behavior (tit for tat) can triumph in complex interactive environments against sophisticated rules. It seems that sometimes less is more, and simple beats sophisticated.  More particularly, when the relevant parameters are opaque or it is uncertain how options should be weighted (i.e. when the hypothesis space is patchy and vague) statistical methods can do quite a bit worse than very simple very unsophisticated categorical rules.[4] In this setting, MSTGs results seem lees surprising.

MSTG makes yet one more interesting point.  It seems that the Stats Truism is true precisely when it operates against a rich, articulated, well-defined, set of options (“relatively” small helps too).[5] As MSTG note “statistical models have proven adequate for properties of language for which the learner’s hypothesis space is known to consist of a very small set of options (5/6).” I will leave it to your imagination to consider what someone who believes in a rich UG that circumscribes the space of possible grammars (me, me, me!) would make of this.

I would like to end with a couple of random observations.

First if anything like this is on the right track it further dooms the kind of big data investigations that I discussed here. What’s relevant is potentially unbounded. There’s no algorithm to determine it. To use Knight-Keynes lingo (see note 3), much of inquiry is uncertain rather than risky. If so, it’s not the kind of problem that big data statistical analysis will substantially crack precisely because inquiry is uncertain, not merely risky.[6]

Second, statistical language learning methods will be relevant in exactly those domains where the mind provides a lot of structure to the hypothesis space.  Absent this kind of UG-like structuring, sophisticated methods will likely fail. Word learning is an unstructured domain. Saussurian arbitrariness (the relation between a concept and the sound that tags it being arbitrary) implies that there isn’t much structure to word the learning context and, as MSTG show, in this domain, statistical learning methods are worse than useless, they are wrong. So, if you like statistical language learning and hope to apply it fruitfully you’d better hope that UG is pretty rich.

Third, there are two reasons for doubting that MSTG’s point will be quickly accepted. First, though simple, Baye’s Law looks a lot more impressive than ‘guess and stick’ or ‘first is best.’  There is a law of academia that requires that complicated trump simple. Simple may be elegant, but complexity, because it is obviously hard and shows off mental muscles (or at least appears to), enhances reputations and can be used to order status hierarchies. Why use simple rules when complex ones are available? Second, truisms are hard to resist because they do seem so true. As such the Stats Truism will not soon be dislodged. However, MSTG’s argument should at least open up our minds enough to consider the possibility that it may be more truthy than true.



[1] I believe that this is currently the dominant view in parsing (but remember I am no expert here). Earlier theories (see here) did not assume that multiple hypotheses were carried forward in time but that once decisions were made about the structure, alternatives were dropped. This was used to account for garden path phenomena and was motivated theoretically by the parsing efficiency of deterministic left corner parsers. 
[2] MSTG have an interesting discussion of how kids recover from a wrong guess (of which there are no doubt many). In effect, they “forget” what they guessed before.  Wrong guesses are easier to forget, correct ones stick in memory longer.  Forgetting allows this “first-is-best” process to reengage and allows the kid to guess again.
[3] Remember Cartwright’s observation (here for discussion) that Hume’s dictum is hard to replicate in the world outside the lab? Here’s a simple example of what she means.
[4] Sophisticated methods triumph when uncertainty can be reduced to risk. We tend to assume that these two concepts are the same. However, two very smart people argued otherwise: Frank Knight and J.M. Keynes argued for the importance of the distinction and urged that we not assimilate the two. The crux of the difference is that whereas risk is calculable, uncertainty is not. Statistical methods are appropriate in evaluating risk but not in taming uncertainty.
[5] Gigerenzer and Haldane also try to identify the circumstances under which more sophisticated procedures come into their own and best the simple rules of thumb.
[6] A point also made by Popper here. We cannot estimate what we don’t know. There are, in the immortal words of D. Rumsfeld “unknown unknowns.”