Comments

Showing posts with label Trueswell. Show all posts
Showing posts with label Trueswell. Show all posts

Friday, October 10, 2014

Two kinds of Poverty of Stimulus arguments

There are two kinds of questions linguists would like to address: (1) Why do we see some kinds of Gs and never see others and (2) Why do kids acquire the particular Gs that they do. GG takes it that the answer to (2) is usefully informed by an answer to (1). One reason for thinking this is that both questions have a similar structure. Kids are exposed to products of a G and on the basis of these products they must infer the structure of the G that produces it. In other words, from a finite set of examples, a Language Acquisition Device (LAD) must infer the correct underlying function, G, that generates these examples. What does ‘correct’ mean? That G is correct which not only covers the finite set of given examples, but also correctly predicts the properties of the unbounded number of linguistic objects that might be encountered. In other words, the “right” G is one that correctly projects all possible unseen data from exposure to the limited input data.[1] GG calls the input examples the ‘primary linguistic data’ (PLD), and contrasts this with ‘linguistic data’ (LD), which comprises the full range of possible linguistic expressions of a given language L (e.g. ‘Who did John see’ is an example of PLD, ‘*Who did John see a man who likes’ is an example of LD). The correct G is that G which covers the PLD and also covers all the non-observed LD. As LD is in effect infinite, and PLD is necessarily finite, there’s a lot of unseen stuff that G needs to cover.[2] 

The very general characterization, let’s call it the Projection Problem (PrP), can cover both (1) and (2) above. Indeed, the standard PoS argument is based on a specific characterization of PrP. How so?

First, a standard PoS argument gives the following characterization of the PLD. It consists of well-formed, “simple,” sound/meaning (SM) pairs generated from a single G. In other words, the data used to infer the right G is “perfect” (i.e. no noise to speak of) but circumscribed (i.e. only “simple” data (see here for some discussion)).[3] Second, it assumes that the data is abundant. Indeed, it is counterfactually presumed that the PLD is presented “all at once,” rather than in smaller incremental chunks.[4] In short, the PoS makes two important assumptions about the PLD: (i) it is restricted to “simple” data, (ii) it is noiseless, homogeneous, and abundant (i.e. there is no room for variance as there would be were the data presented incrementally in smaller bits). Last, the LAD is also assumed to be “perfect” in having no problem in accurately coding the information the PLD contains and no problems computing its structure and relating it to the G that generated it. This idealization eliminates another source of potential noise. Thus, the quality of the data wrt input and intake, is assumed to be flawless.

Given these (clearly idealized) assumptions the PoS question is how does the LAD go from PLD/LAD so described to a G able to generate the full range of data (i.e. both simple and complex)? The idealization isolates the core of the PoS argument: getting from PLD to the “correct” G is massively underdetermined by the PLD even if we assume that the PLD is of immaculate quality. The standard PoS conclusion is that the only way to explain why some kinds of Gs are unattested is to assume that some (logically possible) inductions from PLD to G are formally illict. That’s the Projection Problem as it relates to (1). UG (i.e. formal restrictions on the set of admissible Gs) is the proposed answer.

Next step: assume now that we have a fully developed theory of UG. In other words, let’s assume that we have completely limned the borders of possible Gs. We are still left with question (2). How does the LAD acquire the specific G that it does? How does the LAD use the PLD to select one among the many possible Gs? Note that it appears (at least at first blush) that restricting our attention to selecting the specific G compatible with the given PLD from among the grammatically possible Gs (rather than from all the logically possible Gs) simplifies the problem. There are a whole lot of Gs that LAD need never consider precisely because they are grammatically impossible. And it is conceivable that finding the right G among the grammatically admissible ones requires little more than matching PLD to Gs. So, one possible interpretation of the original Chomsky program is that once UG is fixed, acquisition reduces to simple learning (e.g. once the UG principles are specified, acquisition is little more than standard matching of data to Gs). On this view, UG so restricts the class of accessible Gs that using PLD to search for the right G is relatively trivial. 

There is another possibility, however. Even with the invariant principles fixed (i.e. even once we specified the impossible (kinds of) Gs), the PLD is still too insubstantial to select the right G given PLD (i.e. the PLD still underdetermines choice of the right G). On this second scenario, additional machinery (perhaps some of it domain specific) is required to navigate the remaining space of possible grammatical options. Or another way of putting this: fixing the invariant principles of UG does not suffice to uniquely select a G given PLD? 

There is reason to think Chomsky, at least in Aspects took door number 2 above.[5] In other words, “since the earliest days of generative grammar” (as Chomsky likes to say), it has been assumed that a usable acquisition model will likely need both a way of eliminating the impossible Gs and another (perhaps related, perhaps not) set of principles to guide the LAD to its actual G.[6] So, in addition to invariant principles of UG, GG also deployed markedness principles (i.e. “priors”) to play a hefty explanatory role. So, for example, say the principles of UG delimit the borders of the hypothesis space, Gs within the borders being possible. Acquisition theory (most likely) still requires that the Gs within the borders have some kind of preferential ordering, with some Gs better than others.

To repeat, this is roughly the Aspects view of the world and it is one that fits well with the Bayes conception where in addition to a specification of the hypotheses entertained, some are endowed with higher priors than others. P&P models endorse a similar conception as some parameters, the unmarked ones, are treated as more equal than others. Thus, while the invariant principles and open parameters delimit the space of G options, markedness theory (or the evaluation metric) is responsible for getting an LAD to specific parameter values on the basis of the available PLD.

This division of labor seems reasonable, but is not apodictic.  There is a trading relation between specifying high priors and delimiting the hypothesis space. Indeed, saying that some option is impossible amounts to setting the prior for this option to 0 and saying that it is necessary amounts to setting the prior to 1.  Moreover, given our current state of knowledge, it is unclear what the difference is between assuming that something is impossible given PLD versus saying that it is very improbable. However, it is not unreasonable, IMO, to divide the problem up as above as several kinds of things really do seem unattested while other things though possible are not required.

With this as background, I want to now turn to a kind of PoS argument that builds on (steals from?) a terrific paper that I’ve recently read by Gigerenzer and Brighton (G&B) (here) and that I have been recommending to all and any in my general vicinity in the last week.

G&B discuss the role of biases in inductive learning. The discussion is under the rubric of heuristics. They note that biases/heuristics have commonly been motivated on grounds of reducing computational complexity. As noted several times before in other posts (e.g. here), many inductive theories are computationally intensive if implemented directly. In fact, so intensive as to be intractable.  I’ve mentioned this wrt Bayesian models and several commentators noted (here) that there are reasons to hope that these problems can be finessed using various well-known (in the sense of well-known to those in the know, i.e. not to me) statistical sampling methods/algorithms. These methods can be used to approximate the kinds of solutions the computationally intractable direct Bayesian methods would produce were they tractable. Let’s call these methods “heuristics.” If correct, this constitutes one good cognitive argument for heuristics; they reduce the computational complexity of a problem making its solution tractable.  As G&B note, on this conception, heuristics (and the biases they incorporate) are the price one has to pay for tractability. Or; though it would be best to do the obvious calculation, such calculations are sadly intractable and so we use heuristics to get the calculations done even though this sacrifices (or might sacrifice) some accuracy for tractability. They call this the accuracy-effort tradeoff (AET). As G&B put it:

If you invest less effort the cost is lower accuracy. Effort refers to searching for more information, performing more computation, or taking more time; in fact these typically go together. Heuristics allow for fast and frugal decisions; thus, it is commonly assumed that they are second best approximations of more complex “optimal” computations and serve the purpose of trading off accuracy for effort. If information were free and humans had eternal time, so the argument goes, more information and computation would always be better (109).

G&B note that this is the common attitude towards heuristics/biases.[7]  They exist to make the job doable. And though G&B agree that this might be one reason for them, they think that it is not the most important helpful feature that heuristics/biases have. So what is?  G&B highlight a second feature of biases/heuristics; what they call the “bias-variance dilemma” (BVD).  They describe it as follows:[8]

… achieving a good fit to observations does not necessarily mean we have found a good model, and choosing a model with the best fit is likely to result in poor predictions…(118).

Why? Because

…bias is only one source of error impacting on the accuracy of model predictions. The second source is variance, which occurs when making inferences from finite samples of noisy data. (119).

In other words, a potentially very serious problem is “overfitting,” a problem that flexible models standardly enjoy. In G&B’s words:

The more flexible the model, the more likely it is to capture not only the underlying pattern but unsystematic patterns such as noise…[V]ariance reflects the sensitivity of the induction algorithm to the specific contents of samples, which means that for different samples of the environment, potentially very different models are being induced.  [In such circumstances NH] a biased model can lead to more accurate predictions than an unbiased model. (119)

Hence the dilemma: To best cover the input data set, “model must accommodate a rich class of patterns in order to insure low bias.” But “[t]he price is an increase in variance, as the model will have greater flexibility, this will enable it to accommodate not only systematic patterns but also accidental patterns such as noise” (119-120). Thus a btter fit to the input may have deleterious effects on predicting future data. Hence the BVD:

Combating high bias requires using a rich class of models, while combating high variance requires placing restrictions on this class of models. We cannot remain agnostic and do both unless we are willing to make a bet on what patterns will occur. This is why “general purpose” models tend to be poor predictors of the future when data are sparse (120).

And the moral G&B draw?

The bias-variance dilemma shows formally why a mind can be better off with an adaptive toolbox of biased specialized heuristics. A single, general-purpose tool with many adjustable parameters is likely to be unstable and incur greater prediction error as a result of high variance. (120)

What consequences might the BVD have for work on language? Well, note first of all that it provides the template for an additional kind of PoS argument. In contrast to the standard one reviewed above, this one holds when we relax the standard idealizations reviewed above; in particular, the assumption that the PLD is noise free and that it is provided all-at-once.  We know that these assumptions are false, what the BVD suggests is that when these are relaxed we potentially encounter another kind of inductive problem in which biases can be empirically very useful. I say “suggests” rather than “shows” because as G&B demonstrate quite nicely, whether the problem is a real one, depends on how sparse and noisy the relevant PLD is. 

The severity of the BVD problem in linguistics will likely depend on the particular linguistic case being studied. So for example, work by Gleitman, Trueswell and friends (discussed here, here, here) suggests that at least early word learning occurs in very noisy data sparse environments. This is just the kind that G&B point to as favor shallow non-intensive data analysis. The procedure that Gleitman, Trueswell and friends argue for seems to fit well into this picture.

I’m no expert in the language acquisition literature, but from what I’ve seen, the scenarios that G&B argue promote BVDs are rife in the wild. I sure looks like many people converge to (very close to) the same G despite plausibly having very different individual inputs (isn’t this the basis for the overwhelming temptation to reify languages?  Believe me my Polish parent English PLD was quite a bit different from that of my Montreal peers and we ended up sounding and speaking very much the same). If so, the kinds of biased systems that GG is very comfortable with will be just what G&B ordered. However, whether this always holds or even whether it ever holds is really an empirical question.[9]

G&B contrasts heuristic systems with more standard models, including Bayesian models, exemplar models, multiple regression models etc. that embody Carnap’s “principle of total evidence” (110). From what G&B say (and I have sort of confirmed by doing econometrician on the campus interviews), it appears that most of the current favored approaches to rational decision making embody this principle, at least as an ideal. As a favorite conceit is to assume that cognitively speaking, humans are very rational, indeed optimal decision makers, Carnap’s principle is embodied in most of the common approaches (indeed Bayesians love to highlight this). Theories that embody Carnap’s principle understand “rational decision making as the process of weighing and adding all information” up to computational tractability.  The phenomena that G&B isolates (what the paper dubs “less is more” effects) challenge this vision. These effects, G&B argues, illustrate that it’s just false that more is always better even in the absence of computational constraints. Rather, in some circumstances, the ones that G&B identifies, shallow and blinkered is the way to go. And if this is correct, then the empirical questions will have to be settled on a case by case basis, sometimes favoring total evidence based models and sometimes not. Further, if this is correct (and if Bayesian models are species of total evidence models) then whether a Bayesian approach is apposite in a given cognitive context becomes an empirical question, the answer depending on how well behaved the data samples are.

Third, it would not be surprising (at least to me) were there two (or at least two) kinds of native FL biases, corresponding to the two kinds of PoS arguments discussed above.  It is possible that the biases motivated via the classical PoS argument (the invariances that circumscribe the class of possible Gs) alone suffice to lead the LAD to its specific G. However, this clearly need not be so.  Nor is it obvious (again at least to me) that the principles that operate within the circumscribed class of grammatically possible grammars would operate as well within the wider class of logically possible ones.  Indeed, when specific examples are considered (e.g. ECP effects, island effects, binding effects) the case for the two-prong attack on the PoS problem seems reasonable. In short, there are two different kinds of PoS problems invoking different kinds of mechanisms.

G&B ends with a description of two epistemological scenarios and the worlds where they make sense.[10] Let me recap them, comment very briefly and end.

The first option has a mind with no biases “with an infinitely flexible system of abstract representations.” This massive malleability allows the mind to “reproduce perfectly” “whatever structure the world has.” This mind works best with “large samples of observations” drawn from world that is “relatively stable.” Because such a mind “must choose from an infinite space of representations, it is likely to require resource intensive cognitive processing.” G&B believes that exemplar models and neural networks are excellent models for this sort of mind. (136)

The second mind makes inferences “quickly from a few observations.” The world it lives in changes in unforeseen ways and the data it has access to is sparse and noisy. To overcome this it uses different specialized biases that can “help to reduce the estimation error.” This mind need not have “knowledge of all relevant options, consequences and probabilities both now and in the future” and it “relies on several inference tools rather than a single universal tool.” Last, in this second scenario intensive processing is not required nor favored.  Rather minds come packed with specialized heuristics able to offset the problems that small noisy data brings with it.

You probably know where I am about to go. The first kind of mind seems more than just a tad familiar from the Empiricist literature. “Infinitely flexible” minds that “reproduce perfectly” “whatever structure the world has” sound like the perfect wax tablets waiting to faithfully receive the contours that the world via “large sample of observations” is ready to structure it with. The second with its biases and specialized heuristics has a definite Rationalist flavor. Such minds contain domain specific operations. Sound familiar? 

What G&B adds to the standard Empiricism-Rationalism discussion is not these two conceptions of different minds, but the kinds of advantages we can expect from each given the nature of the input and the “worlds’ that produce it. When a world is well behaved, G&B observes, minds can be lightly structured and wait for the environment to do its work. When it is a blooming buzzing confusion bias really helps. 

There is a lot more in the G&B paper. I found it one of the more stimulating and thought provoking things I’ve read in the last several years. If G&B is correct, the BVD is rich in consequences for language acquisition models that begin to loosen the idealizations characteristic of Plato’s Problem ruminations. Most interestingly, at least to me, coarsening the idealization adds new reasons for assuming that biological systems come packed with rich innately structured minds. In the right circumstances, they don’t only relieve computational burdens, they allow for good inference, indeed better inference than a mind that more carefully tracks the world and intensively computes the consequences of this careful tracking. Interesting, very interesting. Take a look.




[1] This problem goes back to the very beginning GG, see Stanley Peters’ paper “The Projection Problem: How is a grammar to be selected” in Goals of Linguistic Theory. As he noted in his paper, this problem is closely tied to the question of Explanatory Adequacy. The logic outlined above is very clearly articulated in Peters’ paper. He describes the projection problem  as the “problem of providing a general scheme which specifies the grammar (or grammars) tht can be provided by a human upon exposure to a possible set of basic data” (172).
[2] Note that the projection problem can hold for finite sets as well. The issue is how to select the function that covers the unobserved on the basis of the observed (i.e. how to generalize from a small sample to a larger one). How does a system “project” to the unobserved data based on the observed sample. The infinity assumption allows for a clear example of the logic of projection. It is not a necessary feature.
[3] Peters also zeros in on the idea that PLD is “simple.” As he puts it: “as has often been remarked, one rarely hears a fully grammatical sentence of any complexity…One strategy open to him [the LAD, NH] is to put the greatest confidence in short utterances, which are likely to be less complex than longer ones and thus more likely to be grammatical” (175).
[4] As noted here this assumption quite explicit in Aspects is known to be a radical idealization. However, this does not indicate that it has baleful consequences. It does seem that kids in the same linguistic environment come to acquire very similar competences (no doubt the source of our view that languages exist). This despite the reasonable conjecture that they are not exposed (or intake) exactly the same (kinds of) sentences in the same order. This suggests that order of presentation is not that critical and this is follows from the all-at-once idealization. That said, I return to this assumption below. For some useful discussion see Peters where the idealization is defended (p.175).
[5] Again, see Peters for illuminating discussion.
[6] This overstates the case. The evaluation measure did no tell the LAD how to construct a G given PLD. Rather it specified how to order Gs as better or worse given PLD.  In other words, it specifies how to rank two given Gs. Specifying how to actually build these was considered (and probably still is) too ambitious a goal.
[7] I suspect that the general disdain for priors in Bayesian accounts is the belief that they do not fundamentally alter the acquisition scenario. What I mean by this is that though they may accelerate or impede the rate at which one gets to the best result, over enough time the data will overwhelm the priors so that even if one starts, as it were, in the wrong place in the hypothesis space, the optimal solution will be attained. So priors may affect computation and the rate of convergence to the optimum but it cannot fundamentally alter the destination.
[8] By “fit” here G&B mean fit with the input data sets.
[9] So Jeff Lidz noted that perhaps all LADs enjoy a good number of rich learning encounters where sufficient amounts of the same good data is used.  In other words, though the data overall might stink, there are reliable instances where the data is robust and there are where the acquisition action takes place.  This is indeed possible, it seems to me, and this is what makes the BVD problem an empirical, rather than a conceptual, one.
[10] There are actually three, but I ignore the first as it has little real interest.

Monday, November 4, 2013

Learning is to cognition what phlogiston is to chemistry

Last week John Trueswell gave a colloquium talk at UMD that I unfortunately could not attend. I was in Montreal at a workshop in honor of an old prof of mine, Jim McGilvray.  The Montreal gig was great and it was a pleasure to be able to fete Jim in person (he supervised one of my earliest linguistics projects, (my undergrad thesis on a Reichenbachian theory of tense), but I confess that I would have loved to have been at John’s talk as well. I await the day when some clever physicist figures out how to allow someone to be in two places at once. Until that happy time arrives, I thought it appropriate to do some penance for my physical failings by re-reading a terrific paper by Roediger and Arnold (R&A) on the history of one-trial learning experiments. Why the R&A paper? Because, John’s recent work on lexical acquisition is in the one-trial learning tradition whose history they review. I’ve discussed this work already in a couple of places (here and here) but I wanted to bring your attention to the R&A paper again for it highlights some of the more interesting implications of this line of research for topics near and dear to my intellectual prejudices: despite the common conviction that Empiricism has (at least) something going for it, there is a remarkable absence of evidence supporting this very weak view.[1] Here’s what I mean.

Behind every theory there is an inspirational picture. In the mental sciences, the two grand traditions, Rationalism (R) and Empiricism (E), are animated by two contrasting conceptions of the underlying mechanisms of mental life. For Es, the afflatus is the blank wax tablet, learning consisting of imprinting by experience on this tablet and the clarity and distinctness of the resultant concept/idea being a function of the number of repeated imprints. The more the experience, the deeper and clearer the resulting acquired concept/idea.  Here’s Ebbinghause’s version (quoted in R&A: 129):

These relations [between repetition and performance] can be described figuratively by speaking of the series as being more or less deeply engraved on some mental substratum. To carry out this figure: as the number of repetitions increases, the series are engraved more and more deeply and indelibly; if the number of repetitions is small, the inscription is but surface deep and only fleeting glimpses of the tracery can be caught…

Note that on this conception, repeated experience forms the concepts in the mind (e.g. makes the grooves). Repetition is critical for the mind’s main character is its receptivity to external formative forces, the mind itself being structurally rudimentary. On this view, to understand acquisition requires analyzing the fine structure of the input for what minds/brains do in forming mental constructs (ideas, concepts, etc.) is sift and manipulate these input experiences. It is not surprising, that this view focuses on minds’ significant statistical capacities for these are obvious candidate mechanisms for organizing the inputs and separating the significant wheat from the non-significant chaff.

This contrasts with Rish proposals. For these, the mind is very articulated. There is lots of given pre-experiential structure. Thus, the role of experience is not to construct the relevant concepts attained but to kick start them into activation. Experience on this view is a trigger, not an artificer. Not surprisingly, this conception focuses on discovering the natively provided mental structures that experience serves (importantly but modestly) to activate.

On considering these two different raw philosophical pictures, one can understand the intrinsic interest in one-trial learning (OTL). The existence of OTL would be a problem for E but not for R. The empirical question then is whether OTL exists and how common it is. Investigating this requires translating the philosophical pictures into testable theories, and this leads to learning curves.

R&A observe that one of the biggest pieces of evidence for the E view of the world is the classical learning curve; you know the one that rises from low left to high right decelerating as it goes (as below reproduced from R&A p. 128). 




R&A note that this curve perfectly embodies the E conception that the Ebbinghause quote poetically describes. R&A point out two important features of learning curves consonant with the leading E idea.[2] First, “[t]he fact that the learning curve shows a gradual increase in performance is a reflection of the underlying mechanism- the build up of strength- which is itself also gradual.” And second, that this curve is “the same across astonishingly different experimental situations and dependent measures, as well as across species from slugs to humans,” strongly suggesting general “underlying mechanisms” and general “laws of learning” (129). In a word, this curve, it is argued, puts paid to the R idea of triggering and its concomitant conception of a highly structured mind. If learning curves describe the mechanics of learning, then E beats R. [3] Unless this curve is actually an artifact of, e.g. how experimental data is crunched, rather than a description of an underlying mental mechanism. And that’s where the story that R&A tell gets really interesting.

In the late 1950s and early 1960s Irvin Rock and William Estes (these two were big psych shots, look them up) did a series of experiments that showed (at the very least) that this interpretation of the curve as describing an underlying mechanism that is similarly smooth and incremental is premature, and (at the most) that it was false. They showed two things: (i) that these curves were both consistent with an underlying OTL mechanism (i.e. “although the learning curves were continuous, the underlying processes were anything but continuous” (p. 129)) and (ii) that there was very good evidence that OTL is the norm. Let’s discuss each point separately.

Here’s R&A quoting Rock (that’s the “p.186” below) and then commenting (p. 130) wrt (i):

“Another possibility is that repetition is essential because only a limited number of associations can be formed in one trial, and improvement with repetition is only an artifact of working with long lists of items.” (p.186 ). … That subset is learned perfectly, but all the rest of the associations that were presented  are not learned at all….The “artifact” Rock referred to is essentially that of averaging across many subjects learning many lists on many trials: despite the all-or-none nature of the underlying process, the learning curve will be smooth when performance is averaged over these several parameters. (my emphasis;NH)

This is a very important conceptual point for it divorces the big E conclusion that the mechanisms of learning are gradual and driven by repeated environmental inputs (viz. repeated engravings by experience on a mental substratum) from the fact that learning curves have the shape they do and can be found quite generally across tasks and species. Put more pointedly, if this is correct, then the smooth shape of the learning curve implies nothing at all about the smoothness and gradualness of the underlying mechanism.

Rock and Estes not only made this important observation but also then went on to show that in classical cases of “learning”[4] (i.e. acquiring paired associates), there is good evidence against the classical picture. The basic set up was to have two groups, one that learned listed pairs by going through one list again and again (the control group) and a second that learns a list that removes the non-learned pairs so that they are not encountered again. The prediction if E is correct is that the control group will do better than the second group. I will not review Rock’s and Este’s experiments here as that’s what the R&A paper does so well. Suffice it to say, that, as R&A put it, their papers showed that “there was no hint for the continuity/incremental hypothesis in the data” (p.131).

Note that if (ii) is correct, an account that uses an incremental mechanism to derive acquisition data correctly described by the classical learning curve is incorrect.  Let me beat this horse good and dead: Point (i) shows that a classical learning curve is consistent with OTL mechanisms. Point (ii) argues that OTL mechanisms are in fact what we find. Hence, if correct, theories that deploy incremental mechanisms even if they can derive classical learning curves are wrong. This should not be surprising: such curves graph a correlation between trials and responses. The aim of a theory is not to “model” the data but to “model” the mechanisms that generate the data. The E mistake, is to wrongly infer that smooth incremental data implies smooth incremental mechanisms. It doesn’t, though thinking it does is an Eish diathesis. [5]

It goes without saying that both Rock’s and Estes’ results were contested. Methodological problems were purportedly found that confounded the conclusions. However, and this is interesting, no good evidence for the classical theory appeared to be forthcoming. Rather, the critics seemed satisfied with a draw, viz. showing that the Rock/Estes results need not be interpreted as debunking the classical E view. One particularly cute study that R&A report seemed satisfied with the conclusion “that the incremental theory is untestable” or, quoting the critics directly (Underwood and Keppel): “certain theories are not capable of disproof. Certain aspects of the incremental theory seem to be of this nature” (p. 135). If this be a vindication of the classical E view imagine what a refutation would look like.

John’s current work on lexical acquisition develops the Rock/Estes conception. This is what makes it so interesting for people like me. If they are correct, then classical E conceptions of learning don’t exist, or, more modestly, there is precious little evidence in its favor. To me, this has enormous implications for standard approaches to modeling acquisition that assume some form of gradual process taking place, some form of gradual hill climbing or gradual strengthening of connections. Indeed, in some moods (e.g. now), I think that this work, if correct, implies that Gallistel and Matzel are correct and there is no general theory of learning to be had, as there are no mechanisms, mental or neural, that correspond to what the E picture took learning to be.[6] However, for now I would be happy with more modest conclusions: (i) that there is precious little evidence in favor of the E conception, (ii) that there is little evidence in favor of the view that mental mechanisms are gradual and continuous, and (iii) that there is pretty good evidence that we have mechanisms that enable what amounts to one trial “learning” and that this kind of acquisition requires something very much like the classical R conception of the mind.[7] There is no good reason, in other words, for taking the E conception to be the default and every reason to think that the problem of acquisition is largely one of getting the pre-packaged representational formats correct.

Let me end with a request and an exhortation. First the request: does anyone have a poster case of learning not susceptible to the Rock/Estes critique from the psych literature. It would be nice to have one. Second read the R&A papers and the recent papers developing these ideas by John and Lila and Charles and Jesse and their students. If their insights are internalized, we may finally be able to break the grip of E conceptions of mental mechanisms as the default position. One, at least, can always hope; after all we got rid of phlogiston, didn’t we?





[1] I think it was Lila Gleitman who first advanced the following PoS argument: Empiricism must be innate for what else could explain the widespread conviction that it is true despite the dearth of evidence in its favor. Lila also was kind enough to bring the R&A paper to my attention. I should also add that I doubt that Lila would endorse my interpretation of this work as outlined below. In fact, I am sure she wouldn’t given this.
[2] Gallistel and Matzel (see here) note that the LTP view of brains has been taken to similarly support an E picture. G&M argue that the E picture is widely accepted in the neurosciences, despite there being little to recommend it (and a lot to disavow it). R&A’s discussion merges well with G&M’s and supports a similar conclusion.
[3] R&A identify, in passing, an attraction of E views of learning to the formally inclined that comes from the mathematical tractability of learning curves. I quote: “Learning curves (like forgetting curves) are smooth and beautiful, and psychologists with a mathematical bent can have a field day fitting equations to them” (p.128).  One should never underestimate the attraction of a conclusion that fits snugly with your available technology. 
[4] Note the scare quotes: if Rock and Estes (and Gleitman and Trueswell and Company are right) then learning is a hypothesis about the mechanisms of acquisition, not a neutral description of an observed phenomenon. What we observe is change over time given environmental inputs. This change may be due to learning, maturation, growth, or whatever. The cognitive question concerns the mechanism and learning is a proposal for one such.
[5] This is the kind of mistake that modeling which takes the name of the game to be getting the input/output relations right is particularly susceptible to. E conceptions are prone to this kind of misconception given their picture that mental structure mirrors the structure of the input. However, modeling I/O relations confuses the data to be explained for the mechanism that does the explanation. For a related point in a Bayesian context see Glymour on “Osiander’s Psychology” in the comments to Jones & Love’s discussion of modern Bayesianism here and some blogish discussion by me (here). J&L note a similarity between earlier behaviorist conceptions and some modern Bayesian analyses. The above suggests how two programs that appear so different on the surface might nonetheless lead to the same conceptual place via a shared partiality to associationism and/or a misunderstanding of what modeling is supposed to do.
[6] After all, if (roughly) one trial “learning” is the norm then there is little for E like mechanisms to do. Acquisition on this view is more akin to transduction than to mental computation. I would probably be satisfied if it turned out that there was “a little” learning, but this is a topic for another discussion.
[7] After all if we don’t acquire knowledge by carefully sifting through the input it’s because such sifting is not necessary and it wouldn’t be necessary if the knowledge is basically, already, all there.