Comments

Showing posts with label PLD. Show all posts
Showing posts with label PLD. Show all posts

Wednesday, February 14, 2018

The significant gaps in the PoS argument

I admit it, the title was meant to lure some in with the expectation that Hornstein was about to recant and confess (at last) to the holiness (as in full of holes) of the Poverty of Stimulus argument (PoS). If you are one of these, welcome (and gotcha!). I suspect, however, that you will be disappointed for I am here to affirm once again how great an argument form the PoS actually is, and if not ‘holy’ then at least deserving of the greatest reverence. The remarks below are prompted by an observation in a recent essay by Epstein, Kitahara and Seely (EKS’s) (here p. 51, emphasis mine):

…Recognizing the gross disparity between the input and the state attained (knowledge) is the first step one must take in recognizing the fascination surrounding human linguistic capacities. The chasm (the existence of which is still controversial in linguistics, the so-called poverty of the stimulus debate) is bridged by the postulation of innate genetically determined properties (uncontroversial in biology)…

This is a shocking statement! And what makes it shocking is EKS’s completely accurate observation that many linguists, psychologists, neuroscientists, computer scientists and other language scientists still wonder whether there exists a learning problem in the domain of language at all. Yup, after more than 60 years of, IMO, pretty conclusive demonstrations of the poverty of the linguistic stimulus (and several hundred years of people exploring the logic of induction) the question of whether such a gap exists is still not settled doctrine. To my mind, the only thing that this could mean is that skeptics really do not understand what the PoS in the domain of language (or any other really) is all about. If they did, it would be uncontroversial that a significant gap (in fact, as we shall see several) exists between evidence available to fix the capacity and the capacity attained. Of course, recognizing that significant gaps exist does not by itself suffice to bridge them. However, unrecognized problems are particularly difficult to solve (you can prove this to yourself by trying to solve a couple of problems that you do not know exist right now) and so as a public service I would like to rehearse (one more time and with gusto) the various ways that the linguistic input (aka, Primary Linguistic Data (PLD)) to the language acquisition device (LAD (aka child)) underdetermines the structure of the competence attained (knowledge of G that a native speaker has). The deficiencies are severe and failure to appreciate this stems from a misconception of what G acquisition consists in. Let’s review.

There are three different kinds of gaps.

The first and most anodyne relates to the quality of the input. There are several ways that the quality might be problematic. Here are some.

1.     The input PLD is in the form of uttered bits of language. This gap adverts to the fact that there is a slip betwixt phrases/sentences (the structures that we know something about) and the lip(s that utter them). So the PLD in the ambient environment of the LAD is not “perfect.” There are mispronunciations, half thoughts badly expressed, lots of ‘hmms’ and ‘likes’ thrown in for no apparent purpose (except to irritate parents), misperceptions leading to misarticulations, cases of talking with one’s mouth full, and more. In short, the input to the LAD is not ideal. The input data are not perfect examples of the extensional range of sentences and phrases of the language.
2.     The range of input data also falls short. Thus, since forever (Newport, Gleitman and Gleitman did the leg work on this over 30 years ago (if not more)) it has been pointed out that utterances addressed to LADs are largely in imperative or interrogative form. When talking to very young kids, native speakers tend to ask a lot of questions (to which the asker already knows the answer so it is not a “real” question) and issue a log of commands (actually many of the putative questions are also rhetorically disguised commands). Kids, in contrast, are basically declarative utterers. In other words, kids don’t grow up talking motherese, though large chunks of the input has a very stylized phon/syntactic contour (at least in some populations). They don’t sound mothereesish and they eschew use of the kinds of sentences directed at them. So even at a gross level, the match between what LADs hear in their ambient environment and what they do mismatches.

So, the input is not perfect and the attained competence is an idealized version of what the LAD actually has access to. Even the Structuralists appreciated this point, knowing full well that texts of actual speech needed curation to be taken as useful evidence for anything. Anyone who has ever read a non-edited verbatim text of an interview knows that the raw uttered data can be pretty messy. As indeed are data of any kind. The inputs vary in the quality of the exemplar, some being closer approximations to the ideal than others. Or to put this another way: there is a sizeable gap between the set of sentences and the set of utterances. Thus, if we assume that LADs acquire sentential competence, there is a gap to be traversed in building the former set from the latter.

The second gap between input and capacity attained is decidedly more qualitative and significant.  It is a fact that native speakers are linguistically creative in the sense of effortlessly understanding and producing linguistic objects never before experienced. Linguistic creativity requires that what an LAD acquires on the basis of PLD is not a list of sentences/phrases previously encountered in ambient utterances but a way of generating an open ended list of acceptable sentences/phrases. In other words, what is acquired is (at least) some kind of generative procedure (rule system or G) that can recursively specify the open ended set of linguistic objects of the native speaker’s language. And herein lies the second gap: the outputs of the acquisition process is a G or set of rules while the input is products of this G or set of rules AND products of rules and rules (or sentences and the Gs that generate them) are ontologically different kinds of objects. Or to put this another way: an LAD does not “experience” Gs directly but only via their products and there is a principled gap between these products and the generative procedures that generate them. Furthermore, so far as I know nobody has ever shown how to bridge this ontological divide by, say, using conventional analytical methods. For example, so far as I know the standard substitution methods prized by the transitional probability crowd has yet to converge on the actual VP expansion rule (one that includes ate, ate a potato, ate a potato that I bout at the store, ate a potato that I bought at the store that is around the block, ate a potato that I bought at the store that is around the block near the gin joint whose owner has a red buggy which people from Detroit want to buy for a song, etc. All of these are VPs but so far that they are all VPs is not something that artificial G learners have managed to cover. Recursion really is a pain).

In fact, it is worse than this. For any finite set of data specified extensionally there are an infinite number of different functions that can generate those data. This observation goes back to Hume, and has been clear to anyone that has thought about the issue for at least the last several hundred years. Wittgenstein made this point. Goodman made it. Fodor made it. Chomsky made it. Even I, standing on the shoulders of giants) have been known to make it. It is a very old point. The data (always a finite object) cannot speak for themselves in the sense of specifying a unique function that generate those data. Data cannot bootstrap a unique function. The only way to get from data to the functions that generate them is to make concrete assumptions about the specific nature of the induction and in one way or other, this means specifying (listing them, ranking them) the class of admissible functional targets. In the absence of this, data cannot induce functions at all. As Gs or generative procedures just are functions, there is a qualitative irreducible ontological difference between the PLD and the Gs that the LAD acquires. There is no way to get from any data to any function (induce any G from any finite PLD) without specifying in some way the range of potential candidate Gs. Or, to say this another way, if the target is Gs and the input is a finitely specified list of data, there is no way of uniquely specifying the target in terms of the properties of the data list alone.

All of which leads to consideration of a third gap: the evidence required for choosing among plausible competing Gs to which native speakers converge is underdetermined by the PLD available for deciding among competing Gs. So, not only do we need some way of getting native speakers from data to functions that generate that data, but even given a list of plausible options, there exists little evidence in the PLD itself for choosing among these plausible functions/Gs. This is the problem that Chomsky (and moi) has tended to zero in on in discussions of PoS. So for example, a perfectly simple and plausible transformation would manipulate inputs in virtue of their string properties, another in virtue of their hierarchical properties. We have evidence that native speakers do the latter, not the former. Evidence for this conclusion exists in the wider linguistic data (surveyed by the linguist and including complex cases and unacceptable data) but not the Primary linguistic data the child as access to (which is typically quite simple and well formed). Thus, the fact that LADs induce along structure dependent lines is an induction to a particular G from a given list of possible Gs with no basis in the data justifying the induction. So not only do humans project and not only do they do so uniformly, they do so uniformly in the absence of any available evidence that could guide this uniform projection. There are endlessly many qualitatively different kinds of inductions that could get you from PLD to a G and native speakers project in the same way despite no evidence in the PLD constraining this projection. The gap between plausible Gs and the PLD is strongly underdetermined.

It is worth observing that this gap does not require that the set of grammatical objects be infinite (though making the argument is a whole lot easier if something like this is right). The point is that native speakers make systematic conclusions about novel data. That they conclude anything at all implies something like a rule system or generative procedure. That native speakers largely do so in the same way suggests that they are not (always) just guessing.[1] The systematic projection from the data provided in the input to novel examples implies that LADs generalize in the same way from LAD input. Thus, native speakers project the same Gs from finite sample examples of those Gs. And that is the problem: how do native speakers bridge the gap between samples of Gs and the Gs of which they are samples in the same way despite the fact that there are infinitely many ways to project from a finite samples to functions that generate those samples. Answer: the projection is natively constrained in some way (plug in your favorite theory of UG here). The divide between sample data and Gs that generate them is further exacerbated by the fact that native speakers converge on effectively the same kinds of Gs despite having precious little evidence in the PLD for making a decision among plausible Gs.

So there are three gaps: a gap between utterances and sentences, a gap between sentences/phrases and the Gs that generate them and a gap between the Gs that are projected and the evidence to choose among these Gs in the PLD the child has access to while fixing these Gs. Each gap presents its own challenges, though it is the last two that are, IMO, the most serious. Were the problem restricted to getting from exemplars to idealized examples (i.e. utterances to sentences) then the problem would be solvable by conventional statistical massaging. But that is not the problem. It is at most one problem, and not the big one. So are there chasms that must be jumped, and gaps that must be filled? Yup. And does jumping them/filling them the way we do require something like native knowledge? Yup. And will this soon be widely accepted and understood? Don’t bet on it.

Let me end with the following observation. Randy Gallistel recently gave a talk at UMD observing that it is trivial to construct PoS arguments in virtually every domain of animal learning. The problem is not limited to language or to humans. It saturates the animal navigation literature, the animal foraging literature, and the animal decision literature. The fact, as he pointed out, is that in the real world animals learn a whole lot from very little while Eish theories of learning (e.g. associationism, connectionism, etc) assume that animals learn very little from a whole lot. If this is right, then the main problem with Eish theories (which as EKS note still (sadly) dominate the mental and neural sciences and which lie behind the widespread skepticism concerning PoS as a basic fact of mental life in animals) is that the way that it frames the problem of learning has things exactly backwards. And it is for this reason that Es have a hard time spotting the gaps. The failure of Eism, in other words, is not surprising. If you misconceive the problem your solutions will be misconceived as well.



[1] The ‘always’ is a nod to interesting work by Lidz and friends that suggests that sometimes this is exactly what LADs do.

Monday, January 11, 2016

An interesting PoS argument makes it to the "show"

Linguists tend to publish for other linguists. And this is fine. However, it was not always so. There was a time when linguists saw themselves as part of a larger cog-psy community and published in venues frequented by non-linguists. Cognition was a terrific venue for such work and it enabled linguistic discoveries to influence debates about the nature of mind (and, occasionally, even the brain). However, even in these golden years very few linguists published in the leading general science journals and this had the effect of segregating our work from the scientific mainstream. Books like Pinker’s The Language Instinct were effective conduits to the larger scientific community, but really nothing gains scientific street cred like publishing in the big three peer reviewed high impact journals like Science, Nature and PNAS. Moreover, as readers of FoL know, I believe that the single best way to politically advance linguistics and protect it economically is to disseminate our results to our fellow scientists. So, with this as prelude, I am delighted to note a paper that has yesterday appeared in PNAS of that ilk. The paper (here) has three authors: Chung-hye Han, Julien Musolino and Jeff Lidz (HML). Aside from being very GFL (i.e. “good for linguistics”) that such things are being published in PNAS it’s also a very good paper that I heartily recommend you take a look. It even has the virtue of being a mere 6 pages (a page limit we should encourage our own journals to try to approximate). So what’s in it? Here is a quick précis with some comments.

The paper argues that FL requires that speakers adopt (at least in the unmarked case) a single G when exposed to PLD. The reasoning for this conclusion is based on a novel Poverty of Stimulus (PoS) argument. What makes it novel is that the paper outlines how a particular kind of variation can be driven by properties of FL. Let me explain.

The standard PoS argument looks at invariances across Gs and shows how these can be accounted for with some or another proposed innate feature of FL. Variation among Gs is then attributed to the differential inductive effects of the Primary Linguistic Data (PLD). What HML shows is that this same logic allows FL to accommodate some G variation just in case the PLD is insufficient to fix a G parameter in the LAD (aka language acquisition device (i.e. kid)). In such cases, if FL requires an LAD to construct a single G for given PLD, then we expect to find variable Gs in a population of speakers with the following three key features: (i) Speakers exhibit variability wrt a certain set of (relatively recondite) grammatical phenomena, (ii) This variability is attested between but not within speakers, and (iii) The variability is independent between speakers and parents. Let me say a word about each point.

As regards (i), this is the fact HML discovers (actually this possibility is first described in an earlier 2007 HML paper). HML shows that in Korean the height of the verb explains scope of negation effects. These effects are quite obscure and verb height cannot be induced from the inspection of surface forms as they can be in languages like French and English. In effect HML shows argues (IMO, shows) that “children’s acquisition of this knowledge [viz. the scope facts, NH]…is not determined by any aspect of experience…because the experience of the language learner does not contain the necessary environmental trigger” (1).

As regards (ii), HML shows that the variation is consistent within speakers across different negative constructions and over time. In other words, once a LAD’s G fixes the position of a verb (and with it negation) it fixes it there consistently.

As regards (iii), the G variation in the population is effectively random. It is not possible to predict any speaker’s positioning of the verb by examining the Gs of parents (or, for that matter, anyone else). In other words, as there is no data that could fix where the verb sits in a Korean G then the fact that it gets fixed is a product of the structure of FL and so we expect random variation.

This is a very clever argument. Note, that it directly supports the logic of general PoS arguments without assuming G invariance of the output. Or, to put this another way: PoS arguments generally proceed from invariant properties of Gs to features of FL. HML notes that it is consistent with PoS logic that there be variation so long as it is random. The fact that one can find such cases further strengthens PoS logic.

The paper has two other virtues, IMO, more directly relevant to syntactic theory.

First, it provides a model of the kinds of things that syntacticians keep asking for. Syntacticians keep asking whether psycho-ling results can help choose between various alternative syntactic proposals. In principle the answer is, of course, yes. However, paradigms of this are hard to find. HML provides an example where a psycho-ling result could be close to dispositive.  The relevant syntactic alternatives hail from the earliest days of Minimalism when V raising was a hot topic of inquiry.

Cast your mind back to the earliest days of the Minimalist Program (MP), indeed all the way back to 1995 and the Black Book. Chapters 2 and 3 presented alternative theories of V raising. The chapter 2 theory (see 138ff) basically provided a theory in which French (where Vs overtly raise to T) is the unmarked case and English (where V does not overtly raise) is the marked one. The markedness contrast is argued to be a reflection of a leading MP idea (viz. economy). The idea is that overt raising involves fewer operations than overt lowering plus LF raising so an English G is less economical than a French one as a matter of FL principle (not UG incidentally, but FL) for it uses more elaborate derivations in getting V to T.

The cost accounting changes in chapter 3 (roughly 195-198) where English “procrastinating” Gs are the unmarked case. The idea is that LF operations without PF effects are more economical than ones that result in PF “deletions” (an idea, btw, that lingers to the present in current MP accounts that take Gs to relate meanings with sound). This requires rethinking how syntactic morphemes are licensed (checking) and how they enter derivations (fully featurally encumbered), but given this and some other assumptions French overt raising in the chapter 3 theory is less grammatically svelte than covert V to T, and hence less preferred.

HML bears directly on these two theories and argues that they are both wrong.  Were the chapter 2 story right then we would expect that all Korean Gs assigned Korean Vs high positions. Were the chapter 3 theory right they would all have low positions. The fact that both are available and equally so argues that neither option is better than the other. This leaves open the question of what theory allows both Gs to be equally available (see below). I suspect that a single cycle theory with copy deletion can be made to work, but who knows. Now that HML has shown that both are equally fine, we know that neither of the earlier stories can possibly be correct. Note that this does not mean that the HML account is inconsistent with any markedness view of V raising. This brings me to the second virtue of the HML story. It raises an interesting theoretical question.

So far as I know, there is currently no G story for why it is that LADs need choose a singly G for V raising. The data are quite clear that this is what happens, but why this is required is theoretically unclear. Thus, HML raises an interesting grammatical question: what is it about FL/UG that forces a choice? I can imagine some answers having to do with how complex the lexicon is and that having functional heads that optionally assign features to V in one of two ways is more costly than having one that does it only one way. This might have the desired effects if judiciously worked out. However, this then predicts that some Gs, mixed Gs, will be more costly than uniform Gs. This, in effect, makes English the marked case again given the fact that English Gs raise be and have (and maybe modals) but not more “lexical” verbs.[1] At any rate, none of this is a theory, but the HML data raise an interesting theoretical question as well as closing off two reasonable prior alternatives. So it serves as a nice example of how psycho work can impact syntactic theory.

Let me end with one more point. Unlike much publicity concerning linguistics, the HML work offers an excellent example of what linguistics has achieved. This exploits real linguistic advances to make its scientifically interesting point. And this is in contrast to lousy ways of advertising our linguistic wares. One reaction to the invisibility of linguistics in the general scientific culture has been to try to co-opt anything “languagy” to promote linguistics. The word of the year competition at the LSA is an excellent (sad) example.[2]  The idea seems to be that this kind of thing garners media attention and that there is no such thing as bad publicity. I could not disagree more. The word of the year has nothing to do with linguistics, nothing to do with the serious advances GG has made, and relies on no expert/professional knowledge that linguistics brings to the scientific table. As such it does nothing to advertise our scientific bona fides. It’s, IMO, crap.  And using it to advance the visibility of linguistics is both counterproductive and dishonest. I don’t know about you, but scientific overreach (aka, scientism) makes my teeth hurt. This should not be what a professional linguistics organization (the LSA) should be doing to promote linguistics.  What should it be doing? Advertising work like HML, i.e. making this kind of work more widely accessible to the general scientific community. This is what I had hoped the LSA initiative (noted here) was going to do. To date, so far as I can tell, this hope has not been realized. Instead we get words of the year and worse (see here). It’s almost like the LSA is embarrassed by work in real linguistics. Too bad, for as HML indicates, it can sell well.

So, read the HML paper and advertise it to scientific colleagues outside linguistics proper. It is both interesting in itself and good publicity for what we do. It’s real linguistics with a broader reach.




[1] And this might predict that mixed Gs will necessarily only allow Vs robustly indicated in the data to be “different,” (e.g. raise). This is consistent with what we find in English where be and have are pretty PLD robust. It would be interesting to crank these cases through Charles Yang’s forthcoming learner and see what the limits on exceptionality would be for such a story.
[2] Thx to Alexander Williams for some co-venting about this.