Comments

Tuesday, January 12, 2016

Big Data Considered Harmful

An immediate qualification is order. No one in their right mind would claim data is a bad thing. Naturally, we can find out more about language with more data, even though everyone who has worked with a corpus, regardless of its size, often wishes that the pattern we are really interested weren’t so sparsely represented. All the same, spurious statistical correlations (with excellent p-values) are easier and easier to come by with more and more data. And even when they are not spurious, it’s by no means interesting or worth explaining. We all have our favorite examples, from high altitudes begetting ejectives to 37% of English words are nouns. 

These are just the usual caveats with data, so bigger caveats come with bigger data. But I wish to make a stronger point: When language is considered in a psychological setting, it pays to discard a good portion of the data. In fact, Big Data does serious harm to the child so it may even render language unlearnable. Upon reflection, the reasons are obvious but we do need get our hands dirty—with data.

As I noted in these pages, even very young children have a systematic grammar, at least in certain respects.  For instance, a fully productive rule “NP→ D N” suggests that the determiner D (the and a/n) can be interchangeably used with singular nouns (N). In numerous corpus analyses, children produce fairly low values of combinatorial diversity: typically only 20-40% of nouns that appear with either determiner are paired with both, giving rise to the impression that young children do not have abstract rules but rely on memorizing lexically specific combinations from adult input (e.g., Tomasello). Yet a rigorous statistical test shows that even very young children produce the level of diversity which, while low, is expected under a categorical rule that independently combines determiners and nouns.

But let open is how children learn the NP rule. Consider Adam, a boy studied by Roger Brown. In his speech transcripts, Adam produced 3,729 determiner-noun combinations with 780 distinct nouns. Of these only 32.2% appeared with both determiners, which is similar to the expected value of 33.7% under the abstract NP rule. Adam’s mother, who was recorded in the same corpus, produced a diversity measure of 30.3% out of 914 nouns. Even among the 469 nouns used at least twice, which provided opportunities to be used with both determiners, only over half (260) did so. To learn the NP rule, then, Adam must, and apparently did at a very young age, generalize from a small subset of nouns with attested interchangeable determiners to all nouns. 

What could account for this massive leap of faith? Only if Adam attended to a small amount of data. The developmental literature offers the idea of “less is more” (Newport 1990, Elman 1993; see a toy example on word learning here): the maturational constraints place a limit on the processing capacity of young children which may turn out beneficial for language acquisition. Under the sparsity of language distribution, the acquisition of the determiner rule (and by extension, any rules) is impossible if the learner required evidence from all or even most of the participating units. Indeed, if the child can only retain and learn from the most frequent items, the odds of acquiring productive rules improves considerably. Furthermore, under a general law of learning (dubbed the Tolerance Principle), it is easier for rules that operate over a small class of items to clear the productivity threshold. Specifically, theTolerance Principle says that a generalization that could hold over N items cannot tolerate more than θN = N/ln N negative or unattested examples. Small Ns are more tolerant. 

Consider the learning of the NP rule again, except this time he ignores most of what his mother says. If Adam were only to learn from the top 50 most frequent nouns, he would notice that almost all of them — 43 to be precise — are paired with both determiners. On this much small subset of data where N = 50, there is sufficient evidence for generalization: the 7 nouns that appear exclusively with only one determiner are below the tolerance threshold θ50 = 12. For the top N = 100 most frequent nouns, 83 are paired with both determiners: the 17 loners are again below the tolerance threshold θ100 = 22. The effective vocabulary size of children at the age when they show productivity of the DP rule cannot exceed a few hundred words (Fenson et al. 1994). They must have acquired the rule on a very small set of high frequency nouns, almost of which will show interchangeability with both determiners. 

Such examples are abundant in language. In a recent paper, I argue that adjective-like words such as asleep cannot appear attributively because there is robust distributional evidence that links them to PPs. English has at least 41 such adjective-like items, with similar properties: 

abeam, ablaze, abloom, abuzz, across, adrift, afire, aflame, afraid, agape, aghast, agleam, aglitter, aglow, aground, ahead, ajar, akin, alight, alike, alive, alone, amiss, amok, amuck, apart, aplenty, around, ashamed, ashore, askew, aslant, asleep, astern, astir, atilt, awake, aware, awhirl, away, awash

In a six million words of child-directed English, roughly a year of speech (for some children), only 8 adjectives are attested in the crucial syntactic context that identifies them as PP-like. Now if all 41 items are represented in the corpus, there is no way for the child to succeed. Fortunately, only 12 made an appearance at all: from 8 to 12 is easily sanctioned but from 8 to 41? Forget about it. Even we have a gigantic data set, it’s highly unlikely that a majority of the 41 will bear the PP-like signature. Conclusion: our knowledge of the 41 is really based on a rule derived on the 12. 


In a world dominated by Zipf’s Law, the genuine regularities of language must be detectable in the left region of rank/frequency. The long tail is essentially noise: it can only overwhelm the signal and should be discarded. No wonder kids pretty much ignore what their parents say.

A very useful new book on evolang

For those that don't yet know, there is a terrific new little book by berwick and Chomsky on language and evolution. I have a blurb on the book recommending it where I say the following:

“Nothing talks like humans do. Nothing even comes close. This sets up an interesting evolutionary problem: how did this unique capacity arise in the species? Unfortunately, approaching this question intelligently requires combining skills that seldom travel in tandem. Linguists know a lot about the principal features of human language but little about how evolution works, and biologists know a lot about how evolution works but little about the distinctive properties of human language. Enter Berwick and Chomsky’s marvelous little book. In a mere four lucid and easily accessible chapters they educate linguists about the central mechanisms driving evolution and bring biologists up to date on the key distinctive features of natural language. Anyone interested in this topic must read this book.”

Nor am I the only open who enjoyed it. Others more educated than me loved it too; Ian Tatersall, Martin Nowak and Stephen Crain. Tatersall. So give yourself a post new year's gift and buy and read the book.

Monday, January 11, 2016

An interesting PoS argument makes it to the "show"

Linguists tend to publish for other linguists. And this is fine. However, it was not always so. There was a time when linguists saw themselves as part of a larger cog-psy community and published in venues frequented by non-linguists. Cognition was a terrific venue for such work and it enabled linguistic discoveries to influence debates about the nature of mind (and, occasionally, even the brain). However, even in these golden years very few linguists published in the leading general science journals and this had the effect of segregating our work from the scientific mainstream. Books like Pinker’s The Language Instinct were effective conduits to the larger scientific community, but really nothing gains scientific street cred like publishing in the big three peer reviewed high impact journals like Science, Nature and PNAS. Moreover, as readers of FoL know, I believe that the single best way to politically advance linguistics and protect it economically is to disseminate our results to our fellow scientists. So, with this as prelude, I am delighted to note a paper that has yesterday appeared in PNAS of that ilk. The paper (here) has three authors: Chung-hye Han, Julien Musolino and Jeff Lidz (HML). Aside from being very GFL (i.e. “good for linguistics”) that such things are being published in PNAS it’s also a very good paper that I heartily recommend you take a look. It even has the virtue of being a mere 6 pages (a page limit we should encourage our own journals to try to approximate). So what’s in it? Here is a quick précis with some comments.

The paper argues that FL requires that speakers adopt (at least in the unmarked case) a single G when exposed to PLD. The reasoning for this conclusion is based on a novel Poverty of Stimulus (PoS) argument. What makes it novel is that the paper outlines how a particular kind of variation can be driven by properties of FL. Let me explain.

The standard PoS argument looks at invariances across Gs and shows how these can be accounted for with some or another proposed innate feature of FL. Variation among Gs is then attributed to the differential inductive effects of the Primary Linguistic Data (PLD). What HML shows is that this same logic allows FL to accommodate some G variation just in case the PLD is insufficient to fix a G parameter in the LAD (aka language acquisition device (i.e. kid)). In such cases, if FL requires an LAD to construct a single G for given PLD, then we expect to find variable Gs in a population of speakers with the following three key features: (i) Speakers exhibit variability wrt a certain set of (relatively recondite) grammatical phenomena, (ii) This variability is attested between but not within speakers, and (iii) The variability is independent between speakers and parents. Let me say a word about each point.

As regards (i), this is the fact HML discovers (actually this possibility is first described in an earlier 2007 HML paper). HML shows that in Korean the height of the verb explains scope of negation effects. These effects are quite obscure and verb height cannot be induced from the inspection of surface forms as they can be in languages like French and English. In effect HML shows argues (IMO, shows) that “children’s acquisition of this knowledge [viz. the scope facts, NH]…is not determined by any aspect of experience…because the experience of the language learner does not contain the necessary environmental trigger” (1).

As regards (ii), HML shows that the variation is consistent within speakers across different negative constructions and over time. In other words, once a LAD’s G fixes the position of a verb (and with it negation) it fixes it there consistently.

As regards (iii), the G variation in the population is effectively random. It is not possible to predict any speaker’s positioning of the verb by examining the Gs of parents (or, for that matter, anyone else). In other words, as there is no data that could fix where the verb sits in a Korean G then the fact that it gets fixed is a product of the structure of FL and so we expect random variation.

This is a very clever argument. Note, that it directly supports the logic of general PoS arguments without assuming G invariance of the output. Or, to put this another way: PoS arguments generally proceed from invariant properties of Gs to features of FL. HML notes that it is consistent with PoS logic that there be variation so long as it is random. The fact that one can find such cases further strengthens PoS logic.

The paper has two other virtues, IMO, more directly relevant to syntactic theory.

First, it provides a model of the kinds of things that syntacticians keep asking for. Syntacticians keep asking whether psycho-ling results can help choose between various alternative syntactic proposals. In principle the answer is, of course, yes. However, paradigms of this are hard to find. HML provides an example where a psycho-ling result could be close to dispositive.  The relevant syntactic alternatives hail from the earliest days of Minimalism when V raising was a hot topic of inquiry.

Cast your mind back to the earliest days of the Minimalist Program (MP), indeed all the way back to 1995 and the Black Book. Chapters 2 and 3 presented alternative theories of V raising. The chapter 2 theory (see 138ff) basically provided a theory in which French (where Vs overtly raise to T) is the unmarked case and English (where V does not overtly raise) is the marked one. The markedness contrast is argued to be a reflection of a leading MP idea (viz. economy). The idea is that overt raising involves fewer operations than overt lowering plus LF raising so an English G is less economical than a French one as a matter of FL principle (not UG incidentally, but FL) for it uses more elaborate derivations in getting V to T.

The cost accounting changes in chapter 3 (roughly 195-198) where English “procrastinating” Gs are the unmarked case. The idea is that LF operations without PF effects are more economical than ones that result in PF “deletions” (an idea, btw, that lingers to the present in current MP accounts that take Gs to relate meanings with sound). This requires rethinking how syntactic morphemes are licensed (checking) and how they enter derivations (fully featurally encumbered), but given this and some other assumptions French overt raising in the chapter 3 theory is less grammatically svelte than covert V to T, and hence less preferred.

HML bears directly on these two theories and argues that they are both wrong.  Were the chapter 2 story right then we would expect that all Korean Gs assigned Korean Vs high positions. Were the chapter 3 theory right they would all have low positions. The fact that both are available and equally so argues that neither option is better than the other. This leaves open the question of what theory allows both Gs to be equally available (see below). I suspect that a single cycle theory with copy deletion can be made to work, but who knows. Now that HML has shown that both are equally fine, we know that neither of the earlier stories can possibly be correct. Note that this does not mean that the HML account is inconsistent with any markedness view of V raising. This brings me to the second virtue of the HML story. It raises an interesting theoretical question.

So far as I know, there is currently no G story for why it is that LADs need choose a singly G for V raising. The data are quite clear that this is what happens, but why this is required is theoretically unclear. Thus, HML raises an interesting grammatical question: what is it about FL/UG that forces a choice? I can imagine some answers having to do with how complex the lexicon is and that having functional heads that optionally assign features to V in one of two ways is more costly than having one that does it only one way. This might have the desired effects if judiciously worked out. However, this then predicts that some Gs, mixed Gs, will be more costly than uniform Gs. This, in effect, makes English the marked case again given the fact that English Gs raise be and have (and maybe modals) but not more “lexical” verbs.[1] At any rate, none of this is a theory, but the HML data raise an interesting theoretical question as well as closing off two reasonable prior alternatives. So it serves as a nice example of how psycho work can impact syntactic theory.

Let me end with one more point. Unlike much publicity concerning linguistics, the HML work offers an excellent example of what linguistics has achieved. This exploits real linguistic advances to make its scientifically interesting point. And this is in contrast to lousy ways of advertising our linguistic wares. One reaction to the invisibility of linguistics in the general scientific culture has been to try to co-opt anything “languagy” to promote linguistics. The word of the year competition at the LSA is an excellent (sad) example.[2]  The idea seems to be that this kind of thing garners media attention and that there is no such thing as bad publicity. I could not disagree more. The word of the year has nothing to do with linguistics, nothing to do with the serious advances GG has made, and relies on no expert/professional knowledge that linguistics brings to the scientific table. As such it does nothing to advertise our scientific bona fides. It’s, IMO, crap.  And using it to advance the visibility of linguistics is both counterproductive and dishonest. I don’t know about you, but scientific overreach (aka, scientism) makes my teeth hurt. This should not be what a professional linguistics organization (the LSA) should be doing to promote linguistics.  What should it be doing? Advertising work like HML, i.e. making this kind of work more widely accessible to the general scientific community. This is what I had hoped the LSA initiative (noted here) was going to do. To date, so far as I can tell, this hope has not been realized. Instead we get words of the year and worse (see here). It’s almost like the LSA is embarrassed by work in real linguistics. Too bad, for as HML indicates, it can sell well.

So, read the HML paper and advertise it to scientific colleagues outside linguistics proper. It is both interesting in itself and good publicity for what we do. It’s real linguistics with a broader reach.




[1] And this might predict that mixed Gs will necessarily only allow Vs robustly indicated in the data to be “different,” (e.g. raise). This is consistent with what we find in English where be and have are pretty PLD robust. It would be interesting to crank these cases through Charles Yang’s forthcoming learner and see what the limits on exceptionality would be for such a story.
[2] Thx to Alexander Williams for some co-venting about this.

Monday, January 4, 2016

Sorry!

Too much wine over the holidays. I forgot to include the link to Colin's original post in the previous one. I now fixed that. So if you intended to read it but could not find it, go back one post and try again. Happy New Year.

An experiment in open access publishing

Colin Phillips has a very interesting discussion of his three year experiment with an open access journal that he, Matt Wagers and Claudia Felser edited in the Frontiers series (here). The online journal (here) has been quite successful and Colin does something very important and timely; he reflects on what went well, what went less well and WHY. In other words, he brings some first hand empirical experience to bear on a topic that we have discussed at FoL. The whole discussion is very interesting and I strongly recommend it for those interested in the topic.

A key point that Colin makes is that there are virtues other than cost in assessing how a journal is contributing to inquiry; readership, time from submission to publication, Impact Factor all matter, in addition to cost. Moreover, he notes that, interestingly, these are all factors that can be managed better or worse depending on what seem to be easily implementable procedures.

One that seems particularly noteworthy is the apparent fact that most of the submissions are accepted. This is not entirely true for there is a per-submission vetting process that insures that most of the submissions are of a kind to be accepted. However, the fact remains that a large proportion of the papers submitted get into print. The main reason seems to be because "[a]rticles are judged only for soundness, not for impact, i.e. if your study is sound but has minimal novelty or importance it can still be accepted." This makes the reviewing process less contentious and the goal posts easier to identify and hence cuts down on reviewing gamesmanship (though that is not how Colin puts it).

Let me make one observation and again encourage you to read the whole post because it is very good.  One of the downsides of the current review process, IMO, is that it penalizes originality. How so? Well, original work is by its nature contentious and less well-formed than less original work. This is especially so when it comes to theory, where making things clear is very hard and then it is harder still to accommodate all the myriad problems that novelty will face. One way of finessing this problem is to publish most everything. This is what the above journal does (well, depending on how 'soundness' is evaluated: what makes a paper sound?). Another way is to to go the way that the old Cognition did: place your trust in editors (Mehler and Bever) with good taste and let them exercise it. As I have mentioned before, some of the greatest journals were run in this way (Keynes ran a journal as did Planck and both were outstanding). Now, I am not acquainted with Ms Felser, but I do know Colin and Matt quite well and I know them to have excellent taste in questions. So, I suspect that one reason for the success of their journal is this X factor. We need more of this. Not to replace the mainline journals, but to allow idiosyncrasy to sometimes get a hearing. We need editors that evaluate papers in terms of novelty and importance of the ideas. This is not quite impact (at least in the terms that Colin notes are relevant to this currently important metric), but it is a big deal, IMO. It can take time for a new idea to gain a foothold, but when it does, well, you can finish this sentence as well as I can.

So, look at the post. It's really thought provoking. I'd be interested to see comments on this in Colin's blog.

The deep conservatism of biological systems; some lessons for linguists

Linguistic theory and biology have long had a mutually beneficial trade in ideas. Kai Jerne (here) explicitly cites the influence of Generative Grammar on the way he approached the immune system (here). The Principles and Parameters (P&P) theory and the Monod-Jacob theory of biological switches drank from the same stream of ideas and Chomsky has long noted the similarity of his Rationalist take on biology to that of ethologists like Tinbergen, Lorenz and von Frisch.

The most recent instance of the convergence/influence of views comes with the rise of Evo-Devo (ED) adjustments to the standard Darwinian story of natural selection. Here, I want to focus on one property of such accounts that I believe carry an important lesson for grammatical theory; the deep conservation of basic biological mechanisms. Let me explain with an illustration.

One of the big unexpected ED discoveries concerned the eye. The hero of this little tale is Walter Gehring (WG). Here’s a potted history of what WG discovered (based on this very accessible paper (here) by Gehring and Ikeo (G&I)). Prior to Gehring’s work the eye was the poster child for analogous evolution; the idea you can get the same “organ” via different evolutionary routes. Based on comparative and structural studies, the standard account was that “the photo receptor organs … originated independently in at least 40, but possibly up to 65 or more different phyletic lines” (371). The basis for this view? The “obvious” physical differences between the different kinds of eyes. G&I describes the state of play before Gehring’s work as follows:

Their section on ‘the multiple origin of eyes’ begins with the comment that ‘it requires little persuasion to become convinced that the lens eye of a vertebrate and the compound eye of an insect are independent evolutionary developments’. This point has been taught to biology students for over a hundred years. (371)

This turned out to be wrong. In other words, despite the evident surface differences among eyes in different species, nonetheless they all are basic variations of the same basic design. Or put another way, these different eyes all have a common ancestor, all stem from a common prototype, the differences being bells and whistles added to shared basic design. In an important sense then, there is only one eye.[1]

Btw, before saying a bit more about this, it is worth noting an interesting side comment that G&I make on Darwin’s views of natural selection, specifically concerning the origin of the eye prototype. Here’s the relevant wording:

Darwin was highly self-critical in his discussion of the eye prototype and admits that the origin of the prototype cannot be explained by natural selection, because selection can only drive the evolution of an eye once it is partly functional and capable of light detection. Therefore, selection cannot explain the origin of the eye prototype, which for Darwin represents the same problem as the origin of
life. Therefore, both the origin of life and the origin of the eye prototype must have been very rare events, and a polyphyletic origin in over 40 different phyla is not compatible with Darwin’s theory. (371)

In other words, according to Darwin, the emergence of prototypes cannot be accounted for in terms of Natural Selection, rather Natural Selection presupposes such rare events. No prototypes, no tinkering. Moreover, the emergence of prototypes must be very rare events. Sounds more than a bit like Chomsky’s proposals for the origins of recursion doesn’t it? So the next time you hear someone say Chomsky’s views are incompatible with evolution (i.e. that the sudden emergence of Merge makes no evolutionary sense), tell them to go and look up Darwin’s discussion of the emergence of the eye.

Ok, back to the main point.  What’s the new basic view of the evolution of the eye? There is a master control gene responsible for all eyes that is shared across all sighted creatures from Drosophila to mice to humans. In other words, there is really only one eye that are based on the many of the same basic building blocks. P&P indeed!

Any implications of this for linguistic theory? Not logical ones, but very suggestive ones.  Two in particular come to mind.

First, that “evident” surface differences in biology are not reliable signs of different mechanisms. This holds just as true (or should do so) in linguistics as in other parts of biology. I have discussed this recently (no doubt ad nauseum) here and here. But the point is an important one given the temptation to believe what is right before one’s eyes. This may sound like good common sense, but if science has taught us anything it is that things that look very different might nonetheless be exactly the same. In fact, what we consider the highlights of scientific insight basically consist in showing that two things that have entirely antagonistic properties are actually at bottom identical.[2]

Second, that biology finds it very hard to truly innovate. In other words, that there is basically only one way of doing anything. This suggests skepticism whenever a linguistics paper suggests that more than one way of doing anything. For example, there is a current view within syntax that there are two very different kinds of relative clauses formed using two very kinds of operations (matching vs movement). This view is buttressed by showing that these two constructions have different properties and the inference is made that such different properties require different kinds of operations to explain them. That is, of course, one way to explain surface differences. However, perhaps we should resist this inference until dragged kicking and screaming to the conclusion.[3] Why? Because multiplying mechanisms comes at the cost of explanation AND because it looks like biological systems don’t like doing anything two ways. So if you are a thoroughly modern GGer who understands the enterprise in cog/bio terms then you should be reluctant to ever multiply mechanisms. In short, as a general methodological rule of thumb, you should assume that there is only ever one road to Rome.

The proliferation of theoretical mechanisms is an easy way of covering apparently different properties. And, in fact, it is logically possible that variation signals divergent mechanisms. However, the successful sciences, in particular physics, has got to where it is by strongly resisting this mode of inference. It was supposed that this might be good for physics but not for biology where typically many mechanisms that can result in what look like the same effects. We now know that this is not quite so at the smallest cellular levels and ED is showing us that this is likely not even so at the macro level. Not only are living things more or less the same in micro, but it seems that our macro organs (e.g. eyes) are also built in more or less the same ways as well. I suggest that we adopt the same stance within linguistics. Our regulative ideal should be that FL/UG never does anything in two different ways.



[1] Just as there is fundamentally only one language, or more precisely one FL. Or put another way: impressive surface differences among languages do not imply different underlying etiologies.
[2] My good and great friend Elan Dresher once observed that there are really only two kinds of theoretical linguistic papers. The first shows that two things that appear radically different are more or less the same. The second shows that things that are more or less the same are in fact identical. Yup.
[3] For a provocative and interesting discussion resisting these temptations within current syntactic theory, see Dominique Sportiche’s paper Neglect (here). It is very much in the spirit of Gehring’s efforts regarding eyes.

Monday, December 21, 2015

Holidays

I am going to be offline for most of the next two weeks. I don't swear not to put up an occasional post, but I am pretty sure that I won't. So, enjoy the time off. See you again in January.