Comments

Monday, November 19, 2012

Grammatical Zombies



The talk on the street is that linguists should stuff their portfolios with probabilistic grammars as they have intrinsic advantages over the old, stogy under-performing non-probabilsitic grammars of yore.  Partisans assure us that going probabilistic makes the learning problem easier, reduces the need for rich innate structures, and even lowers cholestoral levels. Though I am genetically inclined to skepticism (especially about things that sound too good to be true), debunking these claims has always required talents above my pay grade.  Lucky for me I know people who understand these issues deeply and are able to translate their significance for the unwashed, i.e. me. This is all preamble for the following post by Bob Berwick, someone that I have been able to persuade to post on an occasional basis, (hopefully not too occasionally) especially on technical topics from computational linguistics. The bottom line is that ideas that are too good to be true aren't, though their intellectual half lives can be very very long, as Bob explains below. Btw, I stole the title for this post from Bob's apposite reference to John Quiggin, an Australian economist who has done yeoman's work debunking hard to kill bad ideas in macroeconomics.

                                                             Going off the Gold Standard
                                                                     Robert C. Berwick

You don't have to wander very far along the computational linguistics byways these days to stumble across what economist John Quiggin dubs “zombie ideas” - bad memes like “supply-side economics after the Clinton boom and the Bush bust – repeatedly refuted with evidence and analysis" yet somehow always rising from the dead.  Among these, two stand out, virtually articles of probabilistic faith: One, that Gold’s (1967) celebrated results about the non-learnability of most language families given positive-only evidence become moot, once one “goes probabilistic.”  And two, that “going probabilistic” can be cashed out simply by adding probabilities to context-free rules, rendering probabilistic context-free grammars easier to learn than their context-free counterparts.  Easy enough, in fact, to drive a stake through the heart of Demon Rationalism (aka, prior constraints on grammars).  But is any of this true?  Or is it instead, to quote the now-immortal words of Meygn Kelly on election night, just “math you do as a Republican to make yourself feel better”?
Well, Yes and No.  Yes, in that nearly every thoughtful scholar of language learnability, starting just a few years post Gold-rush, bound up in the foundational work of Jay Horning, Jerry Feldman, and Ken Wexler, in the late 1960s, and up to the present day, has stressed that the tough Gold standards for learning exact identification of languages on all possible positive texts – are cognitively implausible and overly restrictive, so should be relaxed in favor of what's called “probabilistic example presentation” and “probabilistic convergence.”  That was, and still is, the consensus view.  In other words, from the very beginning of serious research into language learnability for natural languages, the probabilistic view has stood front and center.  All agree we’re after a learning standard that ensures much more: feasibility. (Chomsky's words, 1965, pp. 53-54) or what Ken Wexler (1980) termed “easy learnability”, that is, “learnability from fairly restricted primary data, in a sufficiently quick time, with limited use of memory” (p. 18).[1] [2]
In probabilistic presentation, a learner gets sentences drawn from some underlying distribution, and only has to succeed on positive example sequence that it gets with sufficient probability, rather than on all positive example sequences, as with Gold.  And in fact, there’s a technical sense in which this move does the trick.  When learnability on all texts – the Gold standard – is relaxed in favor of what Ken Wexler was first to call “measure-1 learnability” at a Stanford workshop in September, 1970, then, unsurprisingly, moving the goalposts changes the game.  What perhaps is surprising is this change seemingly moves the goalposts almost to the line of scrimmage: now, instead of no interesting language classes being learnable, all interesting language classes are learnable (officially, all recursively enumerable language families).

Why?  Picture sentences being spit out one at a time by some underlying stochastic source (aka, the “I blame the parents” model), so every sentence occurs with some probability. (Assuming independent draws – what’s technically called “i.i.d” – draws from an independent and identically distributed stochastic source).  In that case, a learner cannot count on using some vanishingly rare example sentence as a red flag for their correct target language. So, we don’t want to insist, as Gold did, that the learner must succeed on all texts.  If figuring out that your language is English means having to wait for the one David Foster Wallace sentence that begins “Uninitiated adults who might...” and trundles to a close 90 words later with, “subdued, almost narcotized-looking,” then you know you’re goose is cooked.  However, if a learner knows that all longer and longer sentences always become less and less likely, then it can know with high confidence that e.g., the first 100 sentences are more than enough to sort out what particular target language your source has been pitching.

And this is exactly what Horning (1969) did.  He attached probabilities to context-free grammar rules, to get probabilistic context-free grammars (PCFGs).  So for instance, if there are 2 rules in a grammar expanding NPs, NP → Det N and NP → Pronoun, then we have a probability of 0.5 for each.  Since every rule in a PCFG is assumed to be statistically independent, the probability of any particular sentence is simply the product of the probabilities of the rules deriving it.  As sentences get longer and longer, their probabilities must shrink – they become exponentially less likely. If the learner knows this in advance, then it knows that very long sentences are so unlikely that they can be effectively eliminated, with the problem reduced to essentially that of learning to identify a family of finite languages. But we already knew that finite language families are learnable from positive-only examples.[3]
In fact, as Horning (1969, 80-81) suggests, one need not rely on any conventional probability distribution over context-free grammars and languages; one could substitute, for example, a measure that ascribes exponentially lower probabilities to larger grammars, and search for target grammars in order of increasing length, or even a probability measure that just assigns the value 2–i to the grammar enumerated at step i, Gi (as discussed by Ken Wexler, 1980, fn. 22, 529-532; this is one way to develop the notion of an evaluation metric, something we’ll consider in later posts).  More generally, any so-called uniformly computable probability measure works (as proved in Dan Osherson, Michael Stob, and Scott Weinstein’s book, 1986, chapter 10 pp. 187-188, and by Angluin 1988).
So, are we done? Does adding probabilities to context-free rules really make them easier to learn?  No, there's a catch. TANSTAAFL (viz. There ain’t no such thing as a free lunch!). The measure-1 result holds scant cognitive interest (though it can be said to point the way to what constraints we really do need for human language learning models).  Why? The positive results for measure-1 learnability only work if the learner has extremely strong prior knowledge about the probability distribution that's feeding it sentences.  If you really want the learner’s tabula to be a little more rasa, then as the late Partha Niyogi put the matter, “when one considers statistical learning...all statistical estimation algorithms worth their salt are required to converge in a distribution-free sense” (Niyogi, 2005, p. 67).  So that’s to say, suppose the learner doesn’t have any pre-conceived (and accurate) notion of how the example sentences will be distributed. Problem is, Niyogi then shows, as per Angluin (1988), that distribution-free statistical learning collapses back to Gold learning: if a family of languages is measure-1 learnable in a distribution-free sense, then it is Gold learnable.  Measure-1 learnability doesn't enlarge the class of learnable languages one jot. Oops.
As if that weren’t enough, the assumptions blow away any remaining cognitive fidelity we thought we had won.  For one thing, the statistical independence of PCFG rules is precisely what has been explicitly and strongly rejected by the same folks who embrace Horning, in the recent flood of parsers statistically trained in a supervised way from pre-parsed corpus data.  Rather, such rules are plainly dependent: once you decide a sentence is NP VP, then it’s far more likely that the NP, being a subject, will be expanded as a pronoun, with the reverse true if the NP’s an object – which is just to say that PCFG rules aren’t really statistically independent.  The assumption that sentences are pitched to the learner in i.i.d fashion is plainly wrong.  Nobody believes parents speak this way.
Then too, Horning’s method for selecting grammars – explicit enumeration along with search using Bayes’ rule to hunt for the highest posterior probability grammar given the examples, sometimes collapsing phrase names together and sometimes splitting them apart – proved to be computationally intractable.  (It’s also not a wit different conceptually from what’s been advanced by contemporary Bayesian enthusiasts as far as anyone can tell.) Horning himself remarks on how slow his “550 line Lisp program,”  “EVALUATER”, ran, in part because “the enumerative problem is immense” (pp. 199-120). It could only (partly) find grammars with 2 or 3 nonterminals. To be sure, computation speed has increased by leaps and bounds over the past 35 years from Lisp/36 running on an IBM 360/67. And computer scientists have come up with faster methods to solve Bayesian problems, often approximate – ‘variational Bayes’ and the like.  Nevertheless, the immense space of possible PCFGs – larger than the space of CFGs – confounds search methods to this day.
So in truth, Horning’s enumerative method doesn’t nearly go far enough. PCFGs don’t turn out to be any “easier to learn” than their non-probabilistic counterparts –  perhaps what’s meant here is that it’s  more straightforward to use known techniques to estimate PCFG probabilities.  However, that’s not the same thing as learning the underlying rules themselves.  In fact, estimating language distributions is harder in principle that estimating languages simpliciter, as Niyogi (2005, p.68) observes.  What’s needed, rather, is far, far, stronger medicine, something along the lines of Ken Wexler's “bounded degree of error” approach, in which the possible mistakes a learner can detect with respect to sentences of at most 2 embeddings spans the entire space of all possible errors across a complete family of transformational grammars.  Crucially, Ken’s approach is not enumerative.  Instead, it carries out a search in a tightly-bounded, finite space, with each learning step incrementally improving a single transformational component, rather than selecting and throwing away wholesale entire grammars. But to get this to work, Wexler found out that one must impose far stronger a priori constraints on the family of grammars or languages. 
In our internet era where it sometimes feels as though a veil of ignorance has been drawn across any work that was accomplished before 1990, it is well worth remembering that by that September 1970 workshop held at Stanford, Wexler and colleagues had already set up this model for learning transformational generative grammar where both presentation and convergence were probabilistic, in just the sense described by modern day probability enthusiasts – and found this was not nearly enough.  One had to include constraints on grammars, constraints that wound up to be empirically attested, like locality constraints on how far one could displace phrases.  That work, establishing a positive learnability result, still stands – though rarely cited by statistical enthusiasts despite its probabilistic underpinnings.  In fact, it’s apparently even been recently revived just this past year, in work by Mark Steedman and colleagues (2012), though without any apparent tip of the hat in Ken’s direction.  Mark uses combinatory categorial grammar instead of transformational grammar, but in many other ways follows Ken's model, with the learner using (meaning, surface string) pairs.  Plus ça change.

Next: Why Revolutionary new ideas do appear infrequently.



[1]A partial list of people embracing this view would include Jay Horning (1969); Henry Hamburger and Ken Wexler (1970, 1973, 1975); Bob Matthews (1979); Peter Culicover and Ken (1980); yours truly (1982/1985); Dana Angluin (1988); and Partha Niyogi (2005), among others. Here's Horning, 1969, 19, fn. 1, “In the sequel, when we prove identifiability in the limit with text presentation, it is with a different performance requirement, and a different condition on the presentation.”
[2]Here is the quote from Wexler's presentation at a September, 1970 workshop held at Stanford p.159 “A language class is identifiable in the limit with probability 1 with respect to a probabilistic presentation scheme if there exists a learning procedure such that for any member of the class there is a subset of measure 1, of the set of presentation sequences for which the procedure identifies in the limit the language."

[3]In fact, Horning, in what seems to have gone largely unnoticed by statistical enthusiasts stresses several times that his results hold only for the case of unambiguous context-free grammars: 1969: 32-33: “In the sequel we assume...that unambiguous grammars are desired, and will reject grammars which make any sample string ambiguous.”  Of course, on the assumption that natural languages are ambiguous, this immediately makes his results unacceptable from the standpoint of cognitive fidelity.  It’s hard to know just why this has been overlooked for 45 years.  Putting aside the possibility that the people citing Horning in fact haven't really read his thesis, luckily, this omission is of little importance, because Horning simply didn’t know how to place a properly-defined probability measure on the languages generated by ambiguous context-free grammars, a matter that has since been remedied by Osherson, Stob, and Weinstein (1986), who established a far more general result encompassing Horning’s, as noted below.

Sunday, November 18, 2012

‘I’ before ‘E’: asterisks and ungenerability

If (1) is a sentence, then what is (2)? More importantly, what does the asterisk in (2) signify?
(1)   A doctor might have been there
(2) *A doctor might been have there
And do the asterisks in (2-5) have the same significance?
(3) *Was the guest who fed waffles fed the parking meter?
(4) *Colorless green ideas sleep furiously
(5) *The rat a cat the dog chased chased hid
In conversations, I’ve been assured that practitioners know what they’re doing, and that only troublemakers ask such questions. But often enough, my assurers disagree about what the standard “data marks” mean, especially if asked how ‘*’ differs from ‘#’ and ‘??’. I’m certain that many good students are, quite understandably, confused about how these marks get used. And as a philosopher, I get paid to make trouble. (It’s nice work if you can get it.)
Chapter one of Aspects is, in my view, very clear and very right about the distinction between acceptability and grammaticality. So for present purposes, I’ll take it as given that asterisks indicate a kind of unacceptability that ordinary speakers can detect and report, while grammaticality is a theoretical notion that ordinary speakers may not have (pace Michael Devitt’s Ignorance of Language). But one wants to know how data regarding unacceptability can be evidence for or against theories of implementable procedures that connect articulations with meanings. For example, how does the oddity of (3) bear on theories of Human I-Languages?
            As previously discussed, (3) has the meaning of (3a), which is weird but grammatical.
(3a) *The guest who fed waffles was fed the parking meter?
But (3) would be acceptable if it could be understood as having the meaning of (3b).
(3b) #The guest who was fed waffles fed the parking meter?
Here, ‘#’ indicates that (3) cannot be understood as (3b) is understood. Theorists can describe this point about (3) in terms of a grammatical constraint: the auxiliary verb ‘was’ cannot be displaced from the relative clause ‘who was fed waffles’; though it can be displaced from the verb phrase ‘was fed the parking meter’. But whatever the pronunciation, it’s weird to talk about feeding waffles or feeding a guest a parking meter. And since (6) is fine,
                        (6) Was the guest who fed the parking meter fed waffles?           
it seems clear that (3) has a grammatical reading, viz. (3a). But then the unacceptability of (3) reflects a cluster of facts. The perfectly fine thought indicated with the acceptable sentence (3b) can’t be expressed with (3); and the thought that can be expressed, indicated with (3a), is weird.
            In one respect, (3) is like (4): grammatical and hence meaningful as opposed to gibberish, yet apt to elicit a “bizarreness reaction” in competent speakers who know some things about waffles, guests, ideas, sleep, etc. In another respect, (3) is like (2), which can’t be understood as having the meaning of (1). It’s quite interesting that (2) is nearly word salad, as opposed to a comprehensible second way of expressing the thought expressed with (1). Compare (7),
(7)  *The child seems sleeping
which is a degraded but still comprehensible way of saying that the child seems to be sleeping.
Moral: if (3) is like (4) in one respect, and like (2) in another respect, then it’s important to distinguish sources of unacceptability. But it’s hard to see how one can make the requisite distinctions without positing a procedure that generates articulation-meaning pairs, where the generable meanings are in turn related to mental representations that may or may not be reasonable depictions of language-independent reality. To repeat, the “data point” noted with the asterisk in (3) reflects a cluster of underlying facts: the perfectly reasonable question indicated with (3b) cannot be expressed with (3); and the expressable query, indicated with (3a), is bizarre.
Famously, the sources of unacceptability differ across (2), (4), and (5). It turns out that (5) is a sentence that can be paraphrased with (5a), which is long and awkward, but not crazy.
                        (5a)  The rat that was chased by a cat which the dog chased (was a rat that) hid
Prima facie, the anomaly of (5) has to do with memory limitations and center embedding, as opposed to either constraints on generability of expressions or the kind of “conceptual boggle” that attends thoughts of feeding waffles. This reminds us that Human I-languages can and evidently do (generatively) connect pronunciations with meanings that may go unrecognized by those who have the I-languages, even in cases that involve a relatively small number of words. As Chomsky also noted, it takes works to hear all the possible readings of (8).
                        (8)  I almost had my wallet stolen.
Though the point regarding (3) is a little different. Given a way of classifying examples like (5) as “grammatical but hard to parse,” one might set such data points aside and try to construct an algorithm that classifies other strings as acceptable or not. But examples like (4) remind us that some unacceptable strings are easily parsed “linearizations” of generable expressions that exhibit perfectly fine grammatical structure. Of course, as (2) illustrates, many unacceptable strings are not linearizations of any generable expressions. So if the aim is to provide a theory of human linguistic understanding, one needs to specify an algorithm that pairs the pronunciation of (4) with its meaning and doesn’t pair the pronunciation of (2) with any meaning that it doesn’t have.
Put another way, (2) presents a special case of unambiguity: zero meanings, as opposed to even one; whereas (3) has one meaning but not two. That is, (2) is not the linearization of any generable structured expression that connects the pronunciation in question with a meaning. It’s not that (2) is a sentence, albeit an unacceptable one. That’s the situation with regard to (3). But the more interesting thing about (3) is that its pronunciation isn’t linked to the meaning of (3b). And the interesting thing about (2) is that its pronunciation isn’t linked to the meaning of (1).
It’s also true that the pronunciation of (2) isn’t linked to the meaning of (5a). But that’s hardly surprising. For any word-string, there are endlessly many meanings it doesn’t have. Yet some of those non-meanings can be built up from the word meanings in ways that initially seem no more complicated than the ways that actual expression meanings are built up from word meanings. So absence of homophony, with word salad as a special case, can provide valuable clues about the procedures that generate pronunciation-meaning pairs. So if Human Languages are such procedures, then the data that linguists typically use can indeed be used as evidence for/against theories of Human Languages. Those who don’t adopt an I-language perspective need to say what their target of inquiry is such that it still makes sense to be using the same data.
            We probably also need a graded, multi-dimensional notion of grammaticality. For me, (9) is OK though marginal on its only reading. (Did someone find a helpful dog for any vet?)
(9)  Was a vet found a dog that helped?
Perhaps (9) deserves a mark less harsh than an asterisk. But in any case, the clear unacceptability of (10) reflects the unavailability of meaning (10a) and the unacceptability of (10b).
(10)   *Was a vet helped a dog that found?
(10a) #A vet helped a dog that was found?
(10b) *A vet was helped a dog that found?
The unacceptability of ‘was helped a dog that found’, unlike that of ‘fed waffles’, may well be due to grammar. But to make such distinctions, we need to talk about constrained generative procedures, not conditions on logically possible outputs that might or might not be generable.

Friday, November 16, 2012

QR?


 Caveat Lector! What follows is wonkier than what I have posted heretofore.

Chris Barker has a remark in the recent issue of LI where he presents, what seems to me, pretty convincing evidence that quantificational binding (of a pronoun) does not require c-command, wider scope suffices.  More particularly, he adopts Safir’s scope requirement:

(1) A quantifier must scope over any pronoun that it binds.

and operationalizes scope with (2):

(2) A quantifier can take scope over a pronoun only if it can take scope over an
                 existential inserted in the place of the pronoun.

The test in (2) conceptually divorces ‘scope’ from ‘c-command’ and allows one to investigate whether the two move in tandem, i.e. whether a scopes over b iff a c-commands b. The data Barker presents shows pretty clearly that though c-command is a sufficient condition for scope, it is not obviously necessary.[1] 

Barker’s interesting discussion serves a (perfectly reasonable) PR function. He has another way of doing binding/scope that fits well with the conclusion that it is not tied to c-command.[2] He does not argue in detail against theories that have tried to save the Scope/C-command generalization, but his congeries of counter-examples certainly suggests that defenders of the orthodox  have their work cut out for them (including an earlier counterpart of me should I/he be tempted to defend (2)). Though I could quibble with some of the data (there are a lot of cases involving ‘each,’ which is known to be recalcitrant in many ways, e.g. WCO effects with ‘each’ are quite attenuated (his1 mother considers/believes each boy1 to be perfect/His1 mother yelled at each boy1 to stop shouting), I found weight of data and Barker’s and conclusion based on it pretty convincing.  At the very least, he has made a good case to re-open our minds about the relation between c-command and scope.

This said, I want to point out another consequence of Barker’s argument should his conclusion prove correct. It suggests that UG does not concern itself with quantifier scope and that QR, the grammatical vehicle for adjusting phrase structures in order to bring scope and c-command together, is not a rule of grammar.  This view has a pedigree. Tony Kroch’s thesis argues that the grammar does not regulate variable Q-scope. In fact, it was Chomsky’s work on WCO and the noted parallel between (3a,b) that prompted taking quantifier scope to be a species of A’-movement:[3]

            (3)       a. *Who1 does his1 mother love t1
                        b. *His1 mother loves everyone1

(3a) resists a paraphrase as whose mother loves him and (3b) is infelicitous as everyone’s mother loves him.  Chomsky’s very reasonable point was that the facts in (3) could be collapsed if WCO were an LF fact and (3b) was covertly transformed via QR into a structure analogous to that in (3a).

Soon enough, other facts arose that supported this view (the curious can look at chapter 1 of my book Logical Form for a review of some of this).  However, Barker’s paper suggests that this data is not representative and that the grammar does not regulate scope via QR at all. The paper cites examples where quantifiers scope out of subject islands, relative clauses, subject sentences and adjuncts.  If these data are more or less accurate, then it suggests that QR is not a rule of grammar.  More exactly, QR seems unregulated by the kinds of locality conditions we expect movement operations to be subject to.

There remains a way of finessing this conclusion. Maybe covert movement is not subject to locality restrictions of the kind we see for overt movement. This is a venerable assumption.  Those predisposed to single-cycle theories of syntax (that’s you minimalists out there) might find this assumption challenging.  I suspect that there are ways around this (e.g. treating islands as effectively PF linearization phenomena) but that too will take some work.  At any rate, Barker’s conclusion, if correct, has interesting consequences.

I end this rambling with a question: Why were we so easily convinced that binding required c-command? First, because we understood LF as the input to semantic interpretation and took difference in truth conditions to be sufficient indication of difference in LFs. Thus, if (4) is ambiguous, it must have at least two different LFs. QR delivered two different LFs.

            (4)       At least one boy kissed every girl

Second, as May emphasized, QR was the perfect fix for the regress problem that ACD constructions appeared to present.

The first assumption was now far less attractive, at least to me. There is no reason to assume that the grammar alone determines semantic interpretation. Being one factor among many leaves LF/CI interface issues plenty interesting. The second reason stands. If Barker is right, then it looks like we need to rethink the analysis of ACDs. Some have started doing so, attributing ACD licensing to something like extraposition (Baltin, Fox-Nissenbaum), or ‘afterthoughts’ (Chomsky).  At any rate, the free ride that the QR analysis of ACDs enjoyed by piggy backing on the “independently motivated” rule of QR is probably over.

I confess that I have always found QR a bit unsavory. It didn’t really function like other A’-movement operations (too many landing sites), it really was hard to find overt versions of it (something unexpected if analogized to Wh-movement), it enjoyed different properties than other forms of A-movement (it doesn’t require obligatory reconstruction for if it did it could not license ACDs, but quantified RCs cannot generally drag their sentential parts along for the covert ride for if possible principle C would be systematically violated) etc.  Barker’s paper suggests that we need to rethink not only the c-command condition on binding and scope but the standing of QR as a grammatical operation.



[1] Dave Kush (p.c.) noted that the sufficiency of c-command for scope is quite interesting.  Why should it be that a’s c-commanding b is sufficient for a to scope over b. One can imagine systems where the two notions are entirely divorced and that some c-commanding elements cannot scope over elements they c-command.  Interestingly, one of Reinhart’s tests for bound variable readings of pronouns does seem sensitive to c-command:
(i)             John loves his mother but Bill doesn’t (sloppy ok)
(ii)           John’s father loves him but Bill’s mother doesn’t (no sloppy)
(iii)          Everyone from NYC loves its subway. Does everyone from Boston? (no-ish sloppy)
I find the sloppy readings in (ii) and (iii) marginal at best, though I can be convinced that I am wrong about this.
[2] Earlier work in the mid 80s by George Wilson and Jeff King treats certain cases of bound pronouns as flagged variables in a natural deduction system. This approach also dissociates binding and scope from c-command.  
[3] May’s subsequent work also proved influential.

Thursday, November 15, 2012

Poverty of Stimulus Redux



 This paper by Berwick, Pietroski, Yankama and Chomsky (BPYC) offers an excellent succinct review of the logic of the Poverty of Stimulus argument (POS). In addition it provides an up to date critical survey of putative “refutations.”  Let me highlight (and maybe slightly elaborate) on some points that I found particularly useful.

First, they offer an important perspective, one in which the POS is one step in a more comprehensive enterprise. What’s the utility in identifying “innate domain specific factors” of linguistic cognition? It is “part of a larger attempt to…isolate the role of other factors [my emphasis](1209).”  In other words, the larger goal is to resolve linguistic cognition into four possible factors; (i) innate domain specific, (ii) innate domain general, (iii) external stimuli, and (iv) effects of “natural law.” Understanding (i) is a critical first step in better specifying the roles of the other three factors and how they combine to produce linguistic competence. As they put it:

The point of a POS argument is not to replace “learning” with appeals to “innate principles” of Universal Grammar (UG). The goal is to identify factor (i) contributions to linguistic knowledge, in a way that helps characterize those contributions. One hopes for subsequent revision and reduction of the initial characterization, so that 50 years later, the posited UG seems better grounded (1210).

I have the sense that many of those that get antsy with UG and POS do so because they see it as smug explanatory complacency: all linguists do is shove all sorts of stuff into UG and declare the problem of linguistic competence solved!  Wrong. What linguists have done is create a body of knowledge outlining some non-trivial properties of UG. In this light, for example, GB’s principles and parameters can be understood as identifying a dozen or so “laws of grammar,” which can now themselves become the object of further investigation, e.g. Are these laws basic or derived from more basic cognitive and physical factors? (minimalists believe the latter), What more basic factors are these? (Chomsky thinks Merge is the heart of the system and speculation abounds about whether it is domain general or linguistically specific) and so forth.  The principles/laws, in other words, become potential explananda in their own right. There is nothing wrong with reducing proposed innate principles of UG to other factors (in fact, this is now the parlor game of choice in contemporary minimalist syntax). However, to be worthwhile it helps a lot to start with a relatively decent description of the facts that need explaining and the POS has been a vital tool in establishing these facts. Most critiques of the POS fail to appreciate how useful it is for identifying what needs to be explained.

The second section of the BPYC paper recaps the parade case for POS, Aux to Comp movement in Y/N questions. It provides an excellent and improved description of the main facts that need explanation.  It identifies the target of inquiry to be the phenomenon of “constrained homophony,” i.e. humans given a word-string will understand it to have a subset of the possible interpretations logically attributable to it.  Importantly, native speakers find both that strings have the meanings they do and don’t have the meanings they don’t. The core phenomenon is “(un)acceptability under an interpretation." Simple unacceptability (due to pure ungrammaticality with no interpreations) is the special case of having zero readings. Thus what needs explanation is how given ordinary experience native speakers develop a capacity to identify both the possible and impossible interpretations and thus:

…language acquisition is not merely a matter of acquiring a capacity to associate word strings with interpretations.  Much less is it a mere process of acquiring a (weak generative) capacity to produce just the valid word strings of the language (1212).

As BPYC go on to show in §4, the main problem with most of the non-generative counter proposals is that they simply misidentify what needs to be explained. None of the three proposals on offer even discuss the problem of  constrained homophony, let alone account for it. BPYC emphasize this in their discussion of string-based approaches in §4.1. However, the same problem extends to the Reali and Christensen bi-gram/tri-gram/recurrent network models discussed in §4.3 and the Perfors, Tanenbaum and Regier paper in §4.2, though BPYC don’t emphasize this in their discussion of the latter two. 

The take home message from §2 and §4.1 and §4.2 is that even for this relatively simple case of Y/N question formation, critics of the POS have “solved” the wrong problem (often poorly, as BPYC demonstrate).

Sociologically, the most important section of the BPYC paper is §4.2.  Here they review a paper by Perfors, Tenenbaum and Regier (PTR) that has generated a lot of buzz. They effectively show that where PTR is not misleading, it fails to illuminate matters much. The interested reader should take a careful look. I want to highlight two points.

First, contrary to what one might expect given the opening paragraphs PTR do not engage with the original problem that Aux-inversion generated; whether UG requires that grammatical (transformational) rules be structure dependent. PTR address another question: whether the primary linguistic data contains information that would force an ideal learner to choose a Phrase Structure Grammar (PSG) over either a finite list or a right regular grammar (generates only right branching structures).  PTR conclude that given these three options there is sufficient information in the Primary Linguistic Data (PLD) for an ideal learner to choose a PSG type grammar over the other two. Whatever the interest of this substitute question, it is completely unrelated to the original one Chomsky posed. Whether grammars have PSG structure says nothing about whether transformations are structure or linearly dependent processes.  This point was made as soon as the PTR paper saw the light of day (I heard Lasnik and Uriagereka make it at a conference at MIT many years ago where the paper was presented) and it is surprising that the published PTR version did not clearly point out that the authors were going to discuss a completely different problem. I assume it’s because the POS problem is sexier than the one that PTR actually address. Nothing like a little bait and switch to goose interest in one’s work.

Putting this important matter aside, it’s worth asking what PTR’s paper shows on its own terms. Curiously, it does not show that kids actually use PLD to choose among the three competing possibilities.  They cannot show this for PTR is exploring the behavior of ideal learners, not actual ones.  How close kids are to ideal learners is an open question.  And similar to one Chomsky long ago considered. Chomsky’s reasons for moving from evaluation measure models (as in Aspects) to parameter setting ones (as in LGB) revolved around the computational feasibility of providing an ordering of alternative grammars necessary for having usable evaluation metrics.  Chomsky thought (and still thinks) that this is a very tall computational order, one unlikely to be realizable.  The same kind of feasibility issues affects PTR’s idealization. How computationally feasible is to assume that are able to order, compare and decide among the many grammars compatible with the data that are in the hypothesis space? The larger the space of options available for comparison the more demanding the problem typically is. When looking for needles, choose small haystacks. In the end, what PTR shows is that there is usable information in the PLD that were it used could choose PSGs over the two other alternatives. It does not show that kids do or can use this information effectively.

The results are IMHO more modest still. The POS argument has been used to make the rationalist point that linguistically capable minds come stocked full of linguistically relevant information. PTR’s model agrees. The ideal mind comes with the three possible options pre-coded in the hypothesis space. What they show is that given such a specification of the hypothesis space, the PLD could be used to choose among them. So, as presented, the three grammatical options (note the domain specificity: it's grammars in the hypothesis space) are given (i.e. innately specified). What's “learned” (i.e. data driven) is not the range of options but the particular option selected. What pposition does PTR argue against? It appears to be the following position: only by delimiting the hypothesis space of possible grammars so as to exclude all but PSGs can we explain why the grammars attained are PSGs (which of course they are not, as we have known for a long while, but ignore that here). PTR's proposal is that it's ok to widen the options open for consideration to include non-PSG grammars because the PLD suffices to single out PSGs in this expanded space. 

I can’t personally identify any generativist who has held the position PTR targets. There are two dimensions in a Bayesian scenario: (i) the possible options the hypothesis space delimits, (ii) a possible weighting of the given options giving some higher priors than others (making them less marked in linguistic terms).  These dimensions are also part of Chomsky’s abstract description of the options in Aspects chapter 1 and crops up in current work on questions of whether some parameter value is marked.  So far as I can tell, the rationalist ambitions the POS serves are equally well met by theories that limit the size of the hypothesis space and those that widen it but make some options more desirable than others via various kinds of (markedness) measures (viz. priors). Thus, even disregarding the fact that the issue PTR discuss is not what generativists mean by structure dependence, it is not clear how revelatory their conclusions are as their learning scenario assumes exactly the kind of richly structured domain specific innate hypothesis space the POS generally aims to establish. So, if you are thinking that PTR gets you out from under rich domain specific innate structures, think again.  Indeed if anything, PTR pack more into the innate hypothesis space than generativists typically do.

Rabbinic law requires that every Jew rehash the story of the exodus every year on Passover. Why? The rabbis answer: it’s important for everyone in every generation to personally understand the whys and wherefores of the exodus, to feel as if s/he too were personally liberated, lest it’s forgotten how much was gained.  The seductions of empiricism are almost alluring as “the fleshpots of Egypt.” Reading BPYC responsively, with friends, is excellent antidote, lest we forget!