Comments

Tuesday, April 9, 2013

Data Free Commitment; MOOCs Mania

I chased down on of the links from the Dan Little piece I posted yesterday to take a quick peek at the data backing up the recent rush to MOOCing the classroom. This paper by Bowen and Lack (BL) reviews the data, such as there is, reviewing studies on the topic. As BL note, there was not a lot to go on. Right now, MOOCs are a faith based initiative, or as BL put it: "As of the moment, a priori argument will have to continue to bear much of the weigh" "in determining how much more of an investment, and in what form. should be made in this field (4)." They go on to say that there is little "conclusive evidence" about the cost effectiveness of online courses, in fact little good evidence at all. They also observe that much of this is driven by a desire for cost reduction: "achieving the same outcomes at half the cost per student should be seen as a great victory (11)." My worry is that as the outcomes are very hard to measure, as BL show, the focus will entirely shift to reducing cost. Add to this the strong institutional pressure to inflate the "success" of these online methods given the substantial investment being made by "self-interested advocates of online learning (11)" and I see trouble ahead for the reality based community. Of course academics in regard to self interest are not pure as the driven snow either, but I am pretty sure that the resources behind the drive to MOOC and IT education dwarf those of the academy. At any rate, this study is short and is worth a peek for it demonstrates the power of a profitable idea and the gullibility of our academic leaders. The point seems to be to join the stampede, evidence be damned.

Monday, April 8, 2013

Give Nim, Give Me



(Photo credit: Herb Terrace)

Einstein was a very late talker. “The soup is too hot”, as the legend has it, were his first words at the very ripe age of three. The boy genius hadn't seen anything worth commenting on.

The credulity of such tales aside, they do contain a kernel of truth: a child doesn’t have to say something, anything, just because he can. This poses a challenge for the study of child language, since naturalistic production is often the only, and certainly the most accessible, data at hand. A child’s linguistic knowledge may not be fully reflected in their speech, which we have known since Lila’s deconstruction of the telegraphic stage. Some expressions may not show up because we haven’t waited long enough, while others—an extraction violation, for instance—will never be said for they are unsayable.

In recent years, what-you-say-is-what-you-know appears to be gaining popularity, as interests in usage based theories of language are on the rise. Here is a warmup. The expression "give me" is proposed as a frozen phrase (Lieven et al. 1992, Tomasello 2010), rather than syntactically composed, spawning cottage industries such as "formulaic languages", which some regard as a transient stage in language evolution (Wray 1998). True, "give" and "me" make a good tag team: "give me freedom", "give me cheese", "give me now" … “gimme coffee” (an old favorite of mine), and they dwarf other combinations.  Take the speech of Adam, Eve and Sarah from Roger Brown's classic study: the frequencies of "give me", "give him", and "give her" are:

95 (93 give me, 2 gimme): 15 (give him): 12 (give her), or 7.91 : 1.23 : 1

So “give me” does seem especially formulaic ... right? Well, not if you check the frequencies of "me", "him", and "her" from the same three kids:

2870 (me) : 466 (him) : 364 (her), or 7.88 : 1.28 : 1

Nothing much can be concluded from these six numbers but there seems to be pretty good support for the null hypothesis that “give” and pronouns combine completely independently. The Brown data has been around for forty years; it's just nobody had bothered to check. (Use the grep, Luke.)

Nowadays everyone does statistics but we still need reasonable hypotheses to test for and against. Usage based theories have plenty of p-values: one can easily show that the frequencies of "give me/him/her" are statistically significantly different from "chance"--but what is "chance"? If we know anything about the statistics of language, it is that language is not "chance" (Zipf 1949). To make the argument against grammar, one would need to show, at the minimum, that the observed distribution in child language is statistically inconsistent with the predicted distribution of a grammar.  Judiciously chosen null hypotheses are needed, not gut feelings: so long to the "gimme" myth.

A few years ago, Virginia Valian came to Penn to give a talk. It concerned the distribution of determiner-noun combinations in child English. Virginia was the first to show that English children’s determiners are virtually error free (1986), thereby providing evidence for an abstract grammar. Not so quick, the usage-based folks say, because the absence of errors could be the result of children memorizing specific word combinations from adult speech, which would also be error free. (I fully endorse such skepticism.) We need some other statistical benchmarks to show the presence of grammar. 

Diversity is a popular measurement. Suppose there is a rule DP→DN, where D is either a or the, and N stands for a singular noun, yielding "a/the car", "the/a pizza", etc. Shouldn't the interchangeability of "a" and "the", per grammar, be reflected in the diversity of nouns that appear with both of them?  Young children's determiner use, however, only shows 20-40% of diversity (Pine & Lieven 1996); perhaps they just memorize determiner-noun combinations from the adult input (Tomasello 2000, Cognition). 

Along with Stephanie Solt and John Stewart, Virginia showed that mothers' speech contains comparable, and comparably low, diversity measures as their toddlers’ (2009, J. Child Language). After her talk, I pulled out some numbers from the Brown Corpus: not Roger Brown, but the collection of English print materials at Brown University, the grandmother of all modern linguistic corpora. Only 25% of singular nouns that combine with either "a" and "the" combine with both. That's lower than some child samples from Pine & Lieven (1996), so two year olds have a better command of English grammar than professional writers. Now that is absurd. 

One reaction would be to abandon the premise that syntactic diversity is a direct reflection of grammatical complexity. Not a bad idea, and much of the purported evidence for usage based theory vanishes. Another reaction would be to go for the extra credit, by characterizing the statistical profile of syntactic diversity that can be expected from a grammar. If the child used 100 distinct nouns, and paired them with either "a" or "the" 500 times, how many of the 100 will be paired with both, assuming the rule DP→DN is at work? Virginia's work was inspirational. I was also knee deep in Zipfian waters, thanks to the work of Erwin Chan, Constantine Lignos and my colleague Mitch Marcus. They showed that pretty much everywhere you look--words, lemmas, morphological inflections, syntactic rules--language follows Zipf-like distributions, which can be exploited for fun and benefit.

If a sample contains 100 nouns (types), then a good many of them must occur only once since they will inevitably fall on Zipf's long and flat tail: these fellows will never get to meet both determiners.  Even for those that do show up multiple times, they may still be monogamous, just as when you toss a fair coin 3 times, it may land on heads 3 times in a roll. And grammar is no fair coin. Nouns tend to have a favored determiner, even though both combinations are possible. For instance, "the bathroom" is more commonly used than "a bathroom" but we say "a bath" a lot more often than "the bath". These imbalances are probably not a matter of grammar, which presumably does not encode the frequency of bodily needs, but they will conspire to produce low syntactic diversity, and thus the impression of grammatical absence.

After a bit of probability theory exercise [1], we can use a formula to calculate the expected diversity from the sample and vocabulary size (e.g., 500 and 100). The key here is multiplication--the statistical hallmark of independence--of the noun probabilities with determiner-noun combination probabilities, both of which can be well approximated by Zipf’s law. I was surprised to see how well it worked, and in fact had to learn new statistics just to be sure. We mostly use statistics to show one set of values and another (e.g., experimental results vs. "chance") are statistically different, but being different is not the same as being the same. Lin's concordance correlation coefficient (there is an R package, of course), first invented in biostatistics to verify drug effectiveness across trials, confirmed the observation.   In other words, children's syntactic diversity appears exactly what one might expect from a grammar rule, once the general statistical properties of language are taken into account. [2]

Someday we may have a bunch of these statistical profilers, like what evolutionary geneticists use to detect natural selection at the molecular level. Let me make very clear what this work does and does not show. It does show that at least one part of child language makes use of an abstract rule of grammar but it does not mean that all parts of child language do. It does show that children can merge but it does not tell us how they learn what to merge with. It does show the presence of grammatical ability in very young children, but it does not say how that ability got there in the first place, ontogetically or phylogenetically.

Which brings me to Nim Chimpsky and the evolution of language. The continuity between primate language and early child language is believed to hold “the most promising guide to what happened in language evolution” (Hurford 2011, p590), presumably on the apparent formulaic similarities between them. If the numbers worked out for children, who seem to have a grammar after all, perhaps Nim is due for a similar upgrade?  Whatever one thinks of Project Nim--I had to fight back tears--it produced the only publicly available corpus of primate language. Nim acquired about 125 signs of ASL, and produced thousands of multiple sign combinations, the vast majority of which being two sign combinations (Terrace 1979, Nim) These have been described as rule-like constructions, each consisting of two closed class functors such as “give” and “more,” along with open class items such as “apple,” “Nim,” or “eat.”  Signs do not combine with uniform frequency either, with “eat”, “banana”, “me”, “Nim" etc. among the predictable favorites. What’s Nim’s syntactic diversity if he combined signs under a rule? Run the numbers: the poor guy didn’t seem to have a grammar, just as his trainers concluded (Terrace et al. 1979).

Syntactic diversity in human language, usage based learning and Nim Chimpsky.



Moral for the day: the null hypothesis, once properly formulated, may come back to bite your statistical hand.  All very exciting. When I explained this work to some of my non-linguist friends (I do have a few!), their reaction was one of surprise, though not the kind I had in mind. “Why would anyone think kids learn language by copying us?  Just this morning, Maggie said ___”, to be filled by one of the darndest things kids say.  They do wonder about vocabulary, boys vs. girls, and bilingualism, but no one is remotely concerned about the combinatorics of grammar that are, literally, screaming in their faces. Perhaps linguists do worry too much. 

[1] Thanks to Ruochuan Liu and Qiuye Zhao for spotting an error early on.
[2] Could a usage cum memory-retrieval model account for the same finding? I don’t think one knows for sure,  since it has been difficult to pin down the mechanics of usage based learning so it’s unclear what quantitative predictions it makes. I won’t dwell on the matter here but refer you to the paper under discussion, where a concrete proposal (Tamales 2000, Cognitive Linguistics, p77) is tested but came up short. 

More on MOOCs

Sorry for the earlier bad link. Thanks to Marc for mentioning it to me. I think it is now fixed.

Here is a link to a more elaborate discussion of MOOCs, for those interested in the issue and its possible consequences for linguistics teaching.  I tend to think that there are places where IT technology could really enhance education, but they are unlikely to be substitutes (in the best of worlds, complements) for face time where most of the inauguration into "thinking" will take place.  But, this is in the best of worlds, not the one that we are living in and the one that excites our academic "leaders." Most of our presidents, provosts, deans, etc are under pressure to control costs (or so I believe) and in this context the IT appeal is, I believe, pretty obvious.  At any rate, take a look at the Little post and the Bowen papers (if this topic interests you) for they (especially the latter) will undoubtedly prove to be influential in the coming debate (we should be so lucky as to actually have a debate).

Sunday, April 7, 2013

Operationalizing the Strong Minimalist Thesis


In the sciences, it takes a lot of work for a new idea to take hold.  Aside from a modicum of conceptual clarity, a conceptual innovation must be operationalized, and this requires offering canonical or paradigmatic models of its application. I mention this because for me one of the recurring difficulties with the Minimalist Program has been figuring out what makes any given proposal/analysis minimalist (i.e. the path from Minimalist Program to Minimalist Theory is often obscure). So, while being a minimalist is all very nice (e.g. it greatly (minimally?) enhances my self-esteem), what is more absent than it should be are clear examples of doing minimalism; examples of what makes a particular analysis/proposal minimalist or how the abstract leading ideas get concretized in everyday work. Fortunately, I have recently read some papers that, I believe, can serve as parade cases of minimalist thinking and provide clear examples of how the Strong Minimalist Thesis can be incarnated. They are the topic of today’s sermon.

First, what’s the Strong Minimalist Thesis (SMT)?  It is the claim that “language is an optimal solution” to interface conditions:

…the human faculty of language FL [is] an optimal solution to minimal design specifications, conditions that must be satisfied for language to be usable at all…for each language L (a state of FL), the expressions generated by L must be “legible” to systems that access these objects at the interface between FL and external systems – external to FL, internal to the person. (DbP 1).

The systems that L interfaces with use its generated objects. The SMT proposes that the objects of L are well designed for the cognitive interfaces that use them in doing what they do. Put another way, the generated objects can be used as is (i.e. without further alteration) to do what needs getting done (viz. the information they contain need not be further repackaged for the interfaces to use them for whatever tasks they set their “hands” to.)[1] What the hell does this mean?  Two papers by Pietroski, Lidz, Halberda and Hunter (PLHH) (here and here) provide a useful concrete model for interpreting these abstract claims. Their discussion centers on the correct representation of the meaning of most. Here’s what PLHH do.

The papers are interested in figuring out how most sentences affect visual perception (i.e. how the visual system uses grammatical information in making a visual judgment). Specifically, how does someone who hears (1) judge whether a certain presented array of dots verifies (1).

(1)  Most of the dots are blue

It’s a given that (1) is true iff the number of blue dots exceeds the number of non-blue dots. The question is how does one represent the italicized information and does it matter to what people do.  The problem becomes interesting in that there are several ways of representing the quantity information in (1) that are not intensionally equivalent (viz. they use different predicates and different relations) despite being truth functionally the same (in Frege speak: they involve different routes to the same truth value). Here are three possible representations for the meaning of most.

            (2)       a. OneToOnePlus*: [{x: D (x)}, [x: Y (x)}] iff some some set s, s Ì {X: D
    (x)} and OneToOne [s, {x: Y (x)}][2]
                        b. |{x: D (x) & Y (x)}| > {x: D (x) & - Y (x)}|
                        c. |{x: D (x) & Y (x)}| > |{ x: D (x)}| - |{x: D(x) & Y (x)}|

For D= ‘dot’ and Y= ‘blue,’ the (2a) representation carries out the evaluation of the dot scene by pairing the blue with the non-blue dots and seeing if there is at least one extra blue dot left over. The second, in (2b), sees if the size of the set of blue dots is greater than the size of the set of non-blue dots and the third, (2c) sees if the size of the set of blue dots is greater than the size of the set of all the dots minus the set of blue dots.  (2a) differs from the others in using a distinct predicate (viz. OnToOne) while (2b) and (2c) differ in the sets are compared, the former directly calculating the set of non-blue dots, the latter never directly numerically evaluating this set (i.e. it does so indirectly by directly subtracting the blue dots from the entire set of dots).

PLHH reason as follows: There are two possibilities when speakers are asked to evaluate dot scenes on hearing (1):  (i) the visual/counting system might find some of these representations more congenial than the others. (ii) Or, the visual/counting system might use any of these three truth functionally equivalent representations given the right circumstances.  If (i) holds then this visual-counting interface favors one representation over the other two. Why? Because “linguistic meanings are related the cognitive systems that are used to evaluate sentences for truth and falsity.” More specifically: “a declarative sentence S is semantically associated with a canonical procedure for determining whether S is true…[and] competent speakers are biased towards strategies that directly reflect canonical specifications of truth conditions.” They dub this thesis the Interface Transparency Thesis (ITT). Put in slightly more “minimalist” terms: in a well designed grammar, its products will supply the information the interfaces need in a way transparent to those needs. More specifically, the interfaces will “use” the kinds of information the linguistic structure directly encodes.  Or, the information that the grammatical representations encode and the information that the interface uses is one and the same.

Before going on, observe that the ITT provides a useful (and, as PLHH demonstrate, usable) interpretation of “optimal solution to interface conditions”: for a given interface (i.e. system that uses language) how transparent is the mapping between the information made available from L and the information that the interface uses to do what it does?  SMT amounts to the hypothesis that a strong transparency holds between the information as coded in L and the information these various interfaces exploit to do what they do.  If considerable transparency holds then SMT is vindicated. If not, not.[3] So among other things, one very useful contribution of PLHH’s papers is that they provide a substantive yet manageable interpretation of the SMT. But that is not all.

I would not be going through all of this were it not the case that PLHH demonstrate that not all representations of most are created equal.  They provide very good reasons to conclude that (2c) is the right semantic representation of most (or, is clearly superior to (2a,b)). Demonstrating this is conceptually simple but the argument is quite complicated and rich. Showing (2c) is superior to (2a/b) requires knowing a lot about properties of the interface, in this case knowing how people count (humans use two counting systems with different properties) and how people “count” what they see.  Luckily, this is a well-studied domain of visual perception (Halberda (one of the ‘H’s) has done a lot of basic work on how humans do this) and so it is possible to contrive visual dot scenes that would favor one or another of the representational formats in (2) and see what happens. The answer is that the information in (2c) is what humans compute, even when things are visually arranged so that (2a) or (2b) would be simple to apply.[4] The bottom line: humans have a bias for (2c) and the source of this bias is reasonably attributed to the fact that the visual-counting system likes the information as represented in (2c), as per the ITT.

Assume that this is correct. Can we go further and explain the properties of the representation (2c) in terms of the properties of this interface?  Let me be clear: PLHH show that one representation is preferred to others. We can attribute this to the meaning of most being (2c) coupled with the ITT.  Given this we can ask the next question: is the representational format of (2c) explicable in terms of the properties of this interface? Recall, the SMT suggests that FL (a late emerging system) is the “optimal solution” to interface requirements. This suggests that the properties of FL are what they are because of the properties of the interfaces that use them.  PLHH show that for some features of (2c) this explanatory chit can be cashed in. Here’s their very interesting argument.

First, note that in (2b,c) different sets are being selected for enumeration (recall (2c) says nothing direct about the non-blue dots).  This said, it’s a fact that humans are very good at selecting positive features in an array (e.g. blue dots or red dots or green dots) but not at negatively specified features (e.g. not-blue dots). This clearly argues against (2b) in that one is directed to select the non-blue dots.  Second, it has been shown that subjects (human adults) “always attend and enumerate the superset of all dots,” which is good news for (2c) as this is a required part of the specified computation. Third, it can be shown that when subjects use the Approximate Number System (ANS), the one used in this task, they can “estimate the cardinality of up to three sets in parallel,” which means that if there are blue dots, red dots, yellow dots, green dots and mauve dots that (2b) could not be used to evaluate (1) in such a scene (i.e. the requirements in (2b) do not scale up very well, whereas those in (2c) do).  In sum, as PLHH put it:

A meaning like [(2c)]…is straightforwardly verified with these resources, since the sets required for verification (one color plus the superset) are easily and automatically attended by the visual system. Moreover, this meaning does not become less plausible as the number of color subsets increases.

In other words, given the ITT in this domain (for which PLHH have provided evidence) and given the properties of the ANS and the visual system, representations like (2c) perfectly fit the structural capacities of the interface.  Thus, the meaning of most as specified in (2c) fits the noted interface specifications to a (SM)T!

The work is gorgeous. But aside from its stand-alone value, it’s really useful for minimalists to contemplate and absorb.  The Interface Transparency Thesis provides a useful concept for investigating how interfaces and grammars “fit.”  If such a fit can be established, it is possible (sometimes) to argue from properties of the interface to properties of the representations.  Minimalist should understand and absorb this two-step tango for it serves to operationalize the SMT, moving it from a frequently annoying slogan to a research problem, something every minimalist should welcome.

Let me end by noting that PLHH are not alone in deploying this argument. Berwick and Weinberg (BW) (here, where the notion ‘transparency’ was also mooted and discussed) develop an earlier version of this argument.[5] Their version of the ITT considers another interface, the parser (i.e. those interfaces that underlie parsing utterances in real time) and asks what grammatical properties would allow for optimal parsing (roughly parsing in linear time). BW showed that parsers with bounded left contexts would serve nicely and argued that grammars that respected some version of cyclicity+subjacency would perfectly fit the bill. So, if we assume that parsers use the structures generated by L to parse then a cyclic+subjacent compliant grammar would be the perfect fit.  The form of argument is exactly the same as that in PLHH, with the relevant interface this time being those that underlie parsing.

So, the upshot: there are now some paradigm cases out there of how to argue for the SMT.  Deploying these arguments requires knowing a lot about grammar and a lot about some interface property. However, as these two cases show, such arguments can be made. Moreover, they can be made convincingly.  It seems that the SMT is not merely a guiding regulative ideal but even one that can be empirically evaluated. Pretty damn good!

The take home message?  One effective way of investigating the SMT is to identify some interface system (the parser, the visual system, the ANS) and see how it uses the grammatical information provided by L.  The SMT leads to the expectation that it uses the information “transparently,” and that the details of how the interface works can explain why the representation looks like it does. This is hard to pull off, for it requires knowing a lot both about the grammar and the interface at issue.  It suggests that future syntacticians will need to have new skill sets and/or be very collaborative.  This will no doubt be demanding. But, hey, who every said that cognitive-biolinguistics would be easy. The most anyone promised was that it would be fun, and, if these cases are any indication, crammed with more than a touch of intellectual beauty as well.




[1] If I understand the notion “covering grammar” correctly, then one might say that the competence grammar is the covering grammar for the relevant interface.  In the best case it is the grammar that every interface uses.
[2] OneToOne is a function pairs individuals in D and Y one to one.
[3] Note the word ‘considerable.’  The relevant evaluation will revolve around some estimation of the degree of transparency and this may be a labile notion.  It may be possible to make these estimations on a case by case basis without having a general measure of transparency.
[4] As PLHH note: humans can in fact apply the predicates in (2a) and (2b) in non-quantificational tasks. Thus their failure to apply them in these “linguistic” contexts cannot be traced to some general human incapacity to deploy them.
[5] In addition Colin Phillips proposal that the grammar is identical to the parser can be interpreted as postulating a very strong transparency assumption for this interface.

The Perils of Parody

There are times when reality severely strains the resources of parody. But some brave souls try and try (god bless their hearts). This one (here) came my way recently. A good friend tells me that s/he found the original far less plausible and much funnier. Judge for yourself.

Floats Like a Butterfly, Stings like a...

There's a new species in town named after a famous linguist. Guess Who? Answer is here.

Monday, April 1, 2013

A Neat New Argument and a Hint of the Future?


Syntacticians often act like parochial snobs.  They are snobs in that some (e.g. me) believe that the kinds of data and explanations offered within syntax are deeper than those offered in other domains. There really is non-trivial syntactic theory and a whole budget of effects that these theories explain.  Syntacticians (e.g. me) can be parochial in believing that only these kinds of questions are worth asking and investigating. A consequence of this is the oft-exhibited habit of evaluating the interest of other linguistic questions in terms of how much they address questions in syntax.  This habit can be especially pronounced when syntacticians consider work in psycholinguistics.  When evaluating psycho work, it is not uncommon for syntacticians to expect psycholinguists to provide evidence/methods for helping to choose among competing cutting edge syntactic alternatives.  As this rarely happens, syntacticians often come to have jaundiced views about the intellectual contributions of psycholinguistic research.

This is clearly a rather perverse way of evaluating psycho research. The issue is not whether psycholinguistics can answer syntactic questions but whether they have their own interesting questions concerning the form and function of FL. Over the years I have made it a habit to sit in on lab meetings of my psycho colleagues and two things have struck me. First, that wrt syntactic theory, to date, psycho techniques have contributed little that purely syntactic methods have not delivered more cleanly and quickly (but see below) and, second, that psycholinguists have found fascinating effects and have developed interesting theories of these effects that greatly expand our understanding of how FL is used in real time, both learning and processing.  Let me elaborate each point just a bit.

A good deal of the questions in language processing and acquisition can proceed just fine in blissful ignorance of the latest findings in syntax.  For example, studying the acquisition or processing of long distance dependencies will not be greatly affected by whether one assumes that movement is actually an expression of Merge (I-merge) or an independent operation in its own right. The differences between these two conceptions is too fine grained to be captured by (at least current) psycho methods.  However, this does not mean that there is nothing worth studying. For example, regardless of how a WH gets to clause initial position, we have an interesting question: how eagerly does the parser try to find the gap the WH is related to? One possibility is that it waits to find a gap and then tries to relate the WH to it. A second option is that it tries to link the WH immediately on sighting a theta marking host predicate without waiting to see if there is a gap after the predicate to fill.  Thus in a sentence like (1), we can ask if the parser tries to interpret who as complement of tell after/before seeing if there is a gap there.

(1)  Who did you tell Bill about

So given the question ‘how “eager” is the parser?’, the next question is how to study its eagerness?  Via a very interesting phenomenon discovered in the 1980s (Laurie Stowe in 1984), the so-called “Filled Gap Effect” (FGE). If you put people in front of a computer and have them read a sentence word by word and measure how they doe this, it turns out that in a sentence like (1) readers will pause longer at Bill than they would in reading a sentence like (2):

(2)  Who did you tell about Bill

This is a very reliable effect. Interpretation? Readers are trying to thematically interpret who when they get to tell and must rescind this when they get to Bill in (1) but not in (2). In other words, readers “prefer” giving a thematic role to who even at the cost of having to rescind this assignment soon after over waiting to see if there is an available role to assign, even if this means just waiting one word. Conclusion: parsers are very eager to interpret uninterpreted material. And this eagerness can be measured and used to probe how parsers use grammars to construct sentence interpretations in real time, i.e. FGE is a marvelous tool for probing the relation between grammars and parsers.

Let me give an illustrative example. Consider the following question: Given that parsers use grammars what is the relation between competence grammars (ones beloved of syntacticians) and parsing grammars?  A strong position is that there is a very high level of “transparency” between the two.[1] What’s this mean? Well, that the categories and relations that the grammar specifies are identical to those that the parser exploits/respects in real time.  For example, the categories that grammars deploy (e.g. DP, VP, CP) are what the parser tries to recover and the conditions the grammar respects (e.g. minimality, c-command, subjacency) the parser does as well. A good deal of work in parsing over the last 15 years has been aimed at specifying the degree of transparency between competence grammars and parsing grammars.  For example, psycholinguists have investigated whether parsers respect c-command in trying to find a bound anaphors possible antecedents and whether parsers display cross-over effects.[2] My colleague, Colin Phillips, has a really beautiful set of results bearing on how parsers “do” movement, particularly does gap filling “obey” islands (see here).  The answer is “yes.” How does he know? Building on earlier work by others, he shows that the FGE only appears if the “gap” sits in a possible movement site.  Gaps within islands do not trigger FGEs (unless they are licensable parasitic gaps (this is Colin’s great find)).  Gaps generable by movement do trigger FGEs.  The argument is subtle and well worth reading (if you’ve read it already, read it again to your kids; it makes for a lively bedtime experience!). All this makes sense if parsing grammars and competence grammars are (largely) the same. Ergo…

Note that these psycho results build on what syntactic research has revealed about competence. It uses these results to address a related very interesting question, viz. the Transparency Thesis (TT) and it does so by exploiting a psycho probe, viz. FGE, manifest in online reading tasks.  Most interestingly in my view, as TT becomes more and more empirically grounded (and the evidence in its favor is already pretty good IMO), syntacticians may finally get what they’ve been asking for: psycholinguistic constraints on adequate competence theories (Syntacticians, careful what you ask for lest you get it!).  After all, if TT is right, then syntactic theories that fail to support transparent parsing grammars should be less valued than those that do. In other words if TT proves empirically tenable, then online psycho results will prove highly relevant to evaluating the empirical adequacy of proposed competence grammars. 

How far away is this day? Well, I want to end by presenting a possible glimpse of the future.  In some recent work Shevaun Lewis, Dave Kush and Brad Larson (LKL) have used FGEs to probe the syntactic derivation of constructions like (3) (see here for some slides):

(3)  What and when will we eat

Not surprisingly, these coordinated WH questions have rather elaborate syntactic properties. I will not detail them for you (read the slides), except to say that LKL are led to analyze these by treating the two WHs rather differently. LKL propose that the inner WH lands in clause initial position via movement while the outer one is base generated there. The evidence for this conclusion is two-fold. First, there is an acceptability contrast between (4a) and (4b), the latter being quite a bit worse than the former (and yes they ran the relevant acceptability judgment studies to show this). Second, the what in (4a) fails to induce an FGE.  If FGEs are diagnostic of movement dependencies (as above) then the absence of these in (4a) is just what we would expect, and apparently receive. To my knowledge this is the first time a technique borrowed from psycho has been pressed into service to support a novel syntactic conclusion.  Terrific!

(4)  a. What and when will we eat something
b. When and what will we eat something

It is the sign of a progressing research program that novel questions and techniques keep springing up. The aim is to conciliate these producing a bigger and bigger coherent picture.  To date, in my view, a great deal of what we have discovered about how FL does syntax has come from the careful analysis of natural language grammars. The LKL results are signaling a slightly different future: I have a dream that is deeply rooted in the Generative enterprise. I have a dream that one day linguistics will rise up and live the true meaning of its biolinguistic roots. I have a dream that one day syntacticians and psycholinguists (and eventually neuroscientists) will use each other’s work to strongly constrain their common research project of understanding how FL is structured and how it is used. I have a dream, and LKL provides a glimpse of that wonderful and glorious future.



[1] The term ‘transparency’ is from Berwick and Weinberg 1984.
[2] For a good review of the first, c.f. Dillon (here). All I have to offer for the second is some hot off the presses work by Kush, Lidz and Phillips. Poster from latest CUNY (here).