I chased down on of the links from the Dan Little piece I posted yesterday to take a quick peek at the data backing up the recent rush to MOOCing the classroom. This paper by Bowen and Lack (BL) reviews the data, such as there is, reviewing studies on the topic. As BL note, there was not a lot to go on. Right now, MOOCs are a faith based initiative, or as BL put it: "As of the moment, a priori argument will have to continue to bear much of the weigh" "in determining how much more of an investment, and in what form. should be made in this field (4)." They go on to say that there is little "conclusive evidence" about the cost effectiveness of online courses, in fact little good evidence at all. They also observe that much of this is driven by a desire for cost reduction: "achieving the same outcomes at half the cost per student should be seen as a great victory (11)." My worry is that as the outcomes are very hard to measure, as BL show, the focus will entirely shift to reducing cost. Add to this the strong institutional pressure to inflate the "success" of these online methods given the substantial investment being made by "self-interested advocates of online learning (11)" and I see trouble ahead for the reality based community. Of course academics in regard to self interest are not pure as the driven snow either, but I am pretty sure that the resources behind the drive to MOOC and IT education dwarf those of the academy. At any rate, this study is short and is worth a peek for it demonstrates the power of a profitable idea and the gullibility of our academic leaders. The point seems to be to join the stampede, evidence be damned.
Tuesday, April 9, 2013
Monday, April 8, 2013
Give Nim, Give Me
(Photo credit: Herb Terrace)
Einstein was a very late talker. “The soup is too hot”, as the legend has it, were his first words at the very ripe age of three. The boy genius hadn't seen anything worth commenting on.
The credulity of such tales aside, they do contain a kernel of truth: a child doesn’t have to say something, anything, just because he can. This poses a challenge for the study of child language, since naturalistic production is often the only, and certainly the most accessible, data at hand. A child’s linguistic knowledge may not be fully reflected in their speech, which we have known since Lila’s deconstruction of the telegraphic stage. Some expressions may not show up because we haven’t waited long enough, while others—an extraction violation, for instance—will never be said for they are unsayable.
In recent years, what-you-say-is-what-you-know appears to be gaining popularity, as interests in usage based theories of language are on the rise. Here is a warmup. The expression "give me" is proposed as a frozen phrase (Lieven et al. 1992, Tomasello 2010), rather than syntactically composed, spawning cottage industries such as "formulaic languages", which some regard as a transient stage in language evolution (Wray 1998). True, "give" and "me" make a good tag team: "give me freedom", "give me cheese", "give me now" … “gimme coffee” (an old favorite of mine), and they dwarf other combinations. Take the speech of Adam, Eve and Sarah from Roger Brown's classic study: the frequencies of "give me", "give him", and "give her" are:
95 (93 give me, 2 gimme): 15 (give him): 12 (give her), or 7.91 : 1.23 : 1
So “give me” does seem especially formulaic ... right? Well, not if you check the frequencies of "me", "him", and "her" from the same three kids:
2870 (me) : 466 (him) : 364 (her), or 7.88 : 1.28 : 1
Nothing much can be concluded from these six numbers but there seems to be pretty good support for the null hypothesis that “give” and pronouns combine completely independently. The Brown data has been around for forty years; it's just nobody had bothered to check. (Use the grep, Luke.)
Nowadays everyone does statistics but we still need reasonable hypotheses to test for and against. Usage based theories have plenty of p-values: one can easily show that the frequencies of "give me/him/her" are statistically significantly different from "chance"--but what is "chance"? If we know anything about the statistics of language, it is that language is not "chance" (Zipf 1949). To make the argument against grammar, one would need to show, at the minimum, that the observed distribution in child language is statistically inconsistent with the predicted distribution of a grammar. Judiciously chosen null hypotheses are needed, not gut feelings: so long to the "gimme" myth.
A few years ago, Virginia Valian came to Penn to give a talk. It concerned the distribution of determiner-noun combinations in child English. Virginia was the first to show that English children’s determiners are virtually error free (1986), thereby providing evidence for an abstract grammar. Not so quick, the usage-based folks say, because the absence of errors could be the result of children memorizing specific word combinations from adult speech, which would also be error free. (I fully endorse such skepticism.) We need some other statistical benchmarks to show the presence of grammar.
Diversity is a popular measurement. Suppose there is a rule DP→DN, where D is either a or the, and N stands for a singular noun, yielding "a/the car", "the/a pizza", etc. Shouldn't the interchangeability of "a" and "the", per grammar, be reflected in the diversity of nouns that appear with both of them? Young children's determiner use, however, only shows 20-40% of diversity (Pine & Lieven 1996); perhaps they just memorize determiner-noun combinations from the adult input (Tomasello 2000, Cognition).
Along with Stephanie Solt and John Stewart, Virginia showed that mothers' speech contains comparable, and comparably low, diversity measures as their toddlers’ (2009, J. Child Language). After her talk, I pulled out some numbers from the Brown Corpus: not Roger Brown, but the collection of English print materials at Brown University, the grandmother of all modern linguistic corpora. Only 25% of singular nouns that combine with either "a" and "the" combine with both. That's lower than some child samples from Pine & Lieven (1996), so two year olds have a better command of English grammar than professional writers. Now that is absurd.
One reaction would be to abandon the premise that syntactic diversity is a direct reflection of grammatical complexity. Not a bad idea, and much of the purported evidence for usage based theory vanishes. Another reaction would be to go for the extra credit, by characterizing the statistical profile of syntactic diversity that can be expected from a grammar. If the child used 100 distinct nouns, and paired them with either "a" or "the" 500 times, how many of the 100 will be paired with both, assuming the rule DP→DN is at work? Virginia's work was inspirational. I was also knee deep in Zipfian waters, thanks to the work of Erwin Chan, Constantine Lignos and my colleague Mitch Marcus. They showed that pretty much everywhere you look--words, lemmas, morphological inflections, syntactic rules--language follows Zipf-like distributions, which can be exploited for fun and benefit.
If a sample contains 100 nouns (types), then a good many of them must occur only once since they will inevitably fall on Zipf's long and flat tail: these fellows will never get to meet both determiners. Even for those that do show up multiple times, they may still be monogamous, just as when you toss a fair coin 3 times, it may land on heads 3 times in a roll. And grammar is no fair coin. Nouns tend to have a favored determiner, even though both combinations are possible. For instance, "the bathroom" is more commonly used than "a bathroom" but we say "a bath" a lot more often than "the bath". These imbalances are probably not a matter of grammar, which presumably does not encode the frequency of bodily needs, but they will conspire to produce low syntactic diversity, and thus the impression of grammatical absence.
After a bit of probability theory exercise [1], we can use a formula to calculate the expected diversity from the sample and vocabulary size (e.g., 500 and 100). The key here is multiplication--the statistical hallmark of independence--of the noun probabilities with determiner-noun combination probabilities, both of which can be well approximated by Zipf’s law. I was surprised to see how well it worked, and in fact had to learn new statistics just to be sure. We mostly use statistics to show one set of values and another (e.g., experimental results vs. "chance") are statistically different, but being different is not the same as being the same. Lin's concordance correlation coefficient (there is an R package, of course), first invented in biostatistics to verify drug effectiveness across trials, confirmed the observation. In other words, children's syntactic diversity appears exactly what one might expect from a grammar rule, once the general statistical properties of language are taken into account. [2]
Someday we may have a bunch of these statistical profilers, like what evolutionary geneticists use to detect natural selection at the molecular level. Let me make very clear what this work does and does not show. It does show that at least one part of child language makes use of an abstract rule of grammar but it does not mean that all parts of child language do. It does show that children can merge but it does not tell us how they learn what to merge with. It does show the presence of grammatical ability in very young children, but it does not say how that ability got there in the first place, ontogetically or phylogenetically.
Which brings me to Nim Chimpsky and the evolution of language. The continuity between primate language and early child language is believed to hold “the most promising guide to what happened in language evolution” (Hurford 2011, p590), presumably on the apparent formulaic similarities between them. If the numbers worked out for children, who seem to have a grammar after all, perhaps Nim is due for a similar upgrade? Whatever one thinks of Project Nim--I had to fight back tears--it produced the only publicly available corpus of primate language. Nim acquired about 125 signs of ASL, and produced thousands of multiple sign combinations, the vast majority of which being two sign combinations (Terrace 1979, Nim). These have been described as rule-like constructions, each consisting of two closed class functors such as “give” and “more,” along with open class items such as “apple,” “Nim,” or “eat.” Signs do not combine with uniform frequency either, with “eat”, “banana”, “me”, “Nim" etc. among the predictable favorites. What’s Nim’s syntactic diversity if he combined signs under a rule? Run the numbers: the poor guy didn’t seem to have a grammar, just as his trainers concluded (Terrace et al. 1979).
![]() |
| Syntactic diversity in human language, usage based learning and Nim Chimpsky. |
Moral for the day: the null hypothesis, once properly formulated, may come back to bite your statistical hand. All very exciting. When I explained this work to some of my non-linguist friends (I do have a few!), their reaction was one of surprise, though not the kind I had in mind. “Why would anyone think kids learn language by copying us? Just this morning, Maggie said ___”, to be filled by one of the darndest things kids say. They do wonder about vocabulary, boys vs. girls, and bilingualism, but no one is remotely concerned about the combinatorics of grammar that are, literally, screaming in their faces. Perhaps linguists do worry too much.
[1] Thanks to Ruochuan Liu and Qiuye Zhao for spotting an error early on.
[2] Could a usage cum memory-retrieval model account for the same finding? I don’t think one knows for sure, since it has been difficult to pin down the mechanics of usage based learning so it’s unclear what quantitative predictions it makes. I won’t dwell on the matter here but refer you to the paper under discussion, where a concrete proposal (Tamales 2000, Cognitive Linguistics, p77) is tested but came up short.
More on MOOCs
Sorry for the earlier bad link. Thanks to Marc for mentioning it to me. I think it is now fixed.
Here is a link to a more elaborate discussion of MOOCs, for those interested in the issue and its possible consequences for linguistics teaching. I tend to think that there are places where IT technology could really enhance education, but they are unlikely to be substitutes (in the best of worlds, complements) for face time where most of the inauguration into "thinking" will take place. But, this is in the best of worlds, not the one that we are living in and the one that excites our academic "leaders." Most of our presidents, provosts, deans, etc are under pressure to control costs (or so I believe) and in this context the IT appeal is, I believe, pretty obvious. At any rate, take a look at the Little post and the Bowen papers (if this topic interests you) for they (especially the latter) will undoubtedly prove to be influential in the coming debate (we should be so lucky as to actually have a debate).
Here is a link to a more elaborate discussion of MOOCs, for those interested in the issue and its possible consequences for linguistics teaching. I tend to think that there are places where IT technology could really enhance education, but they are unlikely to be substitutes (in the best of worlds, complements) for face time where most of the inauguration into "thinking" will take place. But, this is in the best of worlds, not the one that we are living in and the one that excites our academic "leaders." Most of our presidents, provosts, deans, etc are under pressure to control costs (or so I believe) and in this context the IT appeal is, I believe, pretty obvious. At any rate, take a look at the Little post and the Bowen papers (if this topic interests you) for they (especially the latter) will undoubtedly prove to be influential in the coming debate (we should be so lucky as to actually have a debate).
Sunday, April 7, 2013
Operationalizing the Strong Minimalist Thesis
In the sciences, it takes a lot of work for a new idea to
take hold. Aside from a modicum of conceptual
clarity, a conceptual innovation must be operationalized,
and this requires offering canonical or paradigmatic models of its application.
I mention this because for me one of the recurring difficulties with the
Minimalist Program has been figuring out what makes any given proposal/analysis
minimalist (i.e. the path from Minimalist Program to Minimalist Theory is often
obscure). So, while being a minimalist
is all very nice (e.g. it greatly (minimally?) enhances my self-esteem), what
is more absent than it should be are clear examples of doing minimalism; examples of what makes a particular
analysis/proposal minimalist or how the abstract leading ideas get concretized
in everyday work. Fortunately, I have recently read some papers that, I
believe, can serve as parade cases of minimalist thinking and provide clear
examples of how the Strong Minimalist Thesis can be incarnated. They are the
topic of today’s sermon.
First, what’s the Strong Minimalist Thesis (SMT)? It is the claim that “language is an optimal
solution” to interface conditions:
…the human faculty of language FL
[is] an optimal solution to minimal design specifications, conditions that must
be satisfied for language to be usable at all…for each language L (a state of
FL), the expressions generated by L must be “legible” to systems that access these
objects at the interface between FL and external systems – external to FL,
internal to the person. (DbP 1).
The systems that L interfaces with use its generated objects. The SMT proposes that the objects of L
are well designed for the cognitive interfaces that use them in doing what they
do. Put another way, the generated objects can be used as is (i.e. without further alteration) to do what needs getting
done (viz. the information they contain need not be further repackaged for the
interfaces to use them for whatever tasks they set their “hands” to.)[1]
What the hell does this mean? Two papers
by Pietroski, Lidz, Halberda and Hunter (PLHH) (here and here) provide a useful
concrete model for interpreting these abstract claims. Their discussion centers
on the correct representation of the meaning of most. Here’s what PLHH do.
The papers are interested in figuring out how most sentences affect visual perception
(i.e. how the visual system uses grammatical information in making a visual
judgment). Specifically, how does someone who hears (1) judge whether a certain
presented array of dots verifies (1).
(1) Most
of the dots are blue
It’s a given that (1) is true iff the number of blue dots exceeds the number of non-blue dots. The
question is how does one represent the italicized information and does it
matter to what people do. The problem
becomes interesting in that there are several ways of representing the quantity
information in (1) that are not intensionally equivalent (viz. they use
different predicates and different relations) despite being truth functionally
the same (in Frege speak: they involve different routes to the same truth
value). Here are three possible representations for the meaning of most.
(2) a. OneToOnePlus*: [{x: D
(x)}, [x: Y
(x)}] iff some some set s, s Ì {X: D
b.
|{x: D
(x) & Y
(x)}| > {x: D
(x) & - Y
(x)}|
c. |{x: D (x) & Y
(x)}| > |{ x: D
(x)}| - |{x: D(x)
& Y
(x)}|
For D= ‘dot’ and Y= ‘blue,’ the (2a) representation carries out the
evaluation of the dot scene by pairing the blue with the non-blue dots and
seeing if there is at least one extra blue dot left over. The second, in (2b),
sees if the size of the set of blue dots is greater than the size of the set of
non-blue dots and the third, (2c) sees if the size of the set of blue dots is
greater than the size of the set of all the dots minus the set of blue
dots. (2a) differs from the others in
using a distinct predicate (viz. OnToOne) while (2b) and (2c) differ in the
sets are compared, the former directly calculating the set of non-blue dots,
the latter never directly numerically evaluating this set (i.e. it does so indirectly by directly subtracting the blue dots from the entire set of dots).
PLHH reason as follows: There are two possibilities when
speakers are asked to evaluate dot scenes on hearing (1): (i) the visual/counting system might find
some of these representations more congenial than the others. (ii) Or, the
visual/counting system might use any of these three truth functionally
equivalent representations given the right circumstances. If (i) holds then this visual-counting interface
favors one representation over the other two. Why? Because “linguistic meanings
are related the cognitive systems that are used to evaluate sentences for truth
and falsity.” More specifically: “a declarative sentence S is semantically
associated with a canonical procedure for determining whether S is true…[and]
competent speakers are biased towards strategies that directly reflect canonical specifications of truth conditions.”
They dub this thesis the Interface
Transparency Thesis (ITT). Put in slightly more “minimalist” terms: in a
well designed grammar, its products will supply the information the interfaces
need in a way transparent to those needs. More specifically, the interfaces
will “use” the kinds of information the linguistic structure directly encodes. Or, the information that the grammatical
representations encode and the information that the interface uses is one and
the same.
Before going on, observe that the ITT provides a useful (and,
as PLHH demonstrate, usable)
interpretation of “optimal solution to interface conditions”: for a given
interface (i.e. system that uses language) how transparent is the mapping
between the information made available from L and the information that the
interface uses to do what it does? SMT amounts
to the hypothesis that a strong transparency holds between the information as
coded in L and the information these various interfaces exploit to do what they
do. If considerable transparency holds then
SMT is vindicated. If not, not.[3]
So among other things, one very useful contribution of PLHH’s papers is that
they provide a substantive yet manageable interpretation of the SMT. But that
is not all.
I would not be going through all of this were it not the
case that PLHH demonstrate that not all representations of most are created equal. They
provide very good reasons to conclude that (2c) is the right semantic
representation of most (or, is
clearly superior to (2a,b)). Demonstrating this is conceptually simple but the
argument is quite complicated and rich. Showing (2c) is superior to (2a/b)
requires knowing a lot about properties of the interface, in this case knowing
how people count (humans use two counting systems with different properties)
and how people “count” what they see. Luckily,
this is a well-studied domain of visual perception (Halberda (one of the ‘H’s)
has done a lot of basic work on how humans do this) and so it is possible to
contrive visual dot scenes that would favor one or another of the
representational formats in (2) and see what happens. The answer is that the
information in (2c) is what humans compute, even
when things are visually arranged so that (2a) or (2b) would be simple to apply.[4]
The bottom line: humans have a bias for (2c) and the source of this bias is
reasonably attributed to the fact that the visual-counting system likes the
information as represented in (2c), as per the ITT.
Assume that this is correct. Can we go further and explain the properties of the
representation (2c) in terms of the properties of this interface? Let me be clear: PLHH show that one
representation is preferred to others. We can attribute this to the meaning of most being (2c) coupled with the
ITT. Given this we can ask the next
question: is the representational format of (2c) explicable in terms of the
properties of this interface? Recall, the SMT suggests that FL (a late emerging
system) is the “optimal solution” to interface requirements. This suggests that
the properties of FL are what they are because of the properties of the
interfaces that use them. PLHH show that
for some features of (2c) this
explanatory chit can be cashed in. Here’s their very interesting argument.
First, note that in (2b,c) different sets are being selected
for enumeration (recall (2c) says nothing direct about the non-blue dots). This said, it’s a fact that humans are very
good at selecting positive features in an array (e.g. blue dots or red dots or
green dots) but not at negatively specified features (e.g. not-blue dots). This
clearly argues against (2b) in that
one is directed to select the non-blue dots.
Second, it has been shown that subjects (human adults) “always attend and enumerate the superset
of all dots,” which is good news for (2c) as this is a required part of the
specified computation. Third, it can be shown that when subjects use the
Approximate Number System (ANS), the one used in this task, they can “estimate
the cardinality of up to three sets in parallel,” which means that if there are
blue dots, red dots, yellow dots, green dots and mauve dots that (2b) could not
be used to evaluate (1) in such a scene (i.e. the requirements in (2b) do not
scale up very well, whereas those in (2c) do).
In sum, as PLHH put it:
A meaning like [(2c)]…is
straightforwardly verified with these resources, since the sets required for
verification (one color plus the superset) are easily and automatically
attended by the visual system. Moreover, this meaning does not become less
plausible as the number of color subsets increases.
In other words, given the ITT in this domain (for which PLHH
have provided evidence) and given the properties of the ANS and the visual
system, representations like (2c) perfectly fit the structural capacities of
the interface. Thus, the meaning of most as specified in (2c) fits the noted
interface specifications to a (SM)T!
The work is gorgeous. But aside from its stand-alone value,
it’s really useful for minimalists to contemplate and absorb. The Interface Transparency Thesis provides a
useful concept for investigating how interfaces and grammars “fit.” If such a fit can be established, it is
possible (sometimes) to argue from properties of the interface to properties of
the representations. Minimalist should
understand and absorb this two-step tango for it serves to operationalize the
SMT, moving it from a frequently annoying slogan to a research problem,
something every minimalist should welcome.
Let me end by noting that PLHH are not alone in deploying
this argument. Berwick and Weinberg (BW) (here, where the notion ‘transparency’
was also mooted and discussed) develop an earlier version of this argument.[5]
Their version of the ITT considers another interface, the parser (i.e. those
interfaces that underlie parsing utterances in real time) and asks what
grammatical properties would allow for optimal parsing (roughly parsing in
linear time). BW showed that parsers with bounded left contexts would serve
nicely and argued that grammars that respected some version of cyclicity+subjacency
would perfectly fit the bill. So, if we assume that parsers use the structures
generated by L to parse then a cyclic+subjacent compliant grammar would be the
perfect fit. The form of argument is
exactly the same as that in PLHH, with the relevant interface this time being those
that underlie parsing.
So, the upshot: there are now some paradigm cases out there
of how to argue for the SMT. Deploying
these arguments requires knowing a lot about grammar and a lot about some interface property. However, as these two
cases show, such arguments can be made. Moreover, they can be made
convincingly. It seems that the SMT is
not merely a guiding regulative ideal but even one that can be empirically
evaluated. Pretty damn good!
The take home message?
One effective way of investigating the SMT is to identify some interface
system (the parser, the visual system, the ANS) and see how it uses the
grammatical information provided by L.
The SMT leads to the expectation that it uses the information
“transparently,” and that the details of how the interface works can explain
why the representation looks like it does. This is hard to pull off, for it
requires knowing a lot both about the
grammar and the interface at issue. It
suggests that future syntacticians will need to have new skill sets and/or be
very collaborative. This will no doubt
be demanding. But, hey, who every said that cognitive-biolinguistics would be
easy. The most anyone promised was that it would be fun, and, if these cases
are any indication, crammed with more than a touch of intellectual beauty as
well.
[1]
If I understand the notion “covering grammar” correctly, then one might say
that the competence grammar is the covering grammar for the relevant
interface. In the best case it is the
grammar that every interface uses.
[3]
Note the word ‘considerable.’ The
relevant evaluation will revolve around some estimation of the degree of transparency and this may be a
labile notion. It may be possible to
make these estimations on a case by case basis without having a general measure of transparency.
[4]
As PLHH note: humans can in fact apply the predicates in (2a) and (2b) in non-quantificational
tasks. Thus their failure to apply them in these “linguistic” contexts cannot
be traced to some general human incapacity to deploy them.
[5]
In addition Colin Phillips proposal that the grammar is identical to the parser
can be interpreted as postulating a very strong transparency assumption for
this interface.
The Perils of Parody
There are times when reality severely strains the resources of parody. But some brave souls try and try (god bless their hearts). This one (here) came my way recently. A good friend tells me that s/he found the original far less plausible and much funnier. Judge for yourself.
Floats Like a Butterfly, Stings like a...
There's a new species in town named after a famous linguist. Guess Who? Answer is here.
Monday, April 1, 2013
A Neat New Argument and a Hint of the Future?
Syntacticians often act like parochial snobs. They are snobs in that some (e.g. me) believe
that the kinds of data and explanations offered within syntax are deeper than
those offered in other domains. There really is non-trivial syntactic theory
and a whole budget of effects that these theories explain. Syntacticians (e.g. me) can be parochial in believing
that only these kinds of questions are worth asking and investigating. A
consequence of this is the oft-exhibited habit of evaluating the interest of other
linguistic questions in terms of how much they address questions in
syntax. This habit can be especially pronounced
when syntacticians consider work in psycholinguistics. When evaluating psycho work, it is not
uncommon for syntacticians to expect psycholinguists to provide evidence/methods
for helping to choose among competing cutting edge syntactic alternatives. As this rarely happens, syntacticians often
come to have jaundiced views about the intellectual contributions of
psycholinguistic research.
This is clearly a rather perverse way of evaluating psycho research.
The issue is not whether psycholinguistics can answer syntactic questions but whether they have their own interesting
questions concerning the form and function of FL. Over the years I have made it
a habit to sit in on lab meetings of my psycho colleagues and two things have struck
me. First, that wrt syntactic theory, to date, psycho techniques have
contributed little that purely syntactic methods have not delivered more
cleanly and quickly (but see below) and, second, that psycholinguists have
found fascinating effects and have developed interesting theories of these
effects that greatly expand our understanding of how FL is used in real time,
both learning and processing. Let me elaborate
each point just a bit.
A good deal of the questions in language processing and
acquisition can proceed just fine in blissful ignorance of the latest findings
in syntax. For example, studying the
acquisition or processing of long distance dependencies will not be greatly
affected by whether one assumes that movement is actually an expression of
Merge (I-merge) or an independent operation in its own right. The differences
between these two conceptions is too fine grained to be captured by (at least current) psycho methods. However, this does not mean that there is
nothing worth studying. For example, regardless of how a WH gets to clause
initial position, we have an interesting question: how eagerly does the parser
try to find the gap the WH is related to? One possibility is that it waits to
find a gap and then tries to relate the WH to it. A second option is that it
tries to link the WH immediately on sighting a theta marking host predicate without
waiting to see if there is a gap after the predicate to fill. Thus in a sentence like (1), we can ask if the parser tries
to interpret who as complement of tell after/before seeing if there is a
gap there.
(1) Who
did you tell Bill about
So given the question ‘how “eager” is the parser?’, the
next question is how to study its eagerness? Via a
very interesting phenomenon discovered in the 1980s (Laurie Stowe in 1984), the
so-called “Filled Gap Effect” (FGE). If you put people in front of a computer and
have them read a sentence word by word and measure how they doe this, it turns
out that in a sentence like (1) readers will pause longer at Bill than they would in reading a
sentence like (2):
(2) Who
did you tell about Bill
This is a very reliable effect. Interpretation? Readers are
trying to thematically interpret who
when they get to tell and must
rescind this when they get to Bill in
(1) but not in (2). In other words, readers “prefer” giving a thematic role to who even at the cost of having to
rescind this assignment soon after over waiting to see if there is an available
role to assign, even if this means just waiting one word. Conclusion: parsers
are very eager to interpret
uninterpreted material. And this eagerness can be measured and used to probe how
parsers use grammars to construct sentence interpretations in real time, i.e. FGE
is a marvelous tool for probing the relation between grammars and parsers.
Let me give an illustrative example. Consider the following
question: Given that parsers use grammars what is the relation between
competence grammars (ones beloved of syntacticians) and parsing grammars? A strong position is that there is a very
high level of “transparency” between the two.[1]
What’s this mean? Well, that the categories and relations that the grammar
specifies are identical to those that the parser exploits/respects in real time. For example, the categories that grammars
deploy (e.g. DP, VP, CP) are what the parser tries to recover and the
conditions the grammar respects (e.g. minimality, c-command, subjacency) the
parser does as well. A good deal of work in parsing over the last 15 years has
been aimed at specifying the degree of transparency between competence grammars
and parsing grammars. For example,
psycholinguists have investigated whether parsers respect c-command in trying
to find a bound anaphors possible antecedents and whether parsers display
cross-over effects.[2]
My colleague, Colin Phillips, has a really beautiful set of results bearing on
how parsers “do” movement, particularly does gap filling “obey” islands (see
here). The answer is “yes.” How does he
know? Building on earlier work by others, he shows that the FGE only appears if
the “gap” sits in a possible movement
site. Gaps within islands do not trigger FGEs (unless they are licensable parasitic gaps (this is
Colin’s great find)). Gaps generable by
movement do trigger FGEs. The argument is subtle and well worth reading
(if you’ve read it already, read it again to your kids; it makes for a lively
bedtime experience!). All this makes sense if parsing grammars and competence
grammars are (largely) the same. Ergo…
Note that these psycho results build on what syntactic
research has revealed about competence. It uses these results to address a
related very interesting question, viz. the Transparency Thesis (TT) and it
does so by exploiting a psycho probe, viz. FGE, manifest in online reading
tasks. Most interestingly in my view, as
TT becomes more and more empirically grounded (and the evidence in its
favor is already pretty good IMO), syntacticians may finally get what they’ve been
asking for: psycholinguistic constraints on adequate competence theories
(Syntacticians, careful what you ask for lest you get it!). After all, if TT is right, then syntactic
theories that fail to support transparent parsing grammars should be less
valued than those that do. In other words if TT proves empirically tenable,
then online psycho results will prove highly relevant to evaluating the
empirical adequacy of proposed competence grammars.
How far away is this day? Well, I want to end by presenting
a possible glimpse of the future. In
some recent work Shevaun Lewis, Dave Kush and Brad Larson (LKL) have used FGEs
to probe the syntactic derivation of constructions like (3) (see here for some
slides):
(3) What
and when will we eat
Not surprisingly, these coordinated WH questions have rather
elaborate syntactic properties. I will not detail them for you (read the slides), except to say that LKL are led to analyze these by
treating the two WHs rather differently. LKL propose that the inner WH lands in
clause initial position via movement while the outer one is base generated
there. The evidence for this conclusion is two-fold. First, there is an
acceptability contrast between (4a) and (4b), the latter being quite a bit
worse than the former (and yes they ran the relevant acceptability judgment
studies to show this). Second, the what
in (4a) fails to induce an FGE. If FGEs
are diagnostic of movement dependencies (as above) then the absence of these in
(4a) is just what we would expect, and apparently receive. To my knowledge this
is the first time a technique borrowed from psycho has been pressed into
service to support a novel syntactic conclusion. Terrific!
(4) a. What
and when will we eat something
b.
When and what will we eat something
It is the sign of a progressing research program that novel
questions and techniques keep springing up. The aim is to conciliate these
producing a bigger and bigger coherent picture.
To date, in my view, a great deal of what we have discovered about how
FL does syntax has come from the careful analysis of natural language grammars.
The LKL results are signaling a slightly different future: I have a
dream that is deeply rooted in the Generative enterprise. I have a dream that
one day linguistics will rise up and live the true meaning of its biolinguistic
roots. I have a dream that one day syntacticians and psycholinguists (and eventually neuroscientists) will use
each other’s work to strongly constrain their common research project of understanding how FL is structured and how it is used. I have a dream, and LKL
provides a glimpse of that wonderful and glorious future.
Subscribe to:
Posts (Atom)

