Two revisions below thx to Ewan spring a thinko, a slip of the mind.
I have recently urged that we adopt a particular understanding of the Strong Minimalist Thesis (SMT) (here). The version that I favor treats the SMT as a thesis about systems that use grammars and suggests that central features of the grammatical representations that they use will be crucial to explaining why they are efficient. If this proves to be doable, then it is reasonable to describe FL and the grammars it makes available as “well designed” and “computationally efficient.” Stealing from Bob Berwick (here), I will take parsing efficiency to mean real time parsing and (real time) acquisition to mean easy acquisition given the PLD. Put this all together and the SMT is the conjecture that the grammatical format of Gs and UG is critical to allowing parsers, acquirers, producers, etc. to be very good at what they do (i.e. to be well-designed). On this view, grammars are “well designed” or “computationally efficient” in virtue of having properties that allow their users to be good at what they do when such grammars are embedded transparently within these systems.
I have recently urged that we adopt a particular understanding of the Strong Minimalist Thesis (SMT) (here). The version that I favor treats the SMT as a thesis about systems that use grammars and suggests that central features of the grammatical representations that they use will be crucial to explaining why they are efficient. If this proves to be doable, then it is reasonable to describe FL and the grammars it makes available as “well designed” and “computationally efficient.” Stealing from Bob Berwick (here), I will take parsing efficiency to mean real time parsing and (real time) acquisition to mean easy acquisition given the PLD. Put this all together and the SMT is the conjecture that the grammatical format of Gs and UG is critical to allowing parsers, acquirers, producers, etc. to be very good at what they do (i.e. to be well-designed). On this view, grammars are “well designed” or “computationally efficient” in virtue of having properties that allow their users to be good at what they do when such grammars are embedded transparently within these systems.
One particularly attractive virtue of this interpretation
(for me) is that I understand how I could go about empirically investigating
it. I confess that this is not true for
other versions of the SMT that talk about neat fits between grammar and the CI
interface, for example. So far as I can tell, we know rather little about the
CI interface and so the question of fit is, at best, premature. On the other
hand we do know a bit about how parsing works and how acquisition proceeds so
we have something to fit the grammar to.[1]
So how to proceed? In two steps I believe. The first is to
see if use systems (e.g. parsers) actually deploy grammars in real time, i.e.
as they parse. Thus, if it is true that the basic features of grammatical
representations are responsible for how (e.g.) parsers manage to efficiently do
what they do then we should find real time evidence implicating these representations
in real time parsing. Second, we should look for how exactly the implicated
features manage to make things so efficient. Thus, we should look for
theoretical reasons for why parsers that transparently embody, say, Subjacency
like principles, would be efficient. Let
me discuss each of these points in turn.
There is increasing evidence from psycho-ling research indicating
that real time parsing respects grammatical distinctions, even very subtle
ones. Colin Phillips is a leader in this
kind of work and he and his (ex) students (e.g. Brian Dillon, Matt Wagers,
Ellen Lau, Dave Kush, Masaya Yoshida) have produced a body of work that
demonstrates how very closely parsers respect grammatical conditions like
islands, c-command, and local binding domains. And by closely I mean very closely. So, for example, Colin shows (here)
that online parsing respects the grammatical conditions that license parasitic
gaps. So, not only do parsers respect islands, but they even treat
configurations where island effects are amnestied as if they were not islands.
Thus, parsers respect both the general conditions that grammars lay down
regarding islands and the exceptions
to these general conditions that grammars allow. This is what I mean by
‘close.’
There is a recent excellent demonstration of this from
Masaya Yoshida, Lauren Ackerman, Morgan Purier and Rebekah Ward (YLPW) (here
are slides from a recent CUNY talk).[2]
YLPW analyzes the processing of backward sluicing constructions like (1):
(1) I
don’t recall which writer, but the editor notified a writer about a new project
There is an ellipsis “gap” right after which writer that is redeemed by anchoring it to a writer in the following sentence. What
YLPW is looking to determine is whether the elided
gap site is sensitive to online
parsing effects. YLPW uses a plausibility effect as probe as follows.
First, it is well known that a wh in CP triggers an active search for a verb/gap that will give it
an interpretation. ‘Active’ here means that the parser uses a top down
predictive process and is eagerly looking to link the wh to a predicate without first consulting bottom information that
would indicate the link to be ill-advised. YLPW show that the eagerness to
“fill a gap” is as true for implicit gaps within ellipsis sites as it is for
“real” gaps in regular wh sentences. YLPW shows this by demonstrating a
plausibility effect slowdown in
sentences like (2a) parallel to the ones found in (2b):
(2) a. I
don’t remember which writer/which book,
but the editor notified a writer about a new book
b.
I don’t remember which writer/which book
the editor notified GAP about a new book
When the wh is which book then there is a significant
pause at notified in both sentences
in (2), as contrasted with the same sentences where which writer is the antecedent of the gap. This is because parsers, we know, greedily
try and relate the wh to the first
syntactically available position encountered and in the case of which book the wh is not a plausible filler of the gap and the attempted filling
results in a little lingering about the verb (*notify this book about…). If the antecedent is which writer no such pause occurs, for obvious reasons. The plausibility effect, then, is just a
version of the well-known filled gap effect, with a little semantic kicker to
add some frisson. At any rate, the first important discovery is that we find
the same plausibility effect in both (2a) with the gap inside a sluiced
ellipsis site, and (2b) where the gap is “overt.”
The next step is to see if this plausibility/filled gap
effect slowdown occurs when the relevant antecedent for the sluiced ellipsis
site is inside an island. It is well known that ellipsis is not subject to island restrictions.
Thus, if the parser cleaves tightly
to the distinctions the grammar makes (as the SMT would lead us to expect) then
we should find plausibility slowdowns except
when the gap is inside an ellipsis site for the latter are not subject to
island restrictions [and so should induce filled gap/plausibility effects (added: thx Ewan)]. And that’s exactly
what YLPW find. Though plausibility effects are not found at notified in cases like (3) they are found in cases like (4) where the
“gap” is inside a sluice sight.
(3) I
don’t remember which book [the editor
who notified the publisher about
some science book] had recommended to me
(4) I
don’t remember which book, but [the
editor who notified the publisher
about some science book] recommended a new book to me
This is just what we expect from a parser that transparently
embeds a UG like grammar that treats movement but not ellipsis as a product of
(long) movement.
The conclusion: it seems that parsers make just the
distinctions that grammars make when they parse in real time, just as the SMT
would lead us to expect.
So, there is growing evidence that parsers transparently
embed UG like grammars. This readies us
for the second step. Why should they do so?
Here, there is less current research
that bears on the issue. However, there is work from the 80s by Mitch Marcus,
Bob Berwick and Amy Weinberg that showed that a Marcus style parser that
incorporated grammatical features like Subjacency (and, interestingly, Extension)
could parse sentences efficiently (effectively, in real time). This is just what the doctor ordered. It goes
without saying (though I will say it) that this work needs updating to bear
more directly on the SMT and minimalist accounts of FL. However, it provides a
useful paradigm of how one might go
about connecting the discoveries concerning online parsing with computational questions
of parsing efficiency and their relationship to central architectural features
of FL/UG.
The SMT is a bold conjecture. Indeed, it is likely false, at
least in fine detail. This does not, however, detract from its programmatic
utility. The fact is that there is
currently lots of research that can be understood as bearing on its accuracy
and that fruitfully brings together work in syntax, psycholinguistics and
computational linguistics. The SMT, in
other words, is a terrific hypothesis that will generate fascinating work
regardless of its ultimate empirical fate.
That’s what we want from a research program and that’s something that
the Strong Minimalist Thesis is ready to deliver. Were this all that the
Minimalist Program provided, it would have been enough (dayenu!). There is more, but for the nonce, this is more than
enough. Yay, for the Minimalist Program!!!
[1]
Let me modulate this: we know something about some other parts, see here
for discussion of magnitude estimation in the visual domain. Note that this
discussion fits well with the version of the SMT deployed here precisely
because we know something about how this part of the visual system works. We
cannot say as much about most of the other parts of CI. Indeed, we don’t really
know how many “parts” CI has.
[2]
They are running some more experiments, so this work is not yet finished.
Nonetheless, it illustrates the relevant point well, and it is really fun
stuff.