Comments

Showing posts with label Yoshida. Show all posts
Showing posts with label Yoshida. Show all posts

Thursday, April 17, 2014

The SMT again

Two revisions below thx to Ewan spring a thinko, a slip of the mind.

I have recently urged that we adopt a particular understanding of the Strong Minimalist Thesis (SMT) (here).  The version that I favor treats the SMT as a thesis about systems that use grammars and suggests that central features of the grammatical representations that they use will be crucial to explaining why they are efficient. If this proves to be doable, then it is reasonable to describe FL and the grammars it makes available as “well designed” and “computationally efficient.” Stealing from Bob Berwick (here), I will take parsing efficiency to mean real time parsing and (real time) acquisition to mean easy acquisition given the PLD.  Put this all together and the SMT is the conjecture that the grammatical format of Gs and UG is critical to allowing parsers, acquirers, producers, etc. to be very good at what they do (i.e. to be well-designed). On this view, grammars are “well designed” or “computationally efficient” in virtue of having properties that allow their users to be good at what they do when such grammars are embedded transparently within these systems.

One particularly attractive virtue of this interpretation (for me) is that I understand how I could go about empirically investigating it.  I confess that this is not true for other versions of the SMT that talk about neat fits between grammar and the CI interface, for example. So far as I can tell, we know rather little about the CI interface and so the question of fit is, at best, premature. On the other hand we do know a bit about how parsing works and how acquisition proceeds so we have something to fit the grammar to.[1]

So how to proceed? In two steps I believe. The first is to see if use systems (e.g. parsers) actually deploy grammars in real time, i.e. as they parse. Thus, if it is true that the basic features of grammatical representations are responsible for how (e.g.) parsers manage to efficiently do what they do then we should find real time evidence implicating these representations in real time parsing. Second, we should look for how exactly the implicated features manage to make things so efficient. Thus, we should look for theoretical reasons for why parsers that transparently embody, say, Subjacency like principles, would be efficient.  Let me discuss each of these points in turn.


There is increasing evidence from psycho-ling research indicating that real time parsing respects grammatical distinctions, even very subtle ones.  Colin Phillips is a leader in this kind of work and he and his (ex) students (e.g. Brian Dillon, Matt Wagers, Ellen Lau, Dave Kush, Masaya Yoshida) have produced a body of work that demonstrates how very closely parsers respect grammatical conditions like islands, c-command, and local binding domains. And by closely I mean very closely.  So, for example, Colin shows (here) that online parsing respects the grammatical conditions that license parasitic gaps. So, not only do parsers respect islands, but they even treat configurations where island effects are amnestied as if they were not islands. Thus, parsers respect both the general conditions that grammars lay down regarding islands and the exceptions to these general conditions that grammars allow. This is what I mean by ‘close.’

There is a recent excellent demonstration of this from Masaya Yoshida, Lauren Ackerman, Morgan Purier and Rebekah Ward (YLPW) (here are slides from a recent CUNY talk).[2] YLPW analyzes the processing of backward sluicing constructions like (1):

(1)  I don’t recall which writer, but the editor notified a writer about a new project

There is an ellipsis “gap” right after which writer that is redeemed by anchoring it to a writer in the following sentence. What YLPW is looking to determine is whether the elided gap site is sensitive to online parsing effects. YLPW uses a plausibility effect as probe as follows.

First, it is well known that a wh in CP triggers an active search for a verb/gap that will give it an interpretation. ‘Active’ here means that the parser uses a top down predictive process and is eagerly looking to link the wh to a predicate without first consulting bottom information that would indicate the link to be ill-advised. YLPW show that the eagerness to “fill a gap” is as true for implicit gaps within ellipsis sites as it is for “real” gaps in regular wh sentences.  YLPW shows this by demonstrating a plausibility effect slowdown in sentences like (2a) parallel to the ones found in (2b):

(2)  a. I don’t remember which writer/which book, but the editor notified a writer about a new book
b. I don’t remember which writer/which book the editor notified GAP about a new book

When the wh is which book then there is a significant pause at notified in both sentences in (2), as contrasted with the same sentences where which writer is the antecedent of the gap.  This is because parsers, we know, greedily try and relate the wh to the first syntactically available position encountered and in the case of which book the wh is not a plausible filler of the gap and the attempted filling results in a little lingering about the verb (*notify this book about…). If the antecedent is which writer no such pause occurs, for obvious reasons.  The plausibility effect, then, is just a version of the well-known filled gap effect, with a little semantic kicker to add some frisson. At any rate, the first important discovery is that we find the same plausibility effect in both (2a) with the gap inside a sluiced ellipsis site, and (2b) where the gap is “overt.”

The next step is to see if this plausibility/filled gap effect slowdown occurs when the relevant antecedent for the sluiced ellipsis site is inside an island. It is well known that ellipsis is not subject to island restrictions. Thus, if the parser cleaves tightly to the distinctions the grammar makes (as the SMT would lead us to expect) then we should find plausibility slowdowns except when the gap is inside an ellipsis site for the latter are not subject to island restrictions [and so should induce filled gap/plausibility effects (added: thx Ewan)].  And that’s exactly what YLPW find. Though plausibility effects are not found at notified in cases like (3) they are found in cases like (4) where the “gap” is inside a sluice sight.

(3)  I don’t remember which book [the editor who notified the publisher about some science book] had recommended to me
(4)  I don’t remember which book, but [the editor who notified the publisher about some science book] recommended a new book to me

This is just what we expect from a parser that transparently embeds a UG like grammar that treats movement but not ellipsis as a product of (long) movement.

The conclusion: it seems that parsers make just the distinctions that grammars make when they parse in real time, just as the SMT would lead us to expect.

So, there is growing evidence that parsers transparently embed UG like grammars.  This readies us for the second step. Why should they do so?  Here, there is less current research that bears on the issue. However, there is work from the 80s by Mitch Marcus, Bob Berwick and Amy Weinberg that showed that a Marcus style parser that incorporated grammatical features like Subjacency (and, interestingly, Extension) could parse sentences efficiently (effectively, in real time).  This is just what the doctor ordered. It goes without saying (though I will say it) that this work needs updating to bear more directly on the SMT and minimalist accounts of FL. However, it provides a useful paradigm of how one might go about connecting the discoveries concerning online parsing with computational questions of parsing efficiency and their relationship to central architectural features of FL/UG.

The SMT is a bold conjecture. Indeed, it is likely false, at least in fine detail. This does not, however, detract from its programmatic utility.  The fact is that there is currently lots of research that can be understood as bearing on its accuracy and that fruitfully brings together work in syntax, psycholinguistics and computational linguistics.  The SMT, in other words, is a terrific hypothesis that will generate fascinating work regardless of its ultimate empirical fate.  That’s what we want from a research program and that’s something that the Strong Minimalist Thesis is ready to deliver. Were this all that the Minimalist Program provided, it would have been enough (dayenu!). There is more, but for the nonce, this is more than enough. Yay, for the Minimalist Program!!!




[1] Let me modulate this: we know something about some other parts, see here for discussion of magnitude estimation in the visual domain. Note that this discussion fits well with the version of the SMT deployed here precisely because we know something about how this part of the visual system works. We cannot say as much about most of the other parts of CI. Indeed, we don’t really know how many “parts” CI has.
[2] They are running some more experiments, so this work is not yet finished. Nonetheless, it illustrates the relevant point well, and it is really fun stuff.

Wednesday, December 4, 2013

Dreams of a unified theory; a great big juicy problem (Yay!!)

The intricacies of A’-syntax is one of the glories of GB.[1]  The unification of Ross’s islands in terms of subjacency and the discovery of ECP dependencies (especially the adjunct/argument distinction) coupled with wide ranging investigations of these effects in a large variety of different kinds of languages marked a high point in Generative Grammar. This all changed with the Minimalist (M) “Revolution” (yes; these are scare quotes). Thereafter, Island and ECP effects mostly fell from the hot topics list (compare post M work with that done in the 80s and early 90s where it seemed that every other paper/book was about A’-dependencies and their island/ECP restrictions). Moreover, though early M was chock full of discussions of Superiority, an A’-effect, it was mainly theoretically interesting for the light that it threw on Minimality and Shortest Move/Attract rather than how it bore on Islands or the ECP. Indeed, from where I sit, the bulk of the interesting work within M has been on A rather than A’ dependencies.[2]

Moreover, whereas there has been interesting research aiming to unify various grammatical modules, subjacency and ECP have resisted theoretical integration, at least interesting versions thereof. It is possible, indeed easy, to translate bounding theory or barriers into phase terminology.[3] However, there is nothing particularly insightful gained in doing this. It is also possible to unify Islands with Minimality given the right use of features placed in appropriate edge positions, but IMO little has been gained to date in so proceeding. So Island and ECP effects, once the pride of theoretical syntax have become a backwater and a slightly embarrassing one for three related reasons.

First, though it is pretty easy to translate Subjacency (viz. bounding theory) in phase terms, this translation simply duplicates the peccadillos of the earlier approaches (e.g. we stipulated bounding nodes, we now stipulate (strong) phases, we stipulated escape hatches (C yes, D no) we now stipulate phase edges (both which phases have any to use and how many they have)).

Second, ad hoc as this is, it’s good compared to the problems the ECP throws up. For example, the ECP is conceptually a trace licensing requirement. Where does this leave us when we replace traces with copies as M does? Do copies need licensing? Why if they are simply different occurrences of a single expression? Moreover, how do we code the difference between adjuncts versus arguments?  What makes the former so restricted when compared to the latter?

Last, the obvious redundancy between Subjacency and the ECP raises serious M questions. Both involve the same island like configurations yet they are entirely different licensing conditions. Talk of redundancy! One of Subjacency or the ECP is bad enough, but both? Argh!!

So, A’-syntax raises M issues and a natural hope is to dispose of these problems by placing them in someone else’s trash bin. And there have been several attempts, to do just this, e.g. Kluender & Kutas, Sag & Hoffmeister, Hawkins, among others. The idea has been to treat island effects as a reflection of processing complexity, the latter arising when parsers try to relate elements outside an island (fillers) to positions (gaps) within an island.  It is well known that filler/gap dependencies impose a memory/storage cost as the process of relating a filler to a gap requires keeping the filler “live” until it’s discharged in the appropriate position. Interestingly, there is independent psycho-ling evidence that the cost of keeping elements active can depend on the details of the parse quite independently of whether islands are involved (e.g. beginnings of finite clauses induce load, as does the parsing of definites).[4] Island effects, on this view, are just the sum total of these island-independent processing costs. In effect, Islands are just structures where these other costly independently manifested requirements converge. If true, this idea could, with some work, let M off the island hook.[5] Wouldn’t that be nice?

It would be, but I personally doubt that this strategy will work out.  The main problem is that it seems very hard to explain the unacceptability profiles of island effects in processing terms. A recent volume (of which I am co-editor though Jon Sprouse did all the really heavy lifting and deserves all the credit, Experimental Syntax and Island Effects) reviews the basic issues. The main take home message is that when considered in detail, the relevant cited complexity inducers (e.g. definiteness) do not eliminate the structural contributions of islands to the perceived acceptability, though they can modulate it (viz. the super-additive effects of islands remain even if the severity of the unacceptability can be manipulated). Many of the papers in the volume address these issues in detail (see especially those by Jon Sprouse, Matt Wagers, and Colin Phillips). The book also contains good representatives of the processing “complexity” alternative and the interested reader is encouraged to take a look at the papers (WARNING: being a co-editor forbids me in good conscience, from advocating purchase but I believe that many would consider this book a perfect holiday gift even for those with no interest in the relevant intellectual issues, e.g. it’s really heavy and would make a perfect paperweight or door stopper).

A nice companion piece to the papers in the above volume that I have recently read seconds the conclusion that Island Effects have a structural source.  The paper (here) is by Yoshida, Kazanina, Pablos and Sturt (YKPS) and it explores the problem in a very clever way. Here’s a quick review.

YKPS starts from the assumption that if the problem is one of the processing complexities of islands, then any dependency into an island that is computed online (as filler/gap dependencies are) should show island like properties even if these dependencies are not products of movement. They identify forward cataphora (e.g. his1 managers revealed that [island the studio that notified Jeffrey Stewart1 about the new film] selected a novel for the script) as one such dependency. YKPS shows that the indicated referential dependency is calculated online just as filler/gap dependencies are (both are very greedy in fixing the dependency). However, in contrast to movement dependencies, pronoun resolution in forward cataphora does not exhibit island effects. The argument is easy to follow and the conclusion strikes me as pretty solid, but read it and judge for yourself. What I liked about it is that it is a classic example of a typical linguistic argument form: YKPS identifies a dog that doesn’t bark. If parsing complexity is the relevant variable then it needs to explain both why some dependencies exhibit island effects and, just as importantly, why some do not. In other words, negative data counts! The absence of island effects is as much a datum as its presence is, though it is often ignored.[6] As YKPS put it:

Complexity accounts, which attribute island effects to the effect of processing complexity of the online dependency formation process, need to explain why the same complexity does not affect (my emphasis, NH) the formation of cataphoric dependencies. (17)

So, it seems to me that islands are here to stay, even if their presence in UG embarrasses minimalists.

Three points and I end. First, the argument that YKPS presents is another nice example of how psycho-techniques can be used to advance syntactic ends.  How so? Well, it is critical to YKPS’s point that forward cataphora involves the same kind of processing strategies (active filler) as do regular filler/gap dependencies that one finds in movement despite the dependencies being entirely different grammatically. This is what makes it possible to compare the two kinds of processes and conclude from their different behavior wrt islands that structural effects cannot be reduced to parsing complexity (a prima facie very reasonable hypothesis and one that might even be nice were it true!).[7] 

Second, the complexity theory of islands pertains to Subjacency Effects. The far harder problem, as I mentioned earlier, involve ECP effects. Indeed, were Subjacency Effects reduced to complexity effects, the presence of ECP effects in the very same configurations would become even more puzzling, at least to me. At any rate, both problems remain, and await a decent M analysis.

Third, let me end with some personal intellectual history. I taught a course on the old GB A’ material with Howard Lasnik this semester (a great experience, thx Howard) and have become pretty convinced that finding a way to simply recapitulate ECP and Island effects in M terms is by no means trivial.  To see this, I invite you to simply try to translate the GB theories into an M acceptable idiom. Even this is pretty hard to do, and a simple translation still leaves one short of a M acceptable account. Conclusion? This is still a really juicy research topic for the unificationally inclined, i.e. a great Minimalist research topic.



[1] I take GB to be the logical culmination of work that first developed as the Extended Standard Theory. Moreover, I here, again, take GB to be one of several kissing cousins, such as GPSG, LFG, HPSG.
[2] This is a bird’s eye evaluation and there are notable exceptions to this coarse generalization. Here is one very conspicuous exception: how ellipsis obviates island effects. Lasnik and Merchant have turned this into a small very productive industry. The main theoretical effect has been to make us reconsider what makes an island islandy. The ellipsis effects have revived an interpretation that has some roots in Ross, that it is not the illicit dependency that matters but the phonological realization thereof that counts.  Islands, on this view, are PF rather than syntactic effects. At any rate, this is really interesting stuff which has led us to understand Island Effects in new ways.
[3] At least if one allows D to be a phase, something some (e.g. Chomsky) has only grudgingly accepted.
[4] Rick Lewis has some nice models of this based on empirical work by Gibson.
[5] Of course, more work needs doing. For example, one needs to explain, why, ellipsis obviates these processing effects (see note 2).
[6] Note 4 indicates another bit of negative data that needs explanation on the complexity account. One might think, for example, that having to infer structure would add to complexity and thus increase the unacceptability of island violations, contrary to what we in fact find.
[7] Very reasonable indeed as witnessed by Chomsky’s extensive efforts to argue against the supposition that island effects are simple complexity effects in On Wh Movement.