Comments

Showing posts with label Strong Minimalist Thesis. Show all posts
Showing posts with label Strong Minimalist Thesis. Show all posts

Monday, July 21, 2014

Comments on lecture 4-I

I have just finished listening to Chomsky’s fourth lecture and so this will be the last series of posts on them (here). If you have not seen them, let me again suggest that you take the time to watch. They are very good and well worth the (not inconsiderable) time commitment.

In 4, Chomsky’s does three things. First he again tries to sell the style of investigation that the lectures as a whole illustrate. Second, he reviews the motivations and basic results of his way of approaching the Darwin’s Problem.  Third, he proposes ways of tidying up some of the loose ends that the outline in 3 generates (at least they were loose ends that I did not understand).  Let me review each of these points in turn.

1. The central issues and the Strong Minimalist Thesis (SMT)

Chomsky, as is his wont, returns to the key issues as he sees them. There are two of particular importance.

First, he believes that we should be looking for simple theories.  He names this dictum Galileo’s Maxim (GM).  GM asserts (i) that nature is simple and (ii) it is the task of the scientist to prove that it is. Chomsky notes that this is not merely good general methodological advice (which it is), but that in the particular context of the study of FL there are substantive domain specific reasons for adopting it. Namely: Darwin’s Problem (DP). Chomsky claims that DP rests on three observations: (i) That our linguistic competence is not learnable from simple data, (ii) There is no analogue of our linguistic capacity anywhere else in the natural world, and (iii) The capacity for language emerged recently (in the last 100k years or so), emerged suddenly and has remained stable in its properties since its emergence.[1]  These three points together imply that we have a non-trival FL, that it is species specific and that it arose as a result of a very “simple” addition to ancestor’s cognitive repertoire.  So, in addition to the general (i.e. external to the specific practice of linguistics) methodological virtues of looking for simple and elegant theories, DP provides a more substantive (i.e. internal to linguistics) incentive, as simple theories are just the sorts of things that could emerge rapidly in a lineage and remain stable after emerging. 

I very much like this way of framing the central aims of the Minimalist Program (MP).  It reconciles two apparently contradictory themes that have motivated MP. The first theme is that looking for simple theories is just good methodology and so MP is nothing new.  On this reading, MP is just the rational extension of GG theorizing, just the application of general scientific principles/standards of rational inquiry to linguistic investigations.  On this view, MP concerns are nothing new and the standards MP applies to theory evaluation are just the same as they always were.  The second view, one that also seems to be a common theme, is that MP does add a new dimension to inquiry. DP, though always a concern, is now ripe for investigation. And thinking about DP motivates developing simple theories for substantive reasons internal to linguistic investigations, motivations in addition to the standard ones prompted by concerns of scientific hygiene.  On this view, raising DP to prominence changes the relevant standards for theoretical evaluation. Adding DP to Plato’s Problem, then, changes the nature of the problem to be addressed in interesting ways.

This combined view, I think, gets MP right.  It is both novel and old hat.  What Chomsky notes is that at some times, depending on how developed theory is, new questions can emerge or become accented and at those times the virtues of simplicity have a bite that goes beyond general methodological concerns.  Another way of saying this, perhaps, is that there are times (now being one in linguistics) where the value of theoretical simplicity is elevated and the task of finding simple non-trivial coherent theories is the central research project. The SMT is intended to respond to this way of viewing the current project (I comment on this below).

Chomsky makes a second very important point. He notes that our explanatory target should be the kinds of effects that GG has discovered over the last 60 years.  Thus, we should try to develop accounts as to why FL generates an unbounded number of structured linguistic objects (SLO), why it incorporates displacement operations, why it obeys locality restrictions (strict cyclicity, PIC), why there is overt morphology, why there are subject/object asymmetries (Fixed Subject Effects/ECP), why there are EPP effects, etc. So, Chomsky identifies both a method of inquiry  (viz. Galileo’s Maxim) and a target of inquiry (viz. the discovered laws and effects of GG). Theory should aim to explain the second while taking DM very very seriously.

The SMT, as Chomsky sees it, is an example of how to do this (actually, I don’t think he believes it is an example, but the only possible conceptually coherent way to proceed).  Here’s the guts of the SMT: look for the conceptually simplest computational procedures that generate SLOs and that are interpreted at CI and (secondarily) SM.  Embed these conceptually simple operations in a computationally efficient system (one that adheres to obvious and generic principles of efficient computation like minimal search, No Tampering, Inclusiveness, Memory load reduction) and show that from these optimal starting points one can derive a good chunk of the properties that GG has discovered natural language grammars to have.  And, when confronted with apparent counter-examples to the SMT, look harder for a solution that redeems the SMT.  This, Chomsky argues is the right way, today, to do theoretical syntax.

I like almost all of this, as you might have guessed. IMO, the only caveat I would have is that the conceptually simple is often a very hard to discern. Moreover, what Occam might endorse, DP might not. I have discussed before that what’s simple in a DP context might well depend on what was cognitively available to our ancestors prior to the emergence of FL. Thus, there may be many plausible simple starting points that lead to different kinds of theories of FL all of which respond to Chomsky’s methodological and substantive vision of MP. For what it’s worth, contra Chomsky, I think (or, at least believe that it is rational to suggest) that Merge is not simple but complex and that it is composed of a more cognitively primitive operation (viz. Iteration) and a novel part (viz. Labeling). For those who care about this, I discuss what I have in mind further here in part 4 (the finale) of my comments to lecture 3.[2] However, that said, I could not agree with Chomsky’s general approach more. An MP that respects DP should deify GM and target the laws of GG.  Right on.



[1] Chomsky has a nice riff where he notes that though it seems to him (and to any sane researcher) that (i)-(iii) are obviously correct, nonetheless these are highly controversial claims, if judged by the bulk of research on language. He particularly zeros in on big data statistical learning types and observes (correctly in my view) that not only have they not been able to deliver on even the simplest PoS problems (e.g. structure dependence in Y/N questions) but that they are currently incapable of delivering anything of interest given that they have misconstrued the problem to be solved. Chomsky develops this theme further, pointing out that to date, in his opinion, we have learned nothing of interest from these pursuits either in syntax or semantics. I completely agree and have said so here. Still, I get great pleasure in hearing Chomsky’s completely accurate dismissive comments. 
[2] I also discuss this in a chapter co-written with Bill Idsardi forthcoming in a collection edited by Peter Kosta from Benjamins.

Thursday, April 17, 2014

The SMT again

Two revisions below thx to Ewan spring a thinko, a slip of the mind.

I have recently urged that we adopt a particular understanding of the Strong Minimalist Thesis (SMT) (here).  The version that I favor treats the SMT as a thesis about systems that use grammars and suggests that central features of the grammatical representations that they use will be crucial to explaining why they are efficient. If this proves to be doable, then it is reasonable to describe FL and the grammars it makes available as “well designed” and “computationally efficient.” Stealing from Bob Berwick (here), I will take parsing efficiency to mean real time parsing and (real time) acquisition to mean easy acquisition given the PLD.  Put this all together and the SMT is the conjecture that the grammatical format of Gs and UG is critical to allowing parsers, acquirers, producers, etc. to be very good at what they do (i.e. to be well-designed). On this view, grammars are “well designed” or “computationally efficient” in virtue of having properties that allow their users to be good at what they do when such grammars are embedded transparently within these systems.

One particularly attractive virtue of this interpretation (for me) is that I understand how I could go about empirically investigating it.  I confess that this is not true for other versions of the SMT that talk about neat fits between grammar and the CI interface, for example. So far as I can tell, we know rather little about the CI interface and so the question of fit is, at best, premature. On the other hand we do know a bit about how parsing works and how acquisition proceeds so we have something to fit the grammar to.[1]

So how to proceed? In two steps I believe. The first is to see if use systems (e.g. parsers) actually deploy grammars in real time, i.e. as they parse. Thus, if it is true that the basic features of grammatical representations are responsible for how (e.g.) parsers manage to efficiently do what they do then we should find real time evidence implicating these representations in real time parsing. Second, we should look for how exactly the implicated features manage to make things so efficient. Thus, we should look for theoretical reasons for why parsers that transparently embody, say, Subjacency like principles, would be efficient.  Let me discuss each of these points in turn.


There is increasing evidence from psycho-ling research indicating that real time parsing respects grammatical distinctions, even very subtle ones.  Colin Phillips is a leader in this kind of work and he and his (ex) students (e.g. Brian Dillon, Matt Wagers, Ellen Lau, Dave Kush, Masaya Yoshida) have produced a body of work that demonstrates how very closely parsers respect grammatical conditions like islands, c-command, and local binding domains. And by closely I mean very closely.  So, for example, Colin shows (here) that online parsing respects the grammatical conditions that license parasitic gaps. So, not only do parsers respect islands, but they even treat configurations where island effects are amnestied as if they were not islands. Thus, parsers respect both the general conditions that grammars lay down regarding islands and the exceptions to these general conditions that grammars allow. This is what I mean by ‘close.’

There is a recent excellent demonstration of this from Masaya Yoshida, Lauren Ackerman, Morgan Purier and Rebekah Ward (YLPW) (here are slides from a recent CUNY talk).[2] YLPW analyzes the processing of backward sluicing constructions like (1):

(1)  I don’t recall which writer, but the editor notified a writer about a new project

There is an ellipsis “gap” right after which writer that is redeemed by anchoring it to a writer in the following sentence. What YLPW is looking to determine is whether the elided gap site is sensitive to online parsing effects. YLPW uses a plausibility effect as probe as follows.

First, it is well known that a wh in CP triggers an active search for a verb/gap that will give it an interpretation. ‘Active’ here means that the parser uses a top down predictive process and is eagerly looking to link the wh to a predicate without first consulting bottom information that would indicate the link to be ill-advised. YLPW show that the eagerness to “fill a gap” is as true for implicit gaps within ellipsis sites as it is for “real” gaps in regular wh sentences.  YLPW shows this by demonstrating a plausibility effect slowdown in sentences like (2a) parallel to the ones found in (2b):

(2)  a. I don’t remember which writer/which book, but the editor notified a writer about a new book
b. I don’t remember which writer/which book the editor notified GAP about a new book

When the wh is which book then there is a significant pause at notified in both sentences in (2), as contrasted with the same sentences where which writer is the antecedent of the gap.  This is because parsers, we know, greedily try and relate the wh to the first syntactically available position encountered and in the case of which book the wh is not a plausible filler of the gap and the attempted filling results in a little lingering about the verb (*notify this book about…). If the antecedent is which writer no such pause occurs, for obvious reasons.  The plausibility effect, then, is just a version of the well-known filled gap effect, with a little semantic kicker to add some frisson. At any rate, the first important discovery is that we find the same plausibility effect in both (2a) with the gap inside a sluiced ellipsis site, and (2b) where the gap is “overt.”

The next step is to see if this plausibility/filled gap effect slowdown occurs when the relevant antecedent for the sluiced ellipsis site is inside an island. It is well known that ellipsis is not subject to island restrictions. Thus, if the parser cleaves tightly to the distinctions the grammar makes (as the SMT would lead us to expect) then we should find plausibility slowdowns except when the gap is inside an ellipsis site for the latter are not subject to island restrictions [and so should induce filled gap/plausibility effects (added: thx Ewan)].  And that’s exactly what YLPW find. Though plausibility effects are not found at notified in cases like (3) they are found in cases like (4) where the “gap” is inside a sluice sight.

(3)  I don’t remember which book [the editor who notified the publisher about some science book] had recommended to me
(4)  I don’t remember which book, but [the editor who notified the publisher about some science book] recommended a new book to me

This is just what we expect from a parser that transparently embeds a UG like grammar that treats movement but not ellipsis as a product of (long) movement.

The conclusion: it seems that parsers make just the distinctions that grammars make when they parse in real time, just as the SMT would lead us to expect.

So, there is growing evidence that parsers transparently embed UG like grammars.  This readies us for the second step. Why should they do so?  Here, there is less current research that bears on the issue. However, there is work from the 80s by Mitch Marcus, Bob Berwick and Amy Weinberg that showed that a Marcus style parser that incorporated grammatical features like Subjacency (and, interestingly, Extension) could parse sentences efficiently (effectively, in real time).  This is just what the doctor ordered. It goes without saying (though I will say it) that this work needs updating to bear more directly on the SMT and minimalist accounts of FL. However, it provides a useful paradigm of how one might go about connecting the discoveries concerning online parsing with computational questions of parsing efficiency and their relationship to central architectural features of FL/UG.

The SMT is a bold conjecture. Indeed, it is likely false, at least in fine detail. This does not, however, detract from its programmatic utility.  The fact is that there is currently lots of research that can be understood as bearing on its accuracy and that fruitfully brings together work in syntax, psycholinguistics and computational linguistics.  The SMT, in other words, is a terrific hypothesis that will generate fascinating work regardless of its ultimate empirical fate.  That’s what we want from a research program and that’s something that the Strong Minimalist Thesis is ready to deliver. Were this all that the Minimalist Program provided, it would have been enough (dayenu!). There is more, but for the nonce, this is more than enough. Yay, for the Minimalist Program!!!




[1] Let me modulate this: we know something about some other parts, see here for discussion of magnitude estimation in the visual domain. Note that this discussion fits well with the version of the SMT deployed here precisely because we know something about how this part of the visual system works. We cannot say as much about most of the other parts of CI. Indeed, we don’t really know how many “parts” CI has.
[2] They are running some more experiments, so this work is not yet finished. Nonetheless, it illustrates the relevant point well, and it is really fun stuff.

Thursday, April 3, 2014

Chomsky's two hunches

I have been thinking again about the relationship between Plato’s Problem and Darwin’s. The crux of the issue, as I’ve noted before (see e.g. here) is the tension between the two. Having a rich linguo-centric FL makes explaining the acquisition certain features of particular Gs easy (why? Because they don’t have to be learned, they are given/innate). Examples include the stubborn propensity for movement rules to obey island conditions, for reflexives to resist non-local binding etc. However, having an FL with rich language specific architecture makes it more difficult to explain how FL came to be biologically fixed in humans. The problem gets harder still if one buys the claim that human linguistic facility arose in the species in (roughly) only the last 50-100,000 years. If this is true, then the architecture of FL must be more or less continuous with that we find in other domains of cognition, with the addition of a possible tweak or two (language is more or less an app in Jan Koster’sense). In other words, FL can’t be that linguo-centric! This is the essential tension. The principle project of contemporary linguistics (in particular that of the Minimalist Program (MP)), I believe, should be to resolve this tension.  In other words, to show how you can eat your Platonic cake and have Darwin’s too.

How to do this? Well, here’s an unkosher way of resolving the tension. It is not an admissible move in this game to deny Plato’s Problem is a real problem. That does not “resolve” the tension. It denies that there is/was one to begin with. Denying Plato’s Problem in our current setting includes ignoring all the POS arguments that have been deployed to argue in favor of linguo-centric structure for FL. Examples abound and I have been talking about these again in recent posts (here, here). Indeed, most of the structure GB postulates, if an accurate description of FL, is innate or stems from innate mental architecture.  GB’s cousins (H/GPSG, LFG, RG) have their corresponding versions of the GB modules and hence their corresponding linguo-centric innate structures. The interesting MP question is how to combine the fact that FL has the properties GB describes with a plausible story of how these GBish features of FL could have arisen. To repeat: denying that Plato’s Problem is real or denying that FL arose in the species at some time in the relatively recent past does not solve the MP problem, it denies that there is any problem to solve.[1]

There is one (and so far as I can tell, only one) way of squaring this apparent circle: to derive the properties of GB from simpler assumptions.  In other words, to treat GB in roughly the way the theory of Subjacency treats islands: to show that the independent principles and modules are all special cases of a simpler more plausible unified theory. 

This project involves two separate steps.

First, we need to show how to unify the disparate modules. A good chunk of my research over the last 15 years has aimed for this (with varying degrees of success). I have argued (though I have persuaded few) that we should try and reduce all non-local dependencies to “movement” relations. Combine this with Chomsky’s proposal that movement and phrase building devolve to the same operation ((E/I)-Merge) and one gets the result that all grammatical dependencies are products of a single operation, viz. Merge.[2] Or to put this now in Chomsky’s terms, once Merge becomes cognitively available (Merge being the evolutionary miracle, aka, random mutation), the rest of GB does as well for GB is nothing other than a catalogue of the various kinds of Merge dependencies available in a computationally well-behaved system.  

Second, we need to show that once Merge arises, the limitations on the Merge dependencies that GB catalogues (island effects, binding effects, control effects, etc.) follow from general (maybe ‘generic’ is a better term) principles of cognitive computation. If we can assimilate locality principles like the PIC and Minimality and Binding Domain to (plausibly) more cognitively generic principles like Extension (conservativity) or Inclusiveness then it is possible to understand that GB dependencies are what one gets if (i) all operations “live on” Merge and (ii) these operations are subject to non-linguocentric principles of cognitive computation. 

Note that if this can be accomplished, then the tension noted at the outset is resolved. Chomsky’s hunch, the basic minimalist conjecture, is that this is doable; that it is possible to reduce grammatical dependencies to a (at most) one (or two) specifically linguistic operations which when combined with other cognitive operations plus generic constraints on cognitive computation (one’s not particularly linguo-centric) we get the effects of GB.

There is a second additional conjecture that Chomsky advances that bears on the program. This second independent hunch is the Strong Minimalist Thesis (SMT). IMO, it has not been very clear how we are to understand the SMT. The slogan is that FL is the “optimal solution to minimal design specifications.”  However, I have never found the intent of this slogan to be particularly clear. Lately, I have proposed (e.g. here) that we understand the SMT in the context of the question one of the four classic questions in Generative Grammar: How are Gs put to use? In particular, the SMT tells us that grammars are well designed for use by the interface.

I want to stress that SMT is an extra hunch about the structure of FL. Moreover, I believe that this reconstruction of the problematic (thanks Hegel) might not (most likely, does not) coincide with how Chomsky understands MP. The paragraphs above argue that reconciling Darwin and Plato requires showing that most of the principles operative in FL are cognitively generic (viz. that they are operative in other non-linguistic cognitive domains). This licenses the assumption that they pre-exist the emergence of FL and so we need not explain why FL recruits them. All that is required is that they “be there” for the taking. The conjecture that FL is optimal computationally (i.e. that it is well-designed wrt to use by the interfaces) goes beyond the evolutionary assumption required to solve the Plato/Darwin tension. The SMT postulates that these evolutionarily available principles are also well designed. This second conjecture, if true, is very interesting precisely because the first Darwinian one can be true without the second optimal design assumption being true. Moreover, if the SMT is true, this might require explanation. In particular, why should evolutionary available mechanisms that FL embodies be well designed for use (especially given that FL is of recent vintage)?[3]

That said, what’s “well designed” mean? Well, here’s a proposal: that the competence constraints that linguists find suffice for efficient parsing and easy learnability. There is actually a lost literature on this conjecture that precedes MP. For example, the work by Marcus and Berwick & Weinberg on parsing, and Culicover & Wexler and Berwick on learnability investigate how the constraints on linguistic representations, when transparently embedded in use systems, can allow for efficient parsing and easy learnability.[4]  It is natural to say that grammatical principles that allow for efficient parsing and easy learning are themselves computationally optimal in a biologically/psychologically relevant sense. The SMT can be (and IMO, should be) understood as conjecturing that FL produces grammars that are computationally optimal in this sense.

Two thoughts to end:

First, this way of conceiving of MP treats it as a very conservative extension of the general generative program. One of the misconceptions “out there” (CSers and Psychologists are particularly prone to this meme)  is that Generativists change their minds and theories every 2 months and that this theoretical Brownian motion is an indication that linguists know squat about FL or UG. This is false. The outlines of MP as necessarily incorporating GB results (with the aim of making them “theorems” in a more general theoretical framework) emphasizes that MP does not abandon GB results but tries to explain them. This what typically takes place in advancing sciences and it is no different in linguistics. Indeed, a good Whig history of Generative Grammar would demonstrate that this conservatism has been characteristic of most of the results from LSLT to MP. This is not the place to show this, but I am planning to demonstrate it anon.

Second, MP rests on two different but related Chomskyan hunches (‘conjectures’ would sound more serious, so I suggest you sue this term when talking to the sciency types on the prestigious parts of campus): first that it is possible to resolve the tension between Plato and Darwin without doing damage to the former and that the results will be embeddable in use systems that are computationally efficient.  We currently have schematic outlines for how this might be done (though there are many holes to be filled). Chomsky’s hunch is that this project can be completed.

IMO, we have made some progress towards showing that this is not a vain hope, in fact that things are better than one might have initially thought (especially if one is a pessimist like me).[5] However, realizing this ambitious program requires a conservative attitude towards past results. In particular, MP does not imply that GB is passe. Going beyond explanatory adequacy does not imply forgetting about explanatory adequacy. Only cheap minimalism forgets what we have found, and as my mother repeatedly wisely warned me “cheap is expensive in the long run.” So, a bit of advice: think babies and bathwaters next time you are tempted to dump earlier GB results for purportedly minimalist ends.



[1] It is important to note that this is logically possible. Maybe the MP project rests on a misdescription of the conceptual lay of the land. As you might imagine, I doubt that this is so. However, it is a logical possibility. This is why POS phenomena are so critical to the MP enterprise. One cannot go beyond explanatory adequacy without some candidate theories that (purport to) have it.
[2] For the record, I am not yet convinced of Chomsky’s way of unifying things via Merge. However, for current purposes, the disagreement is not worth pursuing.
[3] Let me reiterate that I am not interpreting Chomsky here. I am pretty sure that he would not endorse this reconstruction of the Minimalist Problematic. Minimalists be warned!
[4] In his book on learning, Berwick notes that it is a truism in AI that “having the right restrictions on a given representation can make learning simple.” Ditto for parsing. Note that this does not imply that features of use cause features of representations, i.e. this does not imply that demands for efficient parsability cause grammars to have subjacency like locality constraints. Rather, for example, grammars that have subjacency like constraints will allow for simple transparent embeddings into parsers that will compute efficiently and support learning algorithms that have properties that support “easy” learning (See Berwick’s book for lots of details).
[5] Actually, if pressed, I would say that we have made remarkable progress in cashing in Chomsky’s two bets. We have managed to outline plausible theories of FL that unify large chunks of the GB modules and we have begun to find concrete evidence that both parsing, production and language acquisition transparently use the kinds of representations that competence theories have discovered. The project is hardly complete. But, given the ambitious scope of Chomsky’s hunches, IMO we have every reason to be sanguine that something like MP is realizable. This, however, is also fodder for another post at another time.