Comments

Showing posts sorted by relevance for query On Wh movement. Sort by date Show all posts
Showing posts sorted by relevance for query On Wh movement. Sort by date Show all posts

Sunday, February 9, 2014

Where Chris Collins enters the fray

Chris sent me this longish response to some of what has appeared in the blog. In the hope of getting him to become a regularish participant in the ongoing discussions I here post his "Response to Norbert." I feel that he let me off lightly, actually. But this said, I think that I can still finds some points to disagree with. I will restrict these to the comments section and hand the floor over to him. Thx Chris.

*****

Response to Norbert

I read with interest Norbert’s recent post on formalization: “Formalization and Falsification in Generative Grammar”. Here I write some preliminary comments on his post.  I have not read other relevant posts in this sprawling blog, which I am only now learning how to navigate. So some of what I say may be redundant. 

For me the quote by Frege in the Begriffsschrift (pg. 6 of the book “Frege and Godel”) indicates what is important when he analogizes the “ideography” (basically first and second order predicate calculus) to a microscope: “But as soon as scientific goals demand great sharpness of resolution, the eye proves to be insufficient. The microscope, on the other hand, is perfectly suited to precisely such goals, but that is just why it is useless for all others.” Similarly, formalization in syntax is a tool that needs to be employed when needed. It not an absolute necessity and there are many ways of going about things (as I discuss below). By citing Frege, I am in no way claiming that we should aim at the same level of formalization that Frege did.

There is an important connection with the ideas of Rob Chametzky (posted by Norbert) in another place on this blog. As we have seen, Rob divides up theorizing into meta-theoretical, theoretical and analytical.  Analytical work, according to Chametzky is: “concerned with investigating the (phenomena of the) domain in question. It deploys and tests concepts and architecture developed in theoretical work, allowing for both understanding of the domain and sharpening of the theoretical concepts.” It is clear that more than 90% of all linguistics work (maybe 99%) is analytical, and that there is a paucity of true theoretical work.

A good example of analytical work would be Chomsky’s “On Wh-Movement”, which is one of the most beautiful and important papers in the field. Chomsky proposes the wh-diagnostics and relentlessly subjects a series of constructions to those diagnostics uncovering many interesting patterns and facts. The consequence that all these various constructions can be reduced to the single rule of “wh-movement” is a huge advanced, allowing one insight into UG. Ultimately, this paper lead to the Move-Alpha framework, and indirectly to Merge (the simplest and most general operation yet).
However, “On Wh-Movement” is what I would call “semi-formal”. It has semi-formal statements of various conditions and principles, and also lots of assumptions are left implicit. As a consequence it has the hallmark property of semi-formal work: there are no theorems and no proofs. Formalization is stating a theory clearly and formally enough that one can establish conclusively (i.e., with a proof) the relations between various aspects of the theory and between claims of the theory and claims of alternative theories.

Certainly, it would have been a waste of time to fully formalize “On Wh-Movement”. It would have expanded the text 10-20 fold at least, and added nothing. This is something that I think Pullum completely missed in his 1989 paper on formalization. The semi-formal nature of syntactic theory, also found in such classics as “Infinite Syntax” by Ross and “On Raising” by Postal, has led to a huge explosion of knowledge that people outside of linguistics/syntax cannot really understand (hence all the lame discussion out there on the internet and Facebook about what the real accomplishments of generative grammar have been), in part because syntacticians are not very good popularizers.
Theoretical work, according to Rob is:  “is concerned with developing and investigating primitives, derived concepts and architecture within a particular domain of inquiry.” There are many good examples of this kind of work in the minimalist literature. I would say Uriagereka’s original work on multi-spell-out qualifies and so does Epstein’s work on c-command, amongst others.

My feeling is that theoretical work (in Chametzky’s sense) is the natural place for formalization in linguistic theory. The reason is that it is possible, using formal assumptions to show clearly the relationship between various concepts, assumptions, operations and principles. For example, it should be possible to show, from formal work, that things like the NTC and Extension condition should really be thought of as theorems proved on the basis of assumptions about UG.  Since NTC and Extension condition are theorems, they can actually be eliminated from UG. And from this, one can wonder if that program can be extended to the full range of what syntacticians normally think about as constraints.
In this, I agree with Norbert who states: “It can lay bare what the conceptual dependencies between our basic concepts are.” Furthermore, as my previous paragraph makes clear, this mode of reasoning is particularly important for pushing the SMT forward. How can we know, with certainty, how some concept/principle/mechanism fits into the SMT? We can formalize and see if we can prove relations between our assumptions about the SMT and the various concepts/principles/mechanisms. Using the ruthless tools of definition, proof and theorem, we can gradually whittle away at UG, until we have the bare essence. I am sure that there are many surprises in store for us. Given the fundamental, abstract and subtle nature of the elements involved, such formalization is probably a necessity, if we want to avoid falling into a muddle of unclear conclusions.

A related reason for formalization (in addition to clearly stating/proving relationships between concepts and assumptions) is that it allows one to clarify murky areas. One of the biggest such areas nowadays is whether syntactic dependencies make use of chains, multi-dominance structures or something else entirely (maybe nothing else). Chomsky’s papers, including his recent ones, make references to chains at many points. But other recent work invokes multi-dominance. What are the differences and relations between these theories and are either of them really necessary? What assumptions about UG does multi-dominance or chains entail? I am afraid that without formalization it will be impossible to answer these questions. I am investigating these questions in my seminar this semester.
These questions about syntactic dependencies interact closely with TransferPF (Spell-Out) and TransferLF, which to my knowledge, have not only not been formalized but not even stated in an explicit manner. Investigating the question of whether multi-dominance, chains or some something else entirely (perhaps nothing else) is needed to model human language syntax will require a concomitant formalization of TransferPF and TransferLF, since these are the functions that make use of the structures formed by Merge.

Minimalist syntax calls for formalization in a way that previous syntactic theories did not. First, the nature of the basic operations is simple enough (e.g., Merge) to make formalization a real possibility. The baroque and varied nature of “transformations” in the “On Wh-Movement” framework and preceding work made the prospect for a full formalization more daunting.

Second, the nature of the concepts involved in minimalism, because of their simplicity and generality (e.g., copies, occurrences), are just too fundamental and subtle and abstract to resolve by talking through them in an informal or semi-formal way. With formalization we can hope to state things in such a way to make clear conceptual and empirical properties of the various proposals, and compare and evaluate them. In fact, I have recently being doing a lot of this with my colleagues, because only recently (by helping to write Collins and Stabler 2012) have I seen what the issues are.
So, in the spirit of Frege, formalization should be a tool for ordinary working syntacticians to clarify their ideas and examine them empirically and conceptually.


Monday, February 15, 2016

The wonder of subjacency

I am currently teaching our Grad syntax 2 course and, not surprisingly, it focuses on the Minimalist Program (MP). Given my predilections (and the influence of Howard Lasnik) I find that one can best appreciate MP by starting with Government-Binding Theory (GB). Jairo Nunes, Kleanthes Grohmann and me used GB as backdrop to MP in our intro book (here). But every time I teach this course I become more and more impressed with the virtues of GB. It is a pretty neat little theory, and, for my money, it still provides the best set of analytical tools in linguistics. In fact, were I charged with the task of describing a new construction or writing the G of a language I would render it in a GB idiom, minimalist technology be damned. However, this is not what I wanted to write about here. Rather, I wanted to sing the praises on one particular sub-part of GB that dealt with a topic that has largely fallen out of research favor but that stands as one of the great scientific accomplishments of Generative Grammar (GG). The topic? Islands and Subjacency. What follows is why I consider it such an achievement.

As everyone knows, Chomsky’s aim in developing the theory of Subjacency (S) was to unify Ross’s islands, the latter having been discovered and described about a decade earlier. The locus classicus of this effort is On Wh Movement (OWM) where Chomsky lays out the story in gory detail.  Here’s a question: what did the unification add to Ross’s original discussion?

One thing it added was unification. Looked at theoretically, Ross’s islands are a motely, a list of domains opaque to movement. From the get-go, it was hard to believe that this list was what FL/UG coded. There has to be some underlying method. Chomsky’s goal was to find it. I do not actually recall his theoretical discontent being widely shared across the GG community (but, in my experience GG hardly ever suffers from the mental unease that poor theory regularly generates in Chomsky). At any rate, in unifying Ross’s islands, OWM tries to explain why the islands we find are the islands we have. In fact, OWM tries to tie the existence of islands to general computational considerations thereby providing what is, in retrospect, an excellent paradigm of Minimalist thinking. Here is what OWM says:

… the island constraints can be explained in terms of general and quite reasonable computational properties of formal grammar (i.e. subjacency, a property of cyclic rules that states, in effect, that transformational rules have a restricted domain of potential application; SSC, which states that only the most prominent phrase in an embedded structure is accessible to rules relating it to phrases outside; PIC, which stipulates that clauses are islands subject to the language specific escape hatch..). If this conclusion can be sustained, it will be a significant result, since such conditions as CNPC and the independent wh-island constraint seem very curious and difficult to explain on other grounds. (p. 89; On WH Movement, my emphasis).

So the list like nature of the islands becomes comprehensible when viewed from a more general computational perspective. And this is indeed a virtue.

But, and I want to emphasize this, this unification is not, as it stands, an empirical argument in favor of S. Taken at face value, what OWM demonstrates is that it is possible to unify Ross’s islands on a more rational basis, but just unifying them does not show that this unification is empirically fecund or justified.

Happily, the unification proved to be empirically very fertile indeed. OWM provides two ways that S logic could lead to the discovery of novel data.

First, it provides a general method for discovering which kinds of dependencies should be subject to island effects. OWM has a long and interesting discussion of comparative constructions and notes that given the nature of the unification proposed, comparatives should be formed by movement. This was somewhat unconventional at the time (though the work is based on some earlier work by Richie Kayne that argued for this conclusion). In fact, the most carefully worked out theory of comparatives (due to Bresnan) treated comparatives as products of a deletion operation, rather than as products of movement. If memory serves, there was quite a bit of very vigorous debate on this topic over the next little while, including at the UCSD conference where OWM was originally presented. This debate became quite heated and gave lowly grad students like me an appreciation of the old adage: when elephants fight what gets hurt is the grass. At any rate, this was one consequence of the unification that OWM emphasizes. 
A digression: could Ross’s analysis been used as a diagnostic of movement? This is, in effect, what OWM does. It assumes that if dependency D obeys islands yet allows unbounded dependency (btw, this second conjunct is a critical yet often ignored part of S reasoning) then the dependency must be the product of movement. Could Ross’s theory be interpreted in the same way? Not really. Recall, that for Ross, what makes an island and island is not the movement (movement out of islands was fine for Ross). Rather what makes an island is chopping the resumptive pronoun that movement leaves behind. In other words, for Ross, islands restrict chopping, not movement. For Chomsky, S restricts the movement and resumption is analyzed as a non-movement dependency precisely because it does not show island effects.[1] Given this, comparatives are a very good place to empirically distinguish Ross’s theory from S-theory. Why? Because comparatives have no apparent resumptive analogues like DP movement cases do (*John is taller than Bill is it/that/such). But if there are no resumptives then there can be no chopping and so no expectation of islands. This would make deletion the natural generative operation sub-serving comparatives. Thus OWM’s argument that comparatives are actually products of movement, was an empirical argument for the S view of islands.[2]
The second empirical argument for the OWM unification came from a crop of new islands. Thus, the OWM story implied that complex DPs should be islands for extraction. This implied that we should find subject islands (which we more or less do: *What do pictures of hang in the National Gallery) but also that objects should be islands (which is far less evident: What did Bill paint pictures of). OWM spends some time trying to get out from under the problems that object extraction creates. To the degree that it succeeds, then the predicted presence of subject islands is an empirical plus for S-theory.[3]
The third empirical argument in favor of S is by far the best and, if my recollection is correct, the most wow-inducing. It’s successive cyclic A’ movement. The unification of islands predicted that unbounded movement (movement that shows no island effects) is nonetheless derivationally bounded in that it is made up of a series of small bounded steps. Ross’s theory made no such prediction, Indeed, prior to S-theory there was no reason to believe it to be true. Unbounded dependencies were considered to be perfectly reasonable operations. S-theory implies that, at least for one class of dependencies (i.e. movement), such unboundedness is an illusion. This was (and is) a hell of an implication. There are many many languages where there is little evidence suggesting that anything like this is true (English being a good example of one). But, as we soon discovered (and by ‘we’ I mean GGers), it was TRUE (insert fireworks and brass bands here).
I was a grad student in Cambridge when the empirical evidence started trickling in. Jean Yves Pollock gave versions of the deservedly famous paper he co-wrote with Richie Kayne on stylistic inversion in French. If memory serves, Esther Torrego’s equally excellent paper was floating around when I was still a Cambridge denizen. As most now know, this trickle soon became a torrential stream of results with many languages providing overt evidence for cyclic Wh movement (Irish, Chamorro a.o.) At any rate, that this implication of S-theory was apparently true (or at least had non-obvious data that could be explained by it) was stunning. This is what good science does: its theories imply something unexpected and the unexpected turns out to be the case. It was great. And this was the evidence that really sold S-theory.
Let me emphasize the important argumentative structure: Unifying islands as in OWM implies that all movement, even that which does not manifest island-like properties, is local. Thus islands imply successive cyclic C to C movement. The discovery that this prediction holds is stunning confirmation of the unification of Ross’s islands in OWM and a strong confirmation of S-theory.
So, if anyone asks you what unifying islands brought to the table, successive C to C movement (or edge of domain to edge of next higher domain) is one of the biggies. It served a bit like the Syntactic Structures analysis of affix-hopping and do-support in that it sold S-theory with its aha effect and thereby made it widely accepted.

One last virtue: S-theory served as a bridge to other parts of cogsci. For example, S-theory had very natural interpretations in the context of parsing theories (E.g. Berwick and Weinberg) and learnability theories (e.g. Culicover and Wexler). S-theory served as a grammatical bridge to, IMO, the richest interaction between GG and other parts of cognition witnessed to date. In fact, like the income of most Americans, GG has receded from this high point, which, is really too bad.

So, what did S-theory add? It unified islands, allowed for a refinement of our understanding of movement, led to the postulation of new islands, implied that long movements were made up of short steps and served as a productive bridge to other parts of cognition.

And it has one last virtue of contemporary relevance. It serves (IMO) as and excellent  (maybe even the best) example we have of what linguistic theory should aim for. It is our poster child for for GGs scientific bona fides, which makes it odd that S-theory appears not to be a central part of the grad syntax curriculum anymore. I say this on the basis of very cursory investigation, actually just one or two discussions with recently minted PhDs. For the reasons noted above, this is too bad. It is a beautiful GG discovery and deserves to be regularly trumpeted as one of GGs great achievements. So, next time you are at a party and there is a lull in the conversation, remember the wonders of S-theory.



[1] Note that given current work arguing that resumption involves movement raises interesting questions about S theory. Given my partiality to this excellent idea (Demirdache is a leading exponent of this line of thinking), I think that it is worth revisiting some of the OWM assumptions, though I will refrain from doing so here.
[2] There was even independent dialectal evidence in favor of the movement analysis of comparatives: John is taller than what Bill is. The ‘what’ sure looks like a relative pronoun. This observation was due to Kayne, if I recall correctly (way to go Richie!).
[3] Wh islands another “novel” island, one that in fact Ross argued at length did not exist. As you all know, the status of Wh islands is somewhat variable cross linguistically and even among speakers of the same language. Sprouse’s thesis shows that they more or less display island-like acceptability signatures. However, whatever their status (maybe they are semantic rather than syntactic as some have argued), theoretically, they fell under the OWM unification only if one makes additional assumptions about the structure of C (how many “escape” hatches it contains). This assumption was usefully investigated empirically by Reinhart and Comorovski. They showed that the degree of freedom that S-theory allowed for was in fact empirically consequential, thus providing an indirect argument in favor of the unification along OWM lines.

Monday, February 10, 2014

Where Norbert posts Chris's revised post after screwing things up

In my haste to get Chris's opinions out there, I jumped the gun and posted an early draft of what he intended to see the light of day. So all of you who read the earlier post, fogetabouit!!! You are to ignore all of its contents and concentrate on the revised edition below.  I will try (but no doubt fail) never to screw up again.  So, sorry Chris. And to readers, enjoy Chris's "real" post.

*****

Why Formalize?
I read with interest Norbert’s recent post on formalization: “Formalization and Falsification in Generative Grammar”. Here I write some preliminary comments on his post.  I have not read other relevant posts in this sprawling blog, which I am only now learning how to navigate. So some of what I say may be redundant. Lastly, the issues that I discuss below have come up in my joint work with Edward Stabler on formalizing minimalism, to which I refer the reader for more details.
I take it that the goal of linguistic theory is to understand human language faculty by formulating UG, a theory of the human language faculty. Formalization is a tool toward that goal. Formalization is stating a theory clearly and formally enough that one can establish conclusively (i.e., with a proof) the relations between various aspects of the theory and between claims of the theory and claims of alternative theories.
Frege in the Begriffsschrift (pg. 6 of the Begriffschrift in the book Frege and Godel) analogizes the “ideography” (basically first and second order predicate calculus) to a microscope: “But as soon as scientific goals demand great sharpness of resolution, the eye proves to be insufficient. The microscope, on the other hand, is perfectly suited to precisely such goals, but that is just why it is useless for all others.” Similarly, formalization in syntax is a tool that should be employed when needed. It not an absolute necessity and there are many ways of going about things (as I discuss below). By citing Frege, I am in no way claiming that we should aim for the same level of formalization that Frege aimed for.
There is an important connection with the ideas of Rob Chametzky (posted by Norbert in another place on this blog). As we have seen, Rob divides up theorizing into meta-theoretical, theoretical and analytical.  Analytical work, according to Chametzky is: “concerned with investigating the (phenomena of the) domain in question. It deploys and tests concepts and architecture developed in theoretical work, allowing for both understanding of the domain and sharpening of the theoretical concepts.” It is clear that more than 90% of all linguistics work (maybe 99%) is analytical, and that there is a paucity of true theoretical work.
A good example of analytical work would be Noam Chomsky’s “On Wh-Movement”, which is one of the most beautiful and important papers in the field. Chomsky proposes the wh-diagnostics and relentlessly subjects a series of constructions to those diagnostics uncovering many interesting patterns and facts. The consequence that all these various constructions can be reduced to the single rule of wh-movement is a huge advance, allowing one insight into UG. Ultimately, this paper led to the Move-Alpha framework, which then led to Merge (the simplest and most general operation yet).
 “On Wh-Movement” is what I would call “semi-formal”. It has semi-formal statements of various conditions and principles, and also lots of assumptions are left implicit. As a consequence it has the hallmark property of semi-formal work: there are no theorems and no proofs.
Certainly, it would have been a waste of time to fully formalize “On Wh-Movement”. It would have expanded the text 10-20 fold at least, and added nothing. This is something that I think Pullum completely missed in his 1989 NLLT contribution on formalization. The semi-formal nature of syntactic theory, also found in such classics as “Infinite Syntax” by Haj Ross and “On Raising” by Paul Postal, has led to a huge explosion of knowledge that people outside of linguistics/syntax do not really appreciate (hence all the uninformed and uninteresting discussion out there on the internet and Facebook about what the accomplishments of generative grammar have been), in part because syntacticians are generally not very good popularizers.
Theoretical work, according to Chametzky is:  “is concerned with developing and investigating primitives, derived concepts and architecture within a particular domain of inquiry.” There are many good examples of this kind of work in the minimalist literature. I would say Juan Uriagereka’s original work on multi-spell-out qualifies and so does Sam Epstein’s work on c-command, amongst others.
My feeling is that theoretical work (in Chametzky’s sense) is the natural place for formalization in linguistic theory. One reason is that it is possible, using formal assumptions to show clearly the relationship between various concepts, assumptions, operations and principles. For example, it should be possible to show, from formal work, that things like the NTC, the Extension Condition and Inclusiveness should really be thought of as theorems proved on the basis of assumptions about UG.  If they were theorems, they could be eliminated from UG. One could ask if this program could be extended to the full range of what syntacticians normally think of as constraints.
In this, I agree with Norbert who states: “It can lay bare what the conceptual dependencies between our basic concepts are.” Furthermore, as my previous paragraph makes clear, this mode of reasoning is particularly important for pushing the SMT (Strong Minimalist Thesis) forward. How can we know, with certainty, how some concept/principle/mechanism fits into the SMT? We can formalize and see if we can prove relations between our assumptions about the SMT (assumptions about the interfaces and computational efficiency) and the various concepts/principles/mechanisms. Using the ruthless tools of definition, proof and theorem, we can gradually whittle away at UG, until we have the bare essence. I am sure that there are many surprises in store for us. Given the fundamental, abstract and subtle nature of the elements involved, such formalization is probably a necessity, if we want to avoid falling into a muddle of unclear conclusions.
A related reason for formalization (in addition to clearly stating/proving relationships between concepts and assumptions) is that it allows one to compare competing proposals. One of the biggest such areas nowadays is whether syntactic dependencies make use of chains, multi-dominance structures or something else entirely. Chomsky’s papers, including his recent ones, make references to chains at many points. But other recent work invokes multi-dominance. What are the differences between these theories?  Are either of them really necessary? The SMT makes it clear that one should not go beyond Merge, the lexicon, and the structures produced by Merge. So any additional assumptions needed to implement multi-dominance or chains are suspect. But what are those additional assumptions? I am afraid that without formalization it will be impossible to answer these questions.
Questions about syntactic dependencies interact closely with TransferPF (Spell-Out) and TransferLF, which to my knowledge, have not only not been formalized but not even stated in an explicit manner (other than the initial attempt in Collins and Stabler 2013). Investigating the question of whether multi-dominance, chains or some something else entirely (perhaps nothing else) is needed to model human language syntax will require a concomitant formalization of TransferPF and TransferLF, since these are the functions that make use of the structures formed by Merge. Giving explicit and perhaps formalized statements of TransferPF and TransferLF should in turn lead to new empirical work exploring the predictions of the algorithms used to define these functions.
A last reason for formalization is that it may bring out complications in what appear to be innocuous concepts (e.g., “workspaces”, “occurrences”, “chains”).  It will also help one to understand what alternative theories without these concepts would have to accomplish. In accordance with the SMT, we would like to formulate UG without reference to such concepts, unless they are really needed.
Minimalist syntax calls for formalization in a way that previous syntactic theories did not. First, the nature of the basic operations is simple enough (e.g., Merge) to make formalization a real possibility. The baroque and varied nature of transformations in the “On Wh-Movement” framework and preceding work made the prospect for a full formalization more daunting.
Second, the nature of the concepts involved in minimalism, because of their simplicity and generality (e.g., the notion of copy), are just too fundamental and subtle and abstract to resolve by talking through them in an informal or semi-formal way. With formalization we can hope to state things in such a way to make clear the conceptual and the empirical properties of the various proposals, and compare and evaluate them.

My expectation is that selective formalization in syntax will lead to an explosion of interesting research issues, both of an empirical and conceptual natural (in Chametzky’s terms, both analytical and theoretical). One can only look at a set of empirical problems against the backdrop of a particular set of theoretical assumptions about UG and I-language. The more that these assumptions are articulated; the more one will be able to ask interesting questions about UG.

Wednesday, January 23, 2013

Minimalism and On Wh Movement


The Generative enterprise as Chomsky envisioned it has been based on five separate but related questions:

1.     What does a given native speaker know about his language? Or, what do particular Gs look like?
2.     What makes it possible for native speakers to acquire Gs? Or, what does UG look like?
3.     How did UG arise in the species? Or, what must be added to non-linguistic cognition to get UGs?
4.     How do native speakers use Gs in performance? Or, how are Gs used to produce and comprehend sentences in real time?
5.     How do brains code for G and UG?

These questions, though separate are clearly interconnected. For example, it’s very hard to impossible to consider the properties of UG without knowing how particular Gs are put together.  It’s pointless to engage in questions about how UG could have arisen in the species without knowing anything about UG. It’s hard to study how people use their Gs in the absence of descriptions of the Gs that are used.  And, last, it’s hard to study how brains embody G and UG without knowing anything about brains, Gs or UGs. All of this should be evident.  What is less obvious, but still true, is that even when we know a non-trivial amount about the things we are trying to relate, relating them may be really tough. There are many reasons for this. Let’s consider two.

In the last several posts I considered the relation between (2) and (3) above.  I noted that we have a pretty good description of what UG does, e.g. it regulates movement and construal dependencies, identifies the kinds of phrase structures that are linguistically admissible and where expressions that are displaced can be interpreted both phonetically and semantically. GB is a pretty good effective theory of these relations.  However, it is a problematic fundamental theory. Why? Because it’s hard to see how something with the special purpose predicates and intricate internal structure of GB could have arisen in the species. It’s just too different from what we find in other domains of human cognition and much too elaborately structured. Consequently, it’s entirely opaque how something this apparently complex and cognitively idiosyncratic could have arisen in the species, especially in the (apparently) short time available. This motivates a re-think of GB’s depiction of UG. The Minimalist Program is a research program aimed at reducing GB’s parochialism and intricacy by reanalyzing heretofore language specific operations in more general cognitive terms (e.g. Merge) and unifying conditions on grammatical processes in more generic computational terms (e.g. cyclicity as monotonicity/no tampering). As I’ve suggested in other posts, this strategy recapitulates Chomsky’s in ‘On Wh Movement’ (OWM) and I have suggested that OWM provides a unification strategy worth emulating. What specifically did Chomsky do in OWM?

First, he adopted Ross’s theory as an effective account. Thus, he accepted that that lay of the land described by Ross was roughly correct. Why ‘roughly’? Because, though Chomsky adopted Ross’s descriptions of strong islands, he tweaked the data in light of the details of the unifying Subjacency account. In particular, the relevant island data included Wh islands (pace Ross’s theory, which treated Wh-islands as porous) and was generalized to cover subjects in general (not only sentential subjects as in Ross).  Thus, though OWM largely adopted Ross’s description of the data, it modified it as well. 

Second, Chomsky unified the various islands under the theory of movement by unifying the movement constructions under a ‘Move alpha’ rubric. Whereas in Ross rules were effectively constructions, Chomsky distilled out a movement component common to all the cases subject to the Subjacency Condition (i.e. ‘move alpha’) and proposed that all and only the move alpha operation was subject to the locality requirements eventuating in movement effects when violated.[1] Thus, some constructions were treated as composites; one part move alpha one part specific head with their own particular grammatical contributions (‘criteria’ in Rizzi’s sense). And all that mattered in computing locality was the movement part.  In sum, what makes a construction subject to islands in OWM is that it is composed of a move alpha part. None of its other properties matter.[2]

Chomsky adopted a very strong version of this claim: as noted, all and only move alpha is subject to subjacency restrictions. Arguing for this constitutes the better part of OWM. The most interesting empirical component was reducing various kinds of deletion operations investigated by Bresnan and Grimshaw to constructions mediated by movement, most particularly their analysis of comparatives as deletion operations.[3] At any rate, sitting where we are today, it looks like Chomsky’s reanalysis has won the day and that islands are now taken as virtual diagnostics of a movement dependency.

The third leg of the analysis involved the Subjacency Condition itself. This involved several proposals: one concerning the inventory of bounding nodes (which nodes counted in computing distance), one concerning which nodes had “escape hatches,” aka comp(lementizer)s, and one concerning the fine structure of these comps (in particular how many slots it contained).  In elaborating these in OWM Chomsky noted that movement via escape hatches effectively allowed move alpha to finesse the Specified Subject Condition (SSC) and the Propositional Island Constraint (PIC). These were two proposed universals that were retained in revised form in GB, though in later work move alpha was not thought to be subject to them.[4] At any rate, it is interesting to see what Chomsky took to be the computational implications of Subjacency Theory:

… the island constraints can be explained in terms of general and quite reasonable computational properties of formal grammar (i.e. subjacency, a property of cyclic rules that states, in effect, that transformational rules have a restricted domain of potential application; SSC, which states that only the most prominent phrase in an embedded structure is accessible to rules relating it to phrases outside; PIC, which stipulates that clauses are islands subject to the language specific escape hatch..). If this conclusion can be sustained, it will be a significant result, since such conditions as CNPC and the independent wh-island constraint seem very curious and difficult to explain on other grounds. [my emphasis] (p. 89; OWM).

Note what we have here: an attempt to motivate the particular witnessed properties of the theory of bounding on more general computational grounds. This should sound very familiar to minimalist ears- restricted computational domains, prominent targets of operations (think labels in place of subjects) etc.  As we know, theories embedding subjacency were subsequently built into interesting proposals with non-trivial consequences for efficient parsing (e.g. Marcus, Berwick and Weinberg). 

The upshot? OWM provides a good model for minimalist unification ambitions.  Extract out a common core operation in seemingly disparate “constructions” and propose general conditions on the applications of this common rule to unify the indicated dependencies. Minimalism has already taken steps in this direction. For example, case theory was unified with movement theory in early minimalism (e.g. Chomsky 1993) by having case discharged in the specifier positions of relevant heads.  Phrase structure theory has been unified with movement theory by treating both as products of Merge (E-merge and I-merge being instances of the same operation as applied to different inputs). More ambitiously (stupidly?) still, some (e.g. yours truly, Boeckx, Nunes, Idsardi, Lidz, Kayne, Zwart, Polinsky, Potsdam) have proposed treating construal rules like Control, reflexivization, and pronominal binding as species of movement hence unifying all of these operations with Merge. 

These last mentioned proposals require abandoning two key features of GB’s view of movement: (i) that movement into thematic positions is barredand (ii) all movement result in phonetic gaps at launch sites. [5]  (i) is a plausible consequence of removing D-structure as a “level” and (ii) constitutes a return to earlier versions of Generative Grammar in which transformations included the addition of some designated lexical material (e.g. there, reflexives, bound pronouns, etc.). It is very very unclear at this moment whether this whole range of unifications is possible.[6] I personally find the data uncovered and the arguments made to be highly suggestive and the empirical hurdles to be less daunting than they appear.  However, this is a personal judgment, not one shared by the field as a whole (sadly).  That said, I would argue that extending the OWM reasoning to other parts of the grammar to unify the modules is the right thing to try if one hopes to address Darwin’s problem in (3) above. Why? Because if this kind of unification succeeded, grammars with the phenomenological properties of GB’s version of UG would follow from the addition of Merge to the cognitive repertoire of our apish ancestors.  In other words, we would have isolated a single change sufficient to allow the development of an FL with GB observable properties. Such a theory could be considered fundamental.

Interestingly, such a theory would not only plausibly provide an answer to Darwin’s problem (in (3)), it would also provide one for Boraca’s Problem (in (5)).  Indeed, this post was intended to address this point here (as the title indicates), but as I’ve rambled on long enough, let me delay addressing how Minimalism relates to question (5) until the next post.


[1] Please note: This is a little bit of Whig history and it’s not really fair to Ross. Ross suggested that all island sensitive constructions involved a specific rule of chopping, an operation that deleted the resumptive pronoun that was part of the constructions of interest. Interestingly, the intuition behind this part of Ross’s analysis has enjoyed somewhat of a revival in recent work (see Merchant and Lasnik in particular), which traces island effects as PF phenomena rather than restrictions on derivational operations as Chomsky originally proposed.
[2] This is a fact worth savoring. A priori there is nothing odd about having the specific construction determine locality rather than the created move alpha dependency.  There is nothing contradictory in assuming, e.g. that Focus and question formation would obey islands but that Topicalizations, comparatives and relativizations would not.  But this does not seem to be the way things panned out.
[3] The other really interesting consequence of Subjaency Theory was the implication that all movement was successive cyclic. This prediction was confirmed in work by Kayne and Pollock, Sportiche, and Torrego a.o.  In my opinion, this is still one of the nicest “predictions” any theory in syntax has ever made.
[4] The reason is that in OWM A’-traces were treated as anaphoric elements. This was revised in later work due to some empirical findings due to Lasnik.
[5] To be honest, this is my reconstruction of their results. Kayne and Zwart, for example, retain the theta criterion by assuming that there is a lot more doubling than meets the eye.
[6] Indeed, I believe (actually I think I know this, but let me be coy) that Chomsky is very skeptical about the unification of construal with movement (with the possible exception of reflexivization).  This may account for why binding, when discussed, is suggested to be a CI interface operation fed by the grammar rather than a direct product of the grammar as in all earlier generative accounts.