Comments

Showing posts with label the Merge Hypothesis. Show all posts
Showing posts with label the Merge Hypothesis. Show all posts

Tuesday, January 15, 2019

Movement, islands and the ECP

Some papers reset the research agenda. This one by Lu and Yoshida (L&Y), I believe, is one of those (here is a slide conveying the basic point. The paper is under submission at LI and I assume it will be accepted and rapidly published (if not this is will tell us more about LI than it will about the quality of this paper)). The topic is the island status of Wh-in-situ (WIS) constructions in Chinese. The finding is that using judgment studies of the Sprouse experimental syntax (ES) variety provides evidence for two stunning conclusions: (i) that WISs respect islands and (ii) that there is no evidence for an argument/adjunct distinction wrt WISs. Both data points are theoretically pregnant and this post will largely concentrate on drawing out some of the implications. Many of these are mentioned in the paper (yup, I have a draft), so are not original with me. Let’s start.

L&Y is motivated by the premise that ES provides a useful tool for the refining linguistic judgments. The idea, as Sprouse has convincingly argued, is that grammatical complexity should induce a super-additivity effect in well-constructed judgment experiments (see, e.g. here and here for discussion and here for a nice review of the methodology). Importantly, super-additivity profiles arise in cases where less involved rating studies find nothing indicating un-grammaticality.

Before pressing on, let’s make an important and obvious point: all GGers distinguish (or should distinguish) acceptability from grammaticality. Acceptability is a probe for grammaticality. Acceptability is an observable property of utterances. Grammaticality is an abstract property of I-linguistic mental representations. Grammaticality is inferred from acceptability under the right conditions (given the right controls as realized by the appropriate minimal pairs). All of this is old hat, but a still very stylish and durable hat. 

Happily, for most of what we have done in GG, acceptability closely tracks grammaticality, but we also know that the two notions can and do diverge (see here for some discussion). ES is particularly useful for cases where this happens and the simple judgment elicitation procedure (e.g. ask a native speaker) indicates all is well. Diogo Almeida has dubbed cases of the latter “subliminal.” One of ES’s important contributions to syntax has been the discovery of such subliminal effects (SE), SEs being cases where ES procedures reveal super-additivity effects while more regular elicitation suggests grammaticality. So, for example, we now have many examples where standard elicitation has indicated that a certain dependency in a certain language shows no island sensitivity (i.e. the sentences are judged (highly) acceptable) while ES techniques indicate sensitivity to these same island effects (i.e. the relevant data display super-additivity effects).

We also find the converse: standard techniques indicating a profound difference in acceptability, while ES techniques showing nothing at all.[1]All in all then, ES has provided a useful additional kind of data, one that is often more sensitive to G structure than the quick and dirty (and largely accurate and hence very useful) standard judgment techniques, which sometimes fail to track these. 

So, back to the main point: L&Y is an ES study of WISs in Chinese and it has two important findings: that allWISs in Chinese exhibit relative clause island effects (henceforth RCI) (i.e. they alldisplay the super-additivity profile) and that there is no ES evidence that long “why” movement from an RCI is appreciably worse than long “why” movement absent an RCI (i.e. these cases when contrasted do not show a super-additvity profile). The first result argues that WISs are island sensitive and the second argues that there is no additional ECP effect distinguishing WISs like who/what from WISs like why. If correct, this is very big news, and, IMO, very welcome news. Let me say why.

First, as L&Y emphasizes, this result rules out most of the standard approaches to WIS constructions. In particular the result rules out two kinds of theories: (i) accounts that distinguish between overt movement vs covert movement (e.g. Huang’s) and treat island effects as effectively reflexes of overt movement (say, via a chain condition at SS) and (ii) theories that postulate two different kinds of operations (Movement vs Binding) to license WISs with movement subject to islands and binding exempt from them (as in, say, a Rizzi-Cinque approach to ECP effects). Both such kinds of theories will have problems with the apparent fact that WISs induce super-additivity effects.

It is worth noting, furthermore, that the sensitivity of WISs to islands is not the only example of apparent non-movement generated structures being island compliant. The same holds wrt resumptive pronoun (RP) constructions. These also appear to respect islands despite the absence of the main hallmark of movement (i.e. a gap in the “movement” site).[2]Both this RP data and now the WIS data point to the same conclusion: that island effects are notPF effects.[3]From my reading of the literature, this is the most popular current approach to islands and it has some terrifically interesting evidence in support (in particular the fact that some ellipsis (i.e. sluicing) obviates island violations). However, if L&Y are right, then we may have to rethink this assumption (see note 3 however).

Indeed, I would go further (and here it is NH speaking rather than L&Y). There have long been two general approaches to islands. 

First, we have Chomsky’s view of subjacency elaborated in ‘On wh movment’ that treats islands as reflecting bounds on the computational procedure. Island effects reflect the subjacency condition (aka PIC), which bounds the domain of computation (an idea motivated by the reasonable assumption that bounding a domain of computation makes doing computations more tractable).[4]

The second approach can be traced back to Ross’s thesis (islands restrict chopping rules) but has been developed as part of the linearization industry spurred by Kayne’s seminal work and mooted most explicitly by Uriagereka.[5]

The L&Y results argue pretty strongly, IMO, for Chomsky’s original conception precisely because they appear to hold whether or not the construction involves an obvious phonetic gap (gaps being problematic as they undo linearizations). If this is so, then it argues against linearization based approaches to the problem (leaving, of course a very big question: what to do about sluicing).[6]

We can go further still. The L&Y results also argue for a Merge only syntax. Here is what I mean. IMO, the central empirical thesis of the Minimalist Program (MP) is the Merge Hypothesis (MH). MH is the claim that the only specifically linguistic operation of FL is Merge. This entails that allG dependencies are merge mediated. The strong version of the thesis excludes operations like long distance Agree, which spans the same domains as I-merge but is a different operation. Note that it is natural to suppose that I-merge is movement and Agree is some kind of binding or feature sharing. At any rate, the classical conceptions gain empirical benefit from the “observation” that WISs do not display island effects. Why? Because, we might say, they are licensed by Agree not by I-merge and only the latter (being the MP analogue of movement) is subject to subjacency (or its current analogue). But as L&Y indicates this is precisely the wrong conclusion. WISs are subject to islands. A merge only syntax insists that all A’-dependencies are formed in the same way, via I-Merge, as this as the only way to establish any non-local grammatical dependency. So if WISs are G licensed, then they must be G licensed via I-merge and so will form a natural class with overt Wh movement. And this is what L&Y find. Both show super-additivity effects across islands. Thus, L&Y’s findings are what we should expect from a merge only syntax and it cautions against larding this best of MP theories with Agree/Probe-Goal titivations.[7]

We can milk a second important conclusion from L&Y. It solves a giant problem for MP. Which problem? The problem of unifying subjacency with the ECP. I have suggested elsewhere (see here) that the argument/adjunct asymmetries at the heart of the ECP are very MP problematic. This is so for a variety of reasons. The three that move me most are the fact that the ECP is a trace licensing condition and MP eschews traces, the huge theoretical redundancy between ECP and subjacency, and the “ugliness” of the basic technical machinery required to allow the ECP to track the argument/adjunct asymmetry. One of the nice implications of L&Y is that we need not worry about the problems that the ECP generates for MP because the theoretical apparatus is based on a mistaken description of the data. If L&Y is right, then there is no argument/adjunct asymmetry. Poof, the MP problem disappears and with it the ad-hoc theoretically unmotivated (within MP) technical apparatus required to track it.  

Of course, this overstates matters. It behooves us to go over the ECP data more carefully and see how to resolve the difference in acceptability that the standard literature identified. Why after all if all Whs are created equal do long distance adjuncts resist movement more fiercely than do long distance arguments?[8]L&Y offer a suggestion (no spoiler from me, read the squib when it comes out). But whatever the right answer is, it does not rest on making an invidiousgrammaticaldistinction between the two kinds of dependencies. And this is just what MP needs in order to start distinguishing ECP effects from the ECP theoretical apparatus in GB.

Let me hit this a bit harder. The ugliest parts of theories of A’-dependency within GB arise in response to argument/adjunct asymmetry effects. The technical machinery in Barriers (built on Lasnik and Saito foundations), though successful empirically (IMO, Lasnik and Saito’s theory was considerably more empirically effective than Barriers) had little of the virtual conceptual necessity MPers pine for. Nor were alternative theories (Generalized Binding, Connectedness) much prettier. Nonetheless, we put up with that stuff and developed it theoretically because it appeared to be empirically called for. The right aesthetic conclusion should have been (and actually was) that it was too contrived to be correct. L&Y provides courage for our aesthetic convictions. We should have judged these theories as suspect because ugly, though we would have been empirically premature in drawing that conclusion. Given L&Y, the facts are not what we took them to be despite reflecting very different acceptability profiles. 

There is a moral here, and you can all guess what it is but I cannot resist making it explicit anyhow. L&Y provide evidence for a methodological precept that we fail to respect enough: facts can change, no less than theory can. Or to put this another way: just as we can make theoretical wrong turns that we come to revise, we can adopt empirical generalizations that turn out to be misleading. The standard view is that data is hard and theory is fluffy and when the two clash it is best to revise the theory than rethink the data. L&Y provides a case where this is reversed. And I say hooray!

Let me make one more point and I will end this overly long post (long, and yet, filled with endlessly many loose ends). L&Y exemplifies something that I think is important. It is an empirical paper whose purpose is to directlytest a core theoretical assumption. This is not something that we generally see within syntax. Most papers are not out to test theoretical assumptions. Most papers use theory to explore more data. Theory might be tested but it is generally a by-product of better descriptive coverage. L&Y works differently. It starts from the theory and constructs an empirical intervention to probe it. Moreover, the question is quite precise and the assumptions required to answer it are clear within the confines of the project. This has all the look and smell of an honest to god experiment, a process whereby we query the theory using a relatively well-understood probe. Both empirical methods of exploration are worthwhile, but they are different and it is only relatively recently, I think, that we are seeing examples of the second experimental kind gaining traction. 

Curiously (perhaps), a feature of this second kind of paper is that the paper is short. L&Y is a squib. Empirical explorations in linguistics often read like novellas. L&Y is very definitely a very very short story. I would like to suggest that experimental papers like L&Y reflect the scientific health of linguistics. It is now possible to ask a sharp question, and give a sharp answer. We need more of these kinds of short pointed experimental forays into the data starting from well-formulated theoretical starting points.

That’s it. I have gone on far too long. L&Y is terrific. If correct, it is very important. I personally hope the results stand up. It would go a long way to cleaning up a particularly untidy part of syntactic theory and thereby further vindicating the promise of MP, indeed a particularly strong version of MP, one that endorses a Merge only conception of grammar.


[1]ES techniques have, for example, suggested that adjunct island effects might not be of a piece with other islands as they (often) fail to display super-additivity effects.
[2]To be slightly more careful and squinting at the ES results wrt resumptives the following is more accurate: resumptives uniformly ameliorate fixed subject constraint violations but note “mere” subjacency violations. Thus, resumptives inside islands seem to show the same super-additvity profiles as their moved counterparts.
[3]Perhaps a more felicitous way of putting matters is that RPs and WISs are also products of I-merge and so expected to be subject to islands. One reason for treating islands as PF effects is to capture the distinctionbetween these cases and more conventional examples of overt movement. If, however, they pattern the same then the motivation for treating islands as PF effects weakens. This said, I am pretty sure it is possible to model these cases as formed via movement/I-merge by, for example, treating them as cases of remnant movement, the moved Wh or Q morpheme starts as part of a doubled structure including the RP or WIS. 
[4]See Chomsky’s ‘On wh movement’ for discussion. See herefor more prose.
[5]Versions of this original idea were developed by many people including Hornstein, Lasnik and Uriagereka, and Fox and Pesetsky. The idea centers on the idea that the problem with movement is that it reorders elements and so can come into conflict with the ordering algorithm. In this sense, gaps are a big deal and what distinguish movement from other kinds of long distance dependencies like binding.
[6]Or again, it argues for treating the operation (e.g. movement) as in need of constraint rather than the output of the operation (a gap, or new linear order). What plausibly unifies cases of “overt” WH (as in English), “covert” WH (as in Chinese) and resumptive WH (as in Hebrew) is that they all involve relating an A’ element to a non-local syntactic position that can be arbitrarily far away. IT’s the span that seems to matter, not what sits at the tail of the chain.
[7]RP constructions must also be formed via I-merge and so too all forms of binding. This requirement fits well with the observation that RPs obey islands. Binding, especially pronominal binding, is likely to be more problematic. As they say at this point in a journal paper in reply to referee 2; these are topics for further research.
[8]As I have noted in other places, this description of the ECP contrast is not quite correct, as we have known since Rizzi’s work on minimality. The distinction seems less a matter of argument vs adjunct than object centered vs non object centered quantification. But this is a topic for another time.

Monday, June 5, 2017

The wildly successful minimalist program II

In an earlier post (here) I mentioned that I was asked to write about the Minimalist Program (MP) for a volume aimed at comparing various linguistic  “frameworks.” I am personally skeptical about “framework differences.” I personally find them to be more notational variants of common themes (indeed, when pressed, I have been known to complain that H/GPSG, LFG, RG, are all “dialects” of GB) than actual conceptually divergent perspectives. The main reason is that most linguistic theory lacks any real depth and most of the frameworks have the wherewithal to mimic one another’s (ahem) deep insights. So, where others see ideological divergence, I tend to see slight differences in accent.

That said, I have a further gripe about this kind of volume. Even were I to recognize that these different frameworks empirically competed (or should compete) I do not think that MP should be included in the race. MP is not an alternative to GB (or its various dialects) but presupposes that the results of GB are (largely) empirically correct. MP builds on GB results (and its dialects) and aims to conserve these results. Thus, it is not intended (or should not be intended) as a wholesale replacement. For an MPer, GB is wrong the way that Newtonian Gravitation is wrong when compared to General Relativity: the latter theory derives the former as a limiting case. It does not reject the former as misguided, rather it treats it as descriptive rather than fundamental. Indeed, an important part of the argument in favor of General Relativity is that it can derive Newton as a special case. If it could not, that would be an excellent argument that it was fundamentally flawed.

And the MP point? From the MP perspective, as M Antony might have put matters: MPers come to (largely) praise (and incorporate) GB not to bury it. The whole point of MP is to show that the “laws” that GB discovered can be understood in more fundamental MP terms. IMO, it has been less appreciated than it should be how much MP has succeeded in making good on this ambition. So, in this post (and another I will put up later on) I will try to show how far MP has come in making good on its ambitions.

A caveat: if you find GB (and its cousins) to be hopelessly wrongheaded, then you will find a theory that derives its results also hopelessly wrongheaded. If you are one of these, then MP will have, at most, aesthetic interest. If you, like me, take GB to be more or less correct, then MP’s aesthetic virtues will combine with GB’s empirical panache to create a very powerful intellectual rush. It will even create the impression that MP is very much on the right track.

The Merge Hypothesis: Explaining some core features of FL/UG

Here is a list of some characteristic features of FL/UG and its GLs:

(1)  a.   Hierarchical recursion
b.   Displacement (aka, movement)
c.     Gs generate natural formats for semantic interpretation
d.     Reconstruction effects
e.     Movement targets c-commanding positions
f.      No lowering rules
g.     Strict cyclicity
h.     G rules are structure dependent
i.      Antecedents c-command their anaphors
j.      Anaphors never c-command their antecedents (i.e. Principle C effects and Strong Cross Over Effects)
k.     XPs move, X’s don’t, X0s might
l.      Control targets subjects of “defective” (i.e. tns or agreement deficiency) clauses
m.   Control respects the Principle of Minimal Distance
n.     Case and agreement are X0-YP dependencies
o.     Reflexivization and Pronominalization are in complementary distribution
p.    Selection/subcategorization are very local head-head relations
q.     Gs treat arguments and adjuncts differently, with the former less “constrained” than the latter

Note, I am not saying that this exhausts the properties of FL/UG, nor am I saying that all LINGers agree with all of these accurately describe FL/UG.[1] What I am saying is that (1a-q) identify empirically robust(ish) properties of FL/UG and the generative procedures its GLs allow. Put another way, I am claiming (i) that certain facts about human GLs (e.g. that they have hierarchical recursion and movement and binding under c-command and display principle C effects and obligatory control effects, etc.) are empirically well-grounded and (ii) that it is appropriate to ask why FL/UG allows for GLs with these properties and not others. If you buy this, then welcome to the Minimalist Program (MP).

I would go further; not only are the assumptions in (i) reasonable and the question in (ii) appropriate, MP has provided some answers to the question in (ii). One well-known approach to (1a-h), the Merge Hypothesis (MH), unifies all these properties, deriving them from the core generative mechanism Merge. Or more particularly, MH postulates that FL/UG contains a very simple operation (aka, Merge) that suffices to generate unbounded hierarchical structures (1a) and that these Merge generated hierarchical structures will also have the seven additional properties (1b-h). Let’s examine the features of this simple operation and see how it manages to derive these eight properties?

Unbounded hierarchy implies a recursive procedure.[2] MH explains this by postulating a simple operation (“Merge”) that generates the requisite unbounded hierarchical structures. Merge consists of a very simple recursive specification of Syntactic Objects (SO) coupled with the assumption that complex SOs are sets.

(2)  a. If a is a lexical item then a is a SO[3]
b. If a is an SO and b is an SO the Merge(a,b) is an SO

(3)  For a, b, SOs, Merge(a,b)à {a,b}

The inductive step (2b) allows Merge to apply to its own outputs and thus licenses unboundedly “deep” SOs with sets contained within sets contained within sets… The Merge Hypothesis is that the “simplest” conception of this combinatoric operation (the minimum required to generate unbounded hierarchically organized objects) suffices to explain why FL/UG has many of the other properties listed in (1).

In what way is Merge the “simplest” specification of unbounded hierarchy? The operation has three key features: (i) it directly and uniquely targets hierarchy (i.e. the basic complex objects are sets (which are unordered), not strings), (ii) it in no way changes the atomic objects combined in combining them (Inclusiveness), and (iii) it in no way changes the complex objects combined in combining them (Extension). Inclusiveness and Extension together constitute the “No Tampering Condition” (NTC). Thus, Merge recursively builds hierarchy (and only hierarchy) without “tampering” with the inputs in any way save combining them in a very simple way (i.e. just hierarchy no linear information).[4] The key theoretical observation is that if FL/UG has Merge as its primary generative mechanism,[5] then it delivers GLs with properties (1a-h). And if this is right, it provides a proof of concept that it is not premature to ask why FL/UG is structured as it is. In other words, this would be a very nice result given the stated aims of MP. Let’s see how Merge so conceived derives (1a-h).

It should be clear that Gs with Merge can generate unbounded hierarchical dependencies. Given a lexicon containing a finite list of atoms a,b,g,d,… we can, using the definitions in (2) and (3) form structures like (4) (try it!).

            (4)       a. {a, {b, {g, d}}}
                        b.  {{a, b}, {g, d}}
                        c.  {{{a, b}, g}, d}

And given the recursive nature of the operation, we can keep on going ad libitum. So Merge suffices to generate an unbounded number of hierarchically organized syntactic objects.

Merge can also generate structures that model displacement (i.e. movement dependencies). Movement rules code the fact that a single expression can enjoy multiple relations within a structure (e.g. it can be both a complement of a predicate and the subject of a sentence).[6] Merge allows for the derivation of structures that have this property. And this is a very good thing given that we know (due to over 60 years of work in Generative Grammar) that displacement is a key feature of human GLs.

Here’s how Merge does this. Given a structure like (5a) consider how (2) and (3) yield the movement structure (5b). Observe that in (5b), b occurs twice. This can be understood as coding a movement dependency, b being both sister of the SO a and sister of the derived SO {g, {l, {a, b}}}. The derivation is in (6).

(5)       a. {g, {l, {a, b}}}
      b. {b, {g, {l, {a, b}}}}

(6) The SO {g, {l, {a, b}}} and the SO b (within {g, {l, {a, b}}}) merge to      from {b, {g, {l, {a, b}}}}

Note that this derivation assumes that once an SO always an SO. Thus, Merging an SO a to form part of a complex SO b that contains a does not change (tamper with) a’s status as an SO. Because complex SOs are composed of SOs Merge can target a subpart of an SO for further Merging. Thus, NTC allows Merge to generate structures with the properties of movement; structures where a SO is a member of two different “sets.”

Let me emphasize an important point: the key feature that allows Merge to generate movement dependencies (viz. the “once an SO always an SO” assumption) follows from the assumption that Merge does nothing more than take SOs and form them into a unit. It otherwise leaves the combined objects alone. Thus, if some expression is an SO before being merged with another SO then it will retain this property after being Merged given that Merge in no way changes the expressions but for combining them. NTC (specifically the Inclusiveness and Extension Conditions) leaves all properties of the combining expressions intact. So, if a has some property before being combined with b (e.g. being an SO), it will have this property after it is combined with b. As being an SO is a property of an expression, Merging it will not change this and so Merge thus legitimately combine a subpart of an SO to its container.

Before pressing on, a comment: unifying movement and phrase building is an MP innovation. Earlier theories of grammar (and early minimalist theories) treated phrasal dependencies and movement dependencies as the products of entirely different kinds of rules (e.g. phrase structure rules vs transformations/Merge vs Copy+Merge). Merge unifies these two kinds of dependencies and treats them as different outputs of a single operation. As such, the fact that FL yields Gs that contain both unbounded hierarchy and displacement operations is unsurprising. Hierarchy and displacement are flips sides of the same combinatoric coin. Thus, if Merge is the core combinatoric operation FL makes available, then MH explains why FL/UG constructs GLs that have both (1a) and (1b) as characteristic features.

Let’s continue. As should be clear, Merge generated structures like those in (5) and (6) also provides all we need to code the two basic types of semantic dependencies: predicate-argument structures (i.e. thematic dependency) and scope structure. Let me be a bit clearer. The two basic applications of move are those that take two separate SOs and combine them and those that take two SOs with one contained in the other and combines them. The former, E-Merge, is fit for the representation of predicate-argument (aka, thematic structure). The latter, I-Merge, provides an adequate grammatical format for representing operator/variable (i.e. scope) dependencies. There is ample evidence that Gs code for these two kinds of semantic information in simple constructions like Wh-questions. Thus, it is an argument in its favor that Merge as defined in (2) and (3) provides a syntactic format for both. An argument saturates the predicate it E-merges with and scopes over the SO it I-merges with. If this is correct, then Merge provides structure appropriate to explain (1c).

And also (1d). A standard account of Reconstruction Effects (RE) involves allowing a moved expression to function as if it still occupied the position from which it moved. This as-if is redeemed theoretically if the movement site contains a copy of the moved expression. Why does a displaced expression semantically comport itself as if it is in its base position? Because a copy of the moved expression is in the base position. Or, to put this another way, a copy theory of movement would go a long way towards providing the technical wherewithal to account for the possibility of REs. But Merge based accounts of movement like the one above embody a copy theory of movement. Look at (5b): b is in two positions in virtue of being I-merged with its container. Thus, b is a member of the lowest set and the highest. Reconstruction amounts to choosing which “copy” to interpret semantically and phonetically.[7] Reducing movement to I-merge explains why movement should allow REs.[8]

Furthermore, having this option follows from a key assumption concerning Merge. Recall that it eschews tampering. In other words, if movement is a species of Merge then no-tampering requires coding movement with “copies.” To see this contrast how movement is treated in GB.

Within GB, if a moves from its base position to some higher position a trace is left in the launch site. Thus, a GB version of (5b) would look something like (7):

            (7) {b1, {g, {l, {a, t1}}}}

Two features are noteworthy; (i) in place of a copy in the launch site we find a trace and (ii) that trace is co-indexed with the moved expression b.[9] These features are built into the GB understanding of a movement rule. Understood from a Merge perspective, this GB conception is doubly suspect for it violates the Inclusiveness Condition clause of the NTC twice over. It replaces a copy with a trace and it adds indices to the derived structure. Empirically, it also mystifies REs. Traces have no contents. That’s what makes them traces (see note 13). Why are they then able to act as if they did have them? To accommodate such effects GB adds a layer of theory specific to REs (e.g. it invokes reconstruction rules to undo the effects of movement understood in trace theoretic terms). Having copies in place of traces simplifies matters and explains how REs are possible. Furthermore, if movement is a species of Merge (i.e. I-merge) then SOs like (7) are not generable at all as they violate NTC. More specifically, the only kosher way to code movement and obey the NTC is via copies. So, the only way to code movement given a simple conception of syntactic combination like Merge (i.e. one that embodies no tampering) results in a copy theory of movement that serves to rationalize REs without a theoretically bespoke theory of reconstruction. Not bad![10]

So Merge delivers properties (1a-d), and the gifts just keep on coming. It also serves up (6e,f,g) as consequences. This time let’s look at the Extension Condition (EC) codicil to the NTC. EC requires that that inputs to Merge be preserved in the outputs to Merge (any other result would “change” one of the inputs). Thus, if an SO is input to the operation it will be a unit/set in the output as well because Merge does no more than create linguistic units from the inputs. Thus, whatever is a constituent in the input appears as a constituent with the same properties in the output.  This implies (i) that all I-merge is to a c-commanding position, (ii) that lowering rules cannot exist, and (iii) that derivations are strictly cyclic.[11] The conditions that movements be always upwards to c-commanding positions and strictly cyclic thus follows trivially from this simple specification of Merge (i.e. Merge with NTC understood as embodying EC).

An illustration will help clarify this. NTC prohibits deriving structure (8b) from (8a). Here we Merge g with a. The output of this instance of Merge obliterates the fact that {a,b} had been a unit/constituent in (8a), the input to Merge. EC prohibits this. It effectively restricts I-Merge to the root. So restricted, (8b) is not a licit instance of I-Merge (note that {a,b} is not a unit in the output. Nor is (8c) (note that {{a,b},{g,d}} is not a unit in the output). Nor is a derivation that violates the strict cycle (as in (8d)). Only (8e) is a grammatically licit Merge derivation for here all the inputs to the derivation (i.e. g and {{a,b},{g,d}}) are also units in the output of the derivation (i.e. thus the inputs have been preserved (remain unchanged) in the output). Yes a new relation has been added, but no previous ones have been destroyed (i.e. the derivation is info-preserving (viz. monotonic). Repeat the slogan: once an SO always an SO). In deriving (8b-c) one of the inputs (viz. {{a,b},{g,d}}) is no longer a unit in the output and so NTC/EC has been violated.

(8)       a. {{a,b},{g,d}}
b. {{{g,a},b}, {g,d}}
                        c. {{g,{a,b}},{g,d}}
                        d.  {{a, b}, {d, {g,d}}
                        d. {g,{{a,b},{g,d}}}

 In sum, if movement is I-merge subject to NTC then all movement will necessarily be to c-commanding positions, upwards, and strictly cyclic.

It is worth noting that these three features are not particularly recondite properties of FL/UG and find a place in most GG accounts of movement. This makes their seamless derivation within a Merge based account particularly interesting.

Last, we can derive the fact that the rules of grammar are structure dependent ((1h) above), an oft-noted feature of syntactic operations.[12] Why should this be so? Well, if Merge is the sole syntactic operation and then non-structure dependent operations are very hard (impossible?) to state. Why? Because the products of Merge are sets and sets impose no linear requirements on their elements. If we understand a derivation to be a mapping of phrase markers into phrase markers and we understand phrase markers to effectively be sets (i.e. to only specify hierarchical relations) then it is no surprise that rules that leverage linear left-right properties of a string cannot be exploited. They don’t exist for phrase markers eschew this sort of information and thus operations that exploit left/right (i.e. string based) information cannot be defined.  So, why are rules of G structure dependent? Because this is the only structural information that Merge based Gs represent. So, if the basic combinatoric operation that FL/UG allows is Merge, then FL/UGs restriction to structure dependent operations is unsurprising.

This is a good place to pause for a temporary summary: Research in GG over the last 60 years has uncovered several plausible design features of FL/UG. (1a-h) summarizes some uncontroversial examples. All of these properties of FL/UG can be unified if we assume that Merge as outlined in (2) and (3) is the basic combination operation that FL/UG affords. Put simply, the Merge Hypothesis has (1a-h) as consequences.

Let me say the same thing more tendentiously. All agree that a basic feature of FL/UG is that allows for Gs with unbounded hierarchy. A very simple inductive procedure sufficient for specifying this property (1a), also entails many other features of FL/UG (1b-h). What makes this specification simple is that it directly targets hierarchy and requires that the computation be strongly monotonic (embody the NTC). Thus we can explain the fact that FL/UG has these properties by assuming that it embodies a very simple (arguably, the simplest) version of a procedure that any empirically adequate theory of FL/UG would have to embody. Or, given that FL/UG allows for unbounded hierarchical recursion (a non-controversial fact given the fact of Linguistic Productivity), the simplest (or at least, very simple) version of the requisite procedure brings in its train displacement, an adequate format for semantic interpretation, Reconstruction Effects, movement rules that target c-commanding positions, eschew lowering and are strictly cyclic, and G operations that are structure dependent. Thus, if the Merge Hypothesis is true (i.e. if FL/UG has Merge as the basic syntactic operation), it explains why FL/UG has this bushel of properties. In other words, the Merge Hypothesis, provides a plausible first step in answering the basic MP question: why does FL/UG have the properties it has?

Moreover, it is morally certain that something like Merge will be part of any theory of FL/UG precisely because it is so very simple. It is always possible to add bells and whistles to the G rules FL/UG makes available. But any theory hoping to be empirically adequate will contain at least this much structure. After all, what do (2) and (3) specify? They specify a recursive procedure for building hierarchical structures that does nothing but build such structures. Given the fact of Linguistic Productivity and Linguistic Promiscuity any theory of FL/UG will contain at least this much. If it does not contain much more than this much, then (1a-h) results. Not bad.


[1] For example, fans of Dependent Case Theory will reject (1n).
[2] Recall, that LP implies recursion and linguistics has discovered ample evidence that GLs can generate structures of arbitrary depth.
[3] The term lexical item denotes the atoms that are not themselves products of Merge. These roughly correspond to the notion morpheme or word, though these notions are themselves terms of art and it is possible that the naïve notions only roughly corresponds to the technical ones. Every theory of syntax postulates the existence of such atoms. Thus, what is debatable is not their existence but their features.
[4] In my opinion, this line of argument does not require that Merge be the “simplest” possible operation. It suffices that it be natural and simple. The conception of Merge in (2) and (3) meets this threshold.
[5] In the best of all worlds, the sole generative procedure.
[6] A phrase marker is just a list of relations that the combined atoms enjoy. Derivations that map phrase markers into phrase markers allow an expression to enjoy different relations coded in the various relations it enjoys in the varying phrase markers.
[7] Copy is simply a descriptive term here. A more technically accurate variant is “occurrence.” b occurs twice in (5b). The logic, however, does not change.
[8] A full theory of REs would articulate the principles behind choosing which copies to interpret. See Sportiche (forthcoming) for an interesting substantive proposal.
[9] Traces within GB are indexed contentless categories: [1 ec]. 
[10] We could go further: Merge based theory cannot have objects like traces. Traces live on the distinction between Phrase Structure Rules and lexical insertion operations. They are effectively the phrase structure scaffolding without the lexical insertion. But, Merge makes no distinction between structure building and lexical insertion (i.e. between slots and contents). As such, if traces exist, they must be lexical primitives rather than syntactically derived formatives. This would be a very weird conception of traces, inconsistent with the GB rendering in note 12. The same, incidentally, goes for PRO, which we will talk about later on. The upshot: not only would traces violate No tampering, they are indefinable given the “bare phrase structure” nature of movement understood as I-merge.
[11] The first conjunct only holds if there is no inter-arboreal/sidewards movement. For now, let’s assume this to be correct.
[12] For a recent review and defense of the claim see Berwick et. al.