Comments

Showing posts with label phases. Show all posts
Showing posts with label phases. Show all posts

Tuesday, February 16, 2016

More on subjacency

Peter Svenonius (once again) asks the right question and Omer (once again) has interesting things to say about it (see here). Take a look. Here is my take on the issues. Please chime in with yours.

I agree with Omer that there may not be much of a consensus right now about how to deal with successive cyclicity in detail. However, so far as I can tell, there is general agreement that it has something to do with the PIC (this is the current analogue of the subacency principle, Bounding nodes and  the domains they create). 

As you know, there are two extant versions of the PIC, the more favored one being the one wherein a complement of a phase head is rendered inaccessible at the NEXT phase (usually when the next phase head is accessed). This comes close to coding the old idea of subjacent domain (as was early observed one has access to the domain one is in and the next one (i.e. no counting)). The strong version of a phase is less in favor, but it has some charms for it would force something like a "no edge skipping" requirement on the grammar (e.g. the old Rizzi idea that one could "skip" the most immediate Comp would be ruled out). There are purported arguments against this stronger version, but they never struck me as dispositive (and there are Legate style arguments against it). At any rate, there is consensus that phases should derive cyclicity.

How closely is this tied to features, uninterpretable or otherwise? Logically speaking, not that closely so far as I can tell. The issue of features is tied to whether movement is optional or obligatory. If Greed drives movement or uninterpretability then C features might be needed. Yes if Greed is strong and no if something like uninterpretability suffices to drive one to the phase edge as a last resort (Phase balance or Boskovic's take on the same idea). There is some evidence that intermediate Cs can have features, as we know. The generalization to all languages is a standard GG move. So, the idea is not empirically nuts. Of course, WHY this should be true is unclear in the absence of something like Greed, but then maybe this is an argument for a strong version of Greed.

This goes against the current fashion. It seems that nowadays movement is free again (free at last, free at last, thank the lord, free at last!). But then there is no requirement that there be intermediate features to drive movement, nor that there be uninterpretable features on WH to force it to move. The WH moves or it does not. If it does, then it must move to the intermediate C for PIC reasons. If it doesn't then no convergence. More specifically, what one needs are language specific requirements that force a given G to have a WH up top overtly in some languages (something like the old strong feature) or some kind of Rizzi Criterion that is fulfilled in G variable ways. This seems generally assumed in current technology, so no biggie here.

There is one last idea that has been tied to successive cyclicity: Chomsky's current idea about labels. Oddly, for Chomsky, the fact that there are languages where there appears to be agreement in non WH Cs with a moving WH is a big problem. Agreement should obviate further movement. Of course one can get fancier here with different features having different effects on labeling (and so movement), but this begins to hand code in the property we want explained (not a good thing to do).

That's the way things look from where I sit. So, there are several ways of getting edge to edge movement all involving the PIC in some fashion and thereby recoding the old subjacency criterion. I want to emphasize this: this is not a new explanation but a recoding of the old one (not that this is a bad thing).

Two last points: what is less clear to me is how this all hooks up with islands. Chomsky, it seems to me, is reluctant to take islands as G-real phenomena. He seems inclined to take the view he once criticized, viz: that islands are performance residues of complexity. I am skeptical myself, but it is a logical possibility. The Sprouse stuff has convinced me that it is likely false. 

This leaves the question of how to code Islands in phases? That's easy (as anyone who has tried will attest). The problem is that the coding follows from nothing (why are D edges different from C edges? why weak PIC rather than stropping? Why transfer when next pause head chosen rather than next phase completed? Why C and D and v as phases? Why week vs strong phases?). In fact, the coding just recapitulates the machinery in classical Subjacency theory. Or, Minimalism has not given us any insight into the details of subjacency as of this date. So, islands stand as having no good deep explanation beyond the one that Chomsky already provided for Subjacency that I quoted in the body of the earlier post.

Second: we really would love to tie island effects with ECP effects as Barriers and Cinque-Rizzi tried to do. Why? Because the domains for bounding and ECP are so damn similar. It would be really odd (IMO too odd to be tolerable) were these driven by different mechanisms given that their domains are virtually identical. So, we need to find a way of finally addressing ECP questions within MP. In particular we need to find a way of unifying them in ways more conceptually acceptable than the Barriers/Lasnik-Saito theory did.


So island effects are currently no better understood within minimalist theory than they were within GB. The GB story can be smoothly translated into technologically acceptable minimalist terms, but doing so provides no insight. Moreover, some parts of the old theory, the ECP part dealing with adjuncts vs arguments and their differing locality conditions, really has not good minimalist counterparts (does anyone really thing gamma marking is part of FL/UG?). That's how I see things. You? 

Sunday, June 22, 2014

Comments on lecture 2; part deux

In the first post (here), I discussed Chomsky’s version of Merge and the logic behind it.  The main idea is that Merge, the conceptually simplest conception of recursion, has just the properties to explain why NL Gs generate structures with unbounded hierarchical structure, why NLs allow displacement, show reconstruction effects, and why rules of G are structure dependent. Not bad for any story. Really good (especially for DP concerns) if we get all of this from a very simple (nay, simplest) conception. In what follows I turn to a discussion of the last three properties Chomsky identified and see how he aims to account for them. I repeat them here for convenience.

(v)           its operations apply cyclically
(vi)          it can have lots of morphology
(vii)        in externalization only a single “copy” is pronounced

In contrast to the first four properties, the last three do not follow simply from the properties of the conceptually “simplest” combination operation. Rather Chomsky argues that they reflect principles of computational efficiency. Let’s see how.

With respect to (vii), Chomsky assumes that externalization (i.e. “vocalizing” the structures) is computationally costly. In other words, actually saying the structures out loud is hard. How costly? Well, it must be more costly than copy deletion at Transfer is. Here’s why. Given the copy theory as a consequence of Merge, FL must contain a procedure to choose which copy/occurrence is pronounced (note: this is not a conceptual observation but an inference based on the fact that typically only one copy is pronounced). This decision/choice, I assume, requires some computation. I further assume that choosing which copies/occurrences to externalize requires some computation that would not be required were all copies/occurrences pronounced. Chomsky’s assumption is that the cost of choosing is less than the cost of externalizing.  Thus, FL’s choice lowers overall computational cost.

Furthermore, we must also assume that the cost of pronunciation also exceeds the computational cost of being misunderstood for otherwise it would make sense for FL to facilitate parsing by pronouncing all the copies, or at least those that would facilitate a hearer’s parsing of our sentences. None of these assumptions are self-evidently true or false. Plus, the supposition that copy deletion is more computationally efficient than pronouncing them would be does not follow simply from considerations of conceptual simplicity, at least as far as I can tell. It involves substantive assumptions about actual computational costs, for which, so far as I can tell, we have little independent evidence.

One more point: If copy deletion exists in Transfer to the CI interface (as Chomsky argued in his original 1993 paper and that underlies standard accounts of reconstruction effects and that so far as I know is still part of current theory) then in the normal case only a single copy/occurrence makes it to either interface, though which copy is interpreted at CI can be different form the copy spoken at AP (and this is typically how displacement is theoretically described). But if this is correct, then it suggests that Chomsky’s argument here might need some rethinking. Why? If deletion is part of Transfer to CI then copy deletion cannot be simply a fact about the computational cost of externalization, as it applies to the mapping of linguistic objects to the internal thought system as well. It seems that copies per se are the problem, not just copies that must be pronounced.

Before moving on to (v) and (vi) it is worth pausing to note that Chomsky’s discussion here reverberates with pretty standard conceptions of computational efficiency (viz. he is making claims about how hard it is to do something). This moves away from the purely conceptual matters that motivated the discussion of the first four features of FL. There is a very interesting hypothesis that might link the two: that the simplest computational operation will necessarily be embedded in a computationally efficient system. This is along the lines of how I interpreted the SMT in earlier posts (linked to in the first part of this post).  However, whether you think this is feasible, it appears, at least to me, that there are two different kinds of arguments being deployed to SMT ends, a purely conceptual one and a more conventional “resource” argument.

Ok, let’s return to (v) and (vi). Chomsky suggests that considerations of computational efficiency also account for these properties of. In particular, they follow from something like the strict cycle as embodied in phase theory.  So the question is what’s the relation between the strict cycle and efficient computation?

Chomsky supposes that the strict cycle, or something like it, is what we would expect from a computationally well-designed system. There are times that (to me) Chomsky sounds like he seems to be assuming that the conceptually simplest system will necessarily be computationally efficient.[1] I don’t see why. In particular, if I understand the lecture correctly, Chomsky is suggesting that the link between conceptual simplicity and computational efficiency should follow as a matter of natural law. Even if correct, it is clear that this line of reasoning goes considerably beyond considerations of conceptual simplicity. What I mean is that even if one grants that the simplest computational operation will be something like Merge, it does not follow that the simplest system that includes Merge will also incorporate the strict cycle.  Phases then, (Chomsky’s mechanism for realizing the strict cycle) are motivated not on grounds of conceptual simplicity alone but on grounds of efficiency (i.e. a well/optimally designed system will incorporate something like the strict cycle). So far as I can tell Chomsky does not explain the relation (if any) between conceptual simplicity and computationally efficiency, though to be fair, I may be over-interpreting his intent here.

This said how does the strict cycle bear on computational efficiency? It allows computational decisions to be made locally and incrementally. This is a generically nice feature for computational systems to have for it simplifies computations.[2] Chomsky notes that it also simplifies the process of distinguishing two selections of the same expression from the lexicon vs two occurrences of the same expression. How does it simplify it? By making the decision a bounded one. Distinguishing them, he claims, requires recalling whether a given occurrence/copy is a product of E- or I-Merge. If such decisions are made strict cyclically (at every phase) then phases reduce memory demand: because phases are bounded, you need not retain information in memory regarding the provenance of a valued occurrence beyond the phase where an expression’s features are valued.[3] So phases ease the memory burdens that computations impose. Let me note again without further comment, that if this is indeed a motivation for phases, then it presupposes some conception of performance for only in this kind of context do resource issues (viz. memory concerns) arise. God has no need for bounding computation.

Now I have a confession to make.  I could not come up with a concrete example where this logic is realized involving DP copies, given standard views.  It’s easy enough to come up with a relevant case if e.g. reflexivization is a product of movement.[4] If reflexives involve A-chains with two thematically marked “links” then we need to distinguish copies from originals (e.g. Everyone loves himself differs from everyone loves everyone in that the first involves one selection of everyone from the lexicon (and so one chain with two occurrences of everyone) while the second involves two selections of everyone from the lexicon and so two different chains). However, if you don’t assume this, I personally had a hard time finding an example of what’s worrying Chomsky, at least with copies. This might mean that Chomsky is finally coming to his senses and appreciating the beauty of movement theories of Control and Binding OR it might mean that I am a bear of little brain and just couldn’t come up with a relevant case. I know which option I would bet on, even given my little brain, and it’s not the first. So, anyone with a nice illustration is invited to put it in the comments section or send it to me and I will post it. Thanks.

It is not hard to come up with cases that do not involve DPs, but the problem then is not distinguishing copies from originals. Take the standard case of Subject-Predicate agreement for example. Here the unvalued features of T are valued by those of the inherently valued features of the subject DP.  Once valued, the features on T and D are indistinguishable qua features. However, there is assumed to be an important difference between the two, one relevant to the interpretation at the CI interface. Those on D are meaning relevant but those on T are uninterpretable. What, after all, could it mean to say that the past tense is first person and plural?[5] If one assumes that all features at the interfaces must be interpretable at those interfaces if they make it there, then the valued features on T must disappear at Transfer to CI. But if (by assumption) they are indistinguishable from the interpretable ones on D, the computational system must remember how the features got onto T (i.e. by valuation rather or inherently). The ones that get there by valuation in the grammar must be removed or the derivation will not converge. Thus, Gs need to know how features get onto the expressions they sit on and it would be very nice memory-wise if this was a bounded decision.

Before moving on, it’s worth noting that even this version of the argument is hardly straightforward. It assumes that phi-features on T are not-interpretable and that these cause derivations to crash (rather, then, for example, converge as gibberish) (also see note 5). It also requires that deletion not be optional, otherwise there would be derivations where all the good features remained on all of the right objects and all of the uninterpretable ones freely deleted. Nor does it allow Transfer (which, after all, straddles the syntax and CI) to peak at the meaning of T during Transfer, thereby determining which features are interpretable on which items and so which should be deleted and which retained. Note that such a peak-a-boo decision to delete during Transfer would be very local, relying just on the meaning of T and the meaning of phi-features. Were this possible, we could delay Transfer indefinitely. So, to make Chomsky’s argument we must assume that Transfer is completely “blind” to the interpretation of the syntactic objects at every point in the syntactic computation including the one that interfaces with CI. This amounts to a very strong version of the autonomy of syntax thesis; one in which no part of the syntax, even the rules that directly interface with the interpretive interfaces, can see any information that the interfaces contain.[6]

Let’s return to the main point. Must the simplest system imaginable be computationally efficient? It’s not clear. One might imagine that the conceptually “simplest” system would not worry about computational efficiency at all (damn memory considerations!). The simplest system might just do whatever it can and produce whatever structured products it can without complicating FL with considerations of resource demands like memory burdens. True, this might render some products of FL unusable or hard to use (and so we would probably perceive their use as perceive them as unacceptable) but then we just wouldn’t use them (sort of like what we say about self-embedded clauses).  So, for example, we would tend not to use sentences with multiple occurrences of the same expressions where this made life computationally difficult (e.g. you would not talk about two Norberts in the same sentence). Or without phases we might leave to context the determination of whether an expression is a copy or a lexical primitive or we might allow Transfer to see if features on an expression were kosher or not. At any rate, it seems to me that all of these options are as conceptually “simple” as adding phases to FL unless, or course, phases come for free as a matter of “natural law.”  I confess to being skeptical about this supposition. Phases come with a lot of conceptual baggage, which I personally find quite cumbersome (reminds me of Barriers actually, not one of the aesthetic high points in GG (ugh!)). That said, let’s accept that the “simplest” theory comes with phases. 

As Chomsky notes, phases themselves come have complex properties.  For example, phases bring with them a novel operation, feature lowering, which now must be added to the inventory of FL operations. However, feature lowering does not seem to be either a conceptually simple or cognitively/computationally generic kind of operation. Indeed, it seems (at least to me) quite linguistically parochial. This, of course, is not a good thing if one’s sights are set on answering Darwin’s problem.  If so, phases don’t fit snugly with the SMT. This does not mean there are none. It just means that they complicate matters conceptually and pull against Chomsky’s first conceptual argument wrt Merge.

Again, let’s put this all aside and assume that strict cyclicity is a desirable property to have and that phases are an optimal way of realizing this. Chomsky then asks how we identify phases? He argues that we can identify phases by their heads as phase heads are where unvalued features live. Thus a phase is the minimal domain of a phase head with unvalued features.[7] A possible virtue of this way of looking at things is that it might provide a way of explaining why languages contain so much morphology. They are the adventitious by-products for identifying the units/domain of the optimal computational system.  Chomsky notes that what he means by morphology is abstract (a la Vergnaud), so a little more has to be said, especially given that externalization is costly, but it’s an idea in an area where we don’t have many (see here).[8]

One remark: on this reconstruction of Chomsky’s arguments, unvalued features play a very big role. They identify phases, which implement strict cyclicity and are the source of overt morphology.  I confess to being wary here. Chomsky originally introduced unvalued features to replace uninterpretable ones. Now he assumes that features are both +/- valued and +/- interpretable. As unvalued features are always uninterpretatble, this seems like an unwanted redundancy in the feature system.  At any rate, as Chomsky notes, uninterpretable features really do look sort of strange in a perfect system. Why have them only to get rid of them?  Chomsky’s big idea is that they exist to make FL computationally efficient. Color me very unconvinced.

So this is the main lay of the land. I should mention that, as others have pointed out (especially Dennis O), part of Chomsky’s SMT argument here (i.e. the one linked to conceptual simplicity concerns) is different from the interpretation of the SMT that I advanced in other posts (here, here, here).  Thus, my version is definitely NOT the one that Chomsky elaborates when considering these. However, there is a clear second strand dealing with pretty standard efficiency concerns, and here my speculations and his might find some common ground. That said, Chomsky’s proposals rest heavily on certain assumptions about conceptual simplicity, and of a very strong kind. In particular, Chomsky’s argument rests on a very aggressive use of Occam’s razor.  Here’s what I mean. The argument he offers is not that we should adopt Merge because all other notions are too complex to be biologically plausible units of genetic novelty. Rather, he argues that in the absence of information to the contrary, Occamite considerations should rule: choose the simplest (not just a simple) starting point and see where you get. Given that we don’t know much about how operations that describe the phenotype (the computational properties of FL) relate to the underlying biological substrate that is the thing that actually evolved, it is not clear (at least to me) how to weight such strong Occamite considerations. They are not without power, but, to me at least, we don’t really know how to assess whether all things are indeed equal and how seriously to weight this very strong demand for simplicity

Let me end by fleshing this out a bit.  I confess to not being moved by Chomsky’s conceptual simplicity arguments. There are lots of simple starting points (even if some may be simpler than others). Ordered pairs are not that much more conceptually complex than sets. Symmetric operations are not obviously simpler than asymmetric ones, especially given that it appears that syntax abhors symmetry (see Moro and Chomsky). So, the general starting point that we need to start with the conceptually simplest conception of “combination” and that this means an operation that creates sets of expressions seems based on weak considerations. IMO, we should be looking for basic concepts that are simple enough to address DP (and there may be many) and evaluate them in terms of how well they succeed in unifying the various apparently disparate properties of FL. Chomsky does some of this here, and it’s great. But we should not stop here. Let me given an example.

One of the properties that modern minimalist theory has had trouble accounting for is the fact that the unit of syntactic movement/interpretation/deletion is the phrase. We may move heads, but we typically move/delete phrases. Why? Right now standard minimalist accounts have no explanation on hand. We occasionally hear about “pied piping” but more as an exercise in hand waving than in explanation. Now, this feature of FL is not exactly difficult to find in NL Gs. That constituency matters is one of the obvious facts about how displacement/deletion/binding operates. There is a simple story about this that labels and headedness can be used to deliver.[9] If this means that we need a slightly less conceptually simple starting point than sets, then so be it.

More generally: the problem that motivates the minimalist program is DP. To address DP we need to factor out most of the linguistic specific structure of FL and attribute it to more cognitively generic operations (or/and, if Chomsky is right, natural laws).  What’s simple in a DP context is not what is conceptually most basic, but what is simple given what our ancestors had available cognitively about 100k years ago. We need a simple addition to this, not something that is conceptually simple tout court.[10]  In this context it’s not clear to me that adding a set construction operation (which is what Merge amounts to) is the simplest evolutionary alternative. Imagine, for example, that our forbearers already had an itterative concatenation operation.[11]  Might not some addition to this be just as simple as adding Merge in its entirety? Or imagine that our ancestors could combine lexical atoms together into arbitrarily big unstructured sets, might not an addition that allowed that operation to yield structured sets be just as simple in the DP context as adding Merge? Indeed, it might be simpler depending in what was cognitively available in the mental life of our ancestors.  And once we are at it, how “simple” is an operation that forms arbitrary sets from atoms and other sets?  Sets may be simple objects with just the properties we need, but I am not sure that operations that construct them are particularly simple.[12]

Ok, let me end this much too long second post. And moreover, let me end on a very positive note. In the second lecture Chomsky does what we all should be doing when we are doing minimalist syntax. He is interested in finding simple computational systems that derive the basic properties of FL. He concentrates on some very interesting key features: unbounded hierarchy, displacement, reconstruction, etc. and makes concrete proposals (i.e. he offers a minimalist theory) that seem plausible. Whether he is right in detail is less important IMO than that his ambitions and methods are worth copying. He identifies non-trivial properties of FL that GG has discovered over the last 60 years and he tries to explain why they should exist.  This is exactly the right kind of thing MPers should be doing. Is he right? Well, let’s just say that I don’t entirely agree with him (yet!). Does lecture 2 provide a nice example of what MP research should look like. You bet. It identifies real deep properties of FL and sees how to derive them from more general principles and operations. If we are ever to solve Darwin’s problem, we will need simple systems that do just what Chomsky is proposing. 






[1] Note, we want the necessarily here. That it is both simple and efficient does not explain why it need be efficient if simple.
[2] It is also a necessary condition for incrementality in the use systems (e.g. parsing), as Bill Idsardi pointed out to me.  I know that the SMT does not care about use systems according to some (Dennis and William this is a shout-out to you), but this is a curious and interesting fact nonetheless.  Moreover, if I am right that the last three properties do not follow (at least not obviously) from conceptual considerations, it seems that Chomsky might be pursuing a dual route strategy for explaining the properties of FL.
[3] Note that this assumes that there is no syntactic difference between inherent features and features valued in the course of the derivation.
[4] And even this requires a special version of the theory, one like Idsardi and Lidz’s rather than Zwart’s.
[5] However, if v raised to T before Transfer then one might try and link these features to the thematic argument that v licenses. And then it might make lots of sense to say that phi-features are interpretable on T. They would say that the variable of the predicate bound by the subject must have such and such an interpretation. This information might be redundant, but it is not obviously uninterpretable.
[6] The ‘autonomy of syntax’ thesis refers to more than one claim. The simplest one is that syntactic primitives/operations are not reducible to phonetic or semantic ones. This is not  the version adverted to above. This is a more specific version of the thesis; one that requires a complete separation between syntactic and semantic information in the course of a derivation. Note, that the idea that one can add EPP/edge features only if it affects interpretation (the Reinhart-Fox view that Chomsky has at times endorsed) violates this strong version of the autonomy thesis.
[7] Note, we still need to define ‘domain’ here.
[8] Note, incidentally, that Chomsky assumes both that features are +/- valued and that they are +/- interpretable. At one time, the former was considered a substitute for the latter. Now, they are both theoretically required, it seems. As -valued features seem to always be –interpretatble, this seems like an unwanted redundancy. 
[9] I provide a story here based on labels and minimality.
[10] A question: we can define ordered pairs set theoretically. I assume the argument against labels is that ordered sets are conceptually more complex than unordered sets. So {a,b} is conceptually simpler than {a,{a,b}}.  If this is the argument, it is very very subtle. I find it hard to believe that whereas the former is simple enough to be biologically added, the latter is not. Or even that the relative simplicity of the two could possibly matter. Ditto for other operations like concatenation in place of Merge as the simplest operation.  Given how long this post is already, I will refrain from elaborating these points here.
[11] Birds (and mice and other animals) can string “syllables” together (put them together in a left/right order) to make songs. From what I can tell, there is no hard upper bound on how many syllables can be so combined.  These do not display hierarchy, but they may be recursive in the sense that the combination operation can iterate. Might it not be possible that what we find in FL builds on this iteration operation? That the recursion we find in FL is iteration plus something novel (I have suggested labeling is the novelty)? My point here is not that this is correct, but that the question of simplicity in a DP context need not just be a matter of conceptual simplicity.  
[12] How are sets formed? How computationally simple is the comprehension axiom in set theory, for example? It is actually logically quite involved (see here). I ask because Merge is a set forming operation, so the relevant question is how cognitively complex is it to form arbitrary sets. We have been assuming that this is conceptually simple and hence cognitively easy. However, it is worth considering just how easy. The Wikepedia entry suggests that it is not a particularly simple operation. Sets are funny things and what mental powers go into being able to construct them is not all that clear.

Thursday, May 9, 2013

Phases: some questions


One of the nice thing about conferences is that you get to bump into people you haven’t seen for a while. This past weekend, we celebrated our annual UMD Mayfest (it was on prediction in ling sensitive psycho tasks) and, true to form, one of the highlights of the get together was that I was able to talk to Masaya Yoshida (a syntax and psycho dual threat at Northwestern) about islands, subjacency, phases and the argument-adjunct movement asymmetry.  At any rate, as we talked, we started to compare Phase Theory with earlier approaches to strict cyclicity (SC) and it again struck me how unsure I am that the new fangled technology has added to our stock of knowledge.  And, rather than spending hours upon hours trying to figure this out solo, I thought that I would exploit the power of crowds and ask what the average syntactician in the street thinks phases have taught us above and beyond standard GB wisdom.  In other words, let’s consider this a WWGBS (what would GB say) moment (here) and ask what phase wise thinking has added to the discussion.  To set the stage, let me outline how I understand the central features of phase theory and also put some jaundiced cards on the table, repeating comments already made by others. Here goes.

Phases are intended to model the fact that grammars are SC. The most impressive empirical reflex of this is successive cyclic A’-movement.  The most interesting theoretical consequence is that SC grammatical operations bound the domain of computation thereby reducing computational complexity.  Within GB these two factors are the province of bounding theory, aka Subjacency Theory (ST). The classical ST comes in two parts: (i) a principle that restricts grammatical commerce (at least movement) to adjacent domains (viz. there can be at most one bounding node (BN) between the launch site and target of movement) and (ii) a metric for “measuring” domain size (viz. the unit of measure is the BN and these are DP, CP, (vP), and maybe TP and PP).[1] Fix the bounding nodes within a given G and one gets locality domains that undergird SC. Empirically A’-movement applies strictly cyclically because it must given the combination of assumptions (i) and (ii) above.

Now, given this and a few other assumptions and it is also possible to model island effects in a unified way.  The extra assumptions are: (iii) some BNs have “escape hatches” through which a moving element can move from one cyclic domain to another (viz. CP but crucially not DP) (iv) escape hatches can accommodate varying numbers of commuters (i.e. the number of exits can vary; English thought to have just one, while multiple WH fronting languages have many). If we add a further assumption - (v) DP and CP (and vP) are universally BNs but Gs can also select TP and PP as BNs – the theory allows for some typological variation.[2] (i)-(v) constitutes the classical Subjacency theory. Btw, the reconstruction above is historically misleading in one important way.  SC was seen to be a consequence of the way in which island effects were unified. It’s not that SC was modeled first and then assumptions added to get islands, rather the reverse; the primary aim was to unify island effects and a singular consequent of this effort was SC. Indeed, it can be argued (in fact I would so argue) that the most interesting empirical support for the classical theory was the discovery of SC movement.

One of the hot debates when I was a grad student was whether long distance movement dependencies were actually SC. Kayne and Pollock and Torrego provided (at the time surprising) evidence that it was, based on SC inversion operations in French and Spanish.  Chung supplied Comp agreement evidence from Chamorro to the same effect.  This, added to the unification of islands, made ST the jewel in the GB crown, both theoretically and empirically. Given my general rule of thumb that GB is largely empirically accurate, I take it as relatively uncontroversial that any empirically adequate theory of FL must explain why Gs are SC.

As noted in a previous post (here), ST developed and expanded.  But let’s leave history behind and jump to the present. Phase Theory (PT) is the latest model for SC. How does it compare with ST?  From where I sit, PT looks almost isomorphic to it, or at least a version that extends to cover island effects does.  A PT of this ilk has CP, vP and DP as phases.[3] It incorporates the Phase Impenetrabiltiy Condition (PIC) that requires that interacting expressions be in (at most) adjacent phases.[4] Distance is measured from one phase edge to the next (i.e. complements to phase heads are grammatically opaque, edges are not). This differs from ST in that the cyclic boundary is the phase/BN head rather than the MaxP of the Phase/BN head, but this is a small difference technically. PT also assumes “escape hatches” in the sense that movement to a phase edge moves an expression from inside one phase into the next higher phase domain and, as in ST, different phases have different available edges suitable for “escape.”  If we assume that Cs have different numbers of available phase edges and we assume that D has no such available edges at all then we get a theory effectively identical to the ST.  In effect, we traded phase edges for escape hatches and the PIC for (i).[5]
There are a few novelties in PT, but so far as I can tell they are innovations compatible with ST. The two most distinctive innovations regard the nature of derivations and multiple spell out (MSO). Let me briefly discuss each, in reverse order.

MSO is a revival of ideas that go back to Ross, but with a twist.  Uriagereka was the first to suggest that derivations progressively make opaque parts of the derivation by spelling them out (viz. spell out (SO) entails grammatical inaccessibility, at least to movement operations).  This is not new.  ST had the same effect, as SC progressively makes earlier parts of the derivation inaccessible to later parts.  PT, however, makes earlier parts of the derivation inaccessible by disappearing the relevant structure.  It’s gone, sent to the interfaces and hence no longer part of the computation.  This can be effected in various ways, but the standard interpretations of MSO (due to Chomsky and quite a bit different form Uriagereka’s) have coupled SO with linearization conditions in some way (Uriagereka does this as do Fox and Pesetsky, in a different way). This has the empirical benefit of allowing deletion to obviate islands. How? Deletion removes the burden of PF linearization and if what makes an island an island are the burdens of linearization (Uriagereka) or frozen linearizations (Fox and Pesetsky) then as deletion obviates the necessity of linearization, island effects should disappear, as they appear to do (Ross was the first to note this (surprise, surprise) and Merchant and Lasnik have elaborated his basic insight for the last decade!). At any rate, interesting though this is (and it is very interesting IMO), it is not incompatible with ST. Why? Because, ST never said what made an island an island, or more accurately, what made earlier cyclic material unavailable to later parts of the computation. (i.e. it had not real theory of inaccessibility, just a picture) and it is compatible with ST that it is PF concerns that render earlier structure opaque. So, though PT incorporates MSO, it is something that could have been added to ST and so is not an intrinsic feature of PT accounts. In other words, MSO does not follow from other parts of PT any more than it does from ST. It is an add-on; a very interesting one, but an add-on nonetheless.[6]

Note, btw, that MSO accounts, just like STs require a specification of when SO occurs. It occurs cyclically (i.e. either at the end of a relevant phase, or when the next phase head is accessed) and this is how PT models SC. 

The second innovation is that phases are taken to be the units of computation.  In Derivation by Phase, for example, operations are complex and non-markovian within the phase.  This is what I take Chomsky to mean when he says that operations in a phase apply “all at once.” Many apply simultaneously (hence not one “line” at a time) and they have no order of application. I confess to not fully understanding what this means. It appears to require a “generate and filter” view of derivations (e.g. intervention effects are filters rather than conditions on rule application).  It is also the case that SO is a complex checking operation where features are inspected and vetted before being sent for interpretation.  At any rate, the phase is a very busy place: multiple operations apply all at once; expressions E and I merged, features checked and shipped.

This is a novel conception of the derivation, but again, is not inherent in the punctate nature of PT.[7] Thus, PT has various independent parts, one of which is isomorphic to traditional ST and other parts that are logically independent of one another and the ST similar part. That which explains SC is the same as what we find in ST and is independent of the other moving parts. Moreover, the parts of PT isomorphic to ST seem no better motivated (and no less worse) than the analogous features in ST: e.g. why the BNs are just these has no worse answer within ST than the question why the phase heads are just those.

That’s how I see PT.  I have probably skipped some key features. But here are some crowd directed questions: What are the parade cases empirically grounding PT? In other words, what’s the PT analogue of affix hopping? What beautiful results/insights would we loose if we just gave PT up? Without ST we loose an account of island effects and SC. Without PT we loose…? Moreover, are these advantages intrinsic to minimalism or could they have already been achieved in more or less the same form within GB. In other words, is PT an empirical/theoretical advance or just a rebranding of earlier GB technology/concepts (not that there is anything intrinsically wrong with this, btw)?  So, fellow minimalists, enlighten me. Show me the inner logic, the “virtual conceptual necessity” of the PT system as well as its empirical virtues. Show me in what ways we have advanced beyond our earlier GB bumblings and stumblings. Inquiring minimalist minds (or at least one) want to know.



[1] This “history” compacts about a decade of research and is somewhat anachronistic.  The actual history is quite a bit more complicated (thanks Howard).
[2] Actually, if one adds vP as a BN then Rizzi like differences between Italian and English cannot be accommodated. Why? Because, once one moves into an escape hatch movement is thereafter escape hatch to escape hatch, as Rizzi noted for Italian. The option of moving via CP is only available for the first move. Thereafter, if CP is a BN movement must be CP to CP. If vP is added as a BN then it is the first available BN and whether one moves through it or not, all CP positions must be occupied. If this is too much “inside baseball” for you, don’t sweat it. Just the nostalgic reminiscences of a senior citizen.
[3] vP is an addition from Barriers versions of ST, though how it is incorporated into PT is a bit different from how vP acted in ST accounts.
[4] There are two versions of the PIC, one that restricts grammatical commerce to expressions in the same phase and a looser one that allows expressions in adjacent phases to interact. The latter is what is currently assumed (for pretty meager empirical reasons IMO – Nominative object agreement in quirky subject transitive sentences in Icelandic, I think).
[5] As is well known, Chomsky has been reluctant to extend phase status to D. However, if this is not done then PT cannot account for island effects at all and this removes one of the more interesting effects of cyclicity. There has been some allusions to the possibility that islands are not cyclicity effects, indeed not even grammatical effects.  However, I personally find the latter suggestion most implausible (see the forthcoming collection on this edited by Jon Sprouse and yours truly: out sometime in the fall). As for the former, well, if islands are grammatical effects (and like I said, the evidence seems to me overwhelming) then if PT does not extend to cover these then it is less empirically viable than ST.  This does not mean that it is wrong to divorce the two, but it does burden the revisionist with a pretty big theoretical note payable.
[6] MSO is effectively a theory of the PIC. Curiously, from what I gather, current versions of PT have began mitigating the view that SO removes structure by sending it to the interfaces. The problem is that such early shipping makes linearization problematic.  It is also necessitates processes by which spelled out material is “reassembled” so that the interfaces can work their interpretive magic (think binding which is across interfaces, or clausal intonation, which is also defined over the entire sentence, not just a phase).
[7] Nor is the assumption that lexical access is SC (i.e. the numeration is accessed in phase sized chunks). This is roughly motivated on (IMO view weak) conceptual reasons concerning SC arrays reducing computational complexity and empirical facts about Merge over Move (btw: does anyone except me still think that Merge over Move regulates derivations?).