Comments

Showing posts with label Bounding Nodes. Show all posts
Showing posts with label Bounding Nodes. Show all posts

Thursday, May 9, 2013

Phases: some questions


One of the nice thing about conferences is that you get to bump into people you haven’t seen for a while. This past weekend, we celebrated our annual UMD Mayfest (it was on prediction in ling sensitive psycho tasks) and, true to form, one of the highlights of the get together was that I was able to talk to Masaya Yoshida (a syntax and psycho dual threat at Northwestern) about islands, subjacency, phases and the argument-adjunct movement asymmetry.  At any rate, as we talked, we started to compare Phase Theory with earlier approaches to strict cyclicity (SC) and it again struck me how unsure I am that the new fangled technology has added to our stock of knowledge.  And, rather than spending hours upon hours trying to figure this out solo, I thought that I would exploit the power of crowds and ask what the average syntactician in the street thinks phases have taught us above and beyond standard GB wisdom.  In other words, let’s consider this a WWGBS (what would GB say) moment (here) and ask what phase wise thinking has added to the discussion.  To set the stage, let me outline how I understand the central features of phase theory and also put some jaundiced cards on the table, repeating comments already made by others. Here goes.

Phases are intended to model the fact that grammars are SC. The most impressive empirical reflex of this is successive cyclic A’-movement.  The most interesting theoretical consequence is that SC grammatical operations bound the domain of computation thereby reducing computational complexity.  Within GB these two factors are the province of bounding theory, aka Subjacency Theory (ST). The classical ST comes in two parts: (i) a principle that restricts grammatical commerce (at least movement) to adjacent domains (viz. there can be at most one bounding node (BN) between the launch site and target of movement) and (ii) a metric for “measuring” domain size (viz. the unit of measure is the BN and these are DP, CP, (vP), and maybe TP and PP).[1] Fix the bounding nodes within a given G and one gets locality domains that undergird SC. Empirically A’-movement applies strictly cyclically because it must given the combination of assumptions (i) and (ii) above.

Now, given this and a few other assumptions and it is also possible to model island effects in a unified way.  The extra assumptions are: (iii) some BNs have “escape hatches” through which a moving element can move from one cyclic domain to another (viz. CP but crucially not DP) (iv) escape hatches can accommodate varying numbers of commuters (i.e. the number of exits can vary; English thought to have just one, while multiple WH fronting languages have many). If we add a further assumption - (v) DP and CP (and vP) are universally BNs but Gs can also select TP and PP as BNs – the theory allows for some typological variation.[2] (i)-(v) constitutes the classical Subjacency theory. Btw, the reconstruction above is historically misleading in one important way.  SC was seen to be a consequence of the way in which island effects were unified. It’s not that SC was modeled first and then assumptions added to get islands, rather the reverse; the primary aim was to unify island effects and a singular consequent of this effort was SC. Indeed, it can be argued (in fact I would so argue) that the most interesting empirical support for the classical theory was the discovery of SC movement.

One of the hot debates when I was a grad student was whether long distance movement dependencies were actually SC. Kayne and Pollock and Torrego provided (at the time surprising) evidence that it was, based on SC inversion operations in French and Spanish.  Chung supplied Comp agreement evidence from Chamorro to the same effect.  This, added to the unification of islands, made ST the jewel in the GB crown, both theoretically and empirically. Given my general rule of thumb that GB is largely empirically accurate, I take it as relatively uncontroversial that any empirically adequate theory of FL must explain why Gs are SC.

As noted in a previous post (here), ST developed and expanded.  But let’s leave history behind and jump to the present. Phase Theory (PT) is the latest model for SC. How does it compare with ST?  From where I sit, PT looks almost isomorphic to it, or at least a version that extends to cover island effects does.  A PT of this ilk has CP, vP and DP as phases.[3] It incorporates the Phase Impenetrabiltiy Condition (PIC) that requires that interacting expressions be in (at most) adjacent phases.[4] Distance is measured from one phase edge to the next (i.e. complements to phase heads are grammatically opaque, edges are not). This differs from ST in that the cyclic boundary is the phase/BN head rather than the MaxP of the Phase/BN head, but this is a small difference technically. PT also assumes “escape hatches” in the sense that movement to a phase edge moves an expression from inside one phase into the next higher phase domain and, as in ST, different phases have different available edges suitable for “escape.”  If we assume that Cs have different numbers of available phase edges and we assume that D has no such available edges at all then we get a theory effectively identical to the ST.  In effect, we traded phase edges for escape hatches and the PIC for (i).[5]
There are a few novelties in PT, but so far as I can tell they are innovations compatible with ST. The two most distinctive innovations regard the nature of derivations and multiple spell out (MSO). Let me briefly discuss each, in reverse order.

MSO is a revival of ideas that go back to Ross, but with a twist.  Uriagereka was the first to suggest that derivations progressively make opaque parts of the derivation by spelling them out (viz. spell out (SO) entails grammatical inaccessibility, at least to movement operations).  This is not new.  ST had the same effect, as SC progressively makes earlier parts of the derivation inaccessible to later parts.  PT, however, makes earlier parts of the derivation inaccessible by disappearing the relevant structure.  It’s gone, sent to the interfaces and hence no longer part of the computation.  This can be effected in various ways, but the standard interpretations of MSO (due to Chomsky and quite a bit different form Uriagereka’s) have coupled SO with linearization conditions in some way (Uriagereka does this as do Fox and Pesetsky, in a different way). This has the empirical benefit of allowing deletion to obviate islands. How? Deletion removes the burden of PF linearization and if what makes an island an island are the burdens of linearization (Uriagereka) or frozen linearizations (Fox and Pesetsky) then as deletion obviates the necessity of linearization, island effects should disappear, as they appear to do (Ross was the first to note this (surprise, surprise) and Merchant and Lasnik have elaborated his basic insight for the last decade!). At any rate, interesting though this is (and it is very interesting IMO), it is not incompatible with ST. Why? Because, ST never said what made an island an island, or more accurately, what made earlier cyclic material unavailable to later parts of the computation. (i.e. it had not real theory of inaccessibility, just a picture) and it is compatible with ST that it is PF concerns that render earlier structure opaque. So, though PT incorporates MSO, it is something that could have been added to ST and so is not an intrinsic feature of PT accounts. In other words, MSO does not follow from other parts of PT any more than it does from ST. It is an add-on; a very interesting one, but an add-on nonetheless.[6]

Note, btw, that MSO accounts, just like STs require a specification of when SO occurs. It occurs cyclically (i.e. either at the end of a relevant phase, or when the next phase head is accessed) and this is how PT models SC. 

The second innovation is that phases are taken to be the units of computation.  In Derivation by Phase, for example, operations are complex and non-markovian within the phase.  This is what I take Chomsky to mean when he says that operations in a phase apply “all at once.” Many apply simultaneously (hence not one “line” at a time) and they have no order of application. I confess to not fully understanding what this means. It appears to require a “generate and filter” view of derivations (e.g. intervention effects are filters rather than conditions on rule application).  It is also the case that SO is a complex checking operation where features are inspected and vetted before being sent for interpretation.  At any rate, the phase is a very busy place: multiple operations apply all at once; expressions E and I merged, features checked and shipped.

This is a novel conception of the derivation, but again, is not inherent in the punctate nature of PT.[7] Thus, PT has various independent parts, one of which is isomorphic to traditional ST and other parts that are logically independent of one another and the ST similar part. That which explains SC is the same as what we find in ST and is independent of the other moving parts. Moreover, the parts of PT isomorphic to ST seem no better motivated (and no less worse) than the analogous features in ST: e.g. why the BNs are just these has no worse answer within ST than the question why the phase heads are just those.

That’s how I see PT.  I have probably skipped some key features. But here are some crowd directed questions: What are the parade cases empirically grounding PT? In other words, what’s the PT analogue of affix hopping? What beautiful results/insights would we loose if we just gave PT up? Without ST we loose an account of island effects and SC. Without PT we loose…? Moreover, are these advantages intrinsic to minimalism or could they have already been achieved in more or less the same form within GB. In other words, is PT an empirical/theoretical advance or just a rebranding of earlier GB technology/concepts (not that there is anything intrinsically wrong with this, btw)?  So, fellow minimalists, enlighten me. Show me the inner logic, the “virtual conceptual necessity” of the PT system as well as its empirical virtues. Show me in what ways we have advanced beyond our earlier GB bumblings and stumblings. Inquiring minimalist minds (or at least one) want to know.



[1] This “history” compacts about a decade of research and is somewhat anachronistic.  The actual history is quite a bit more complicated (thanks Howard).
[2] Actually, if one adds vP as a BN then Rizzi like differences between Italian and English cannot be accommodated. Why? Because, once one moves into an escape hatch movement is thereafter escape hatch to escape hatch, as Rizzi noted for Italian. The option of moving via CP is only available for the first move. Thereafter, if CP is a BN movement must be CP to CP. If vP is added as a BN then it is the first available BN and whether one moves through it or not, all CP positions must be occupied. If this is too much “inside baseball” for you, don’t sweat it. Just the nostalgic reminiscences of a senior citizen.
[3] vP is an addition from Barriers versions of ST, though how it is incorporated into PT is a bit different from how vP acted in ST accounts.
[4] There are two versions of the PIC, one that restricts grammatical commerce to expressions in the same phase and a looser one that allows expressions in adjacent phases to interact. The latter is what is currently assumed (for pretty meager empirical reasons IMO – Nominative object agreement in quirky subject transitive sentences in Icelandic, I think).
[5] As is well known, Chomsky has been reluctant to extend phase status to D. However, if this is not done then PT cannot account for island effects at all and this removes one of the more interesting effects of cyclicity. There has been some allusions to the possibility that islands are not cyclicity effects, indeed not even grammatical effects.  However, I personally find the latter suggestion most implausible (see the forthcoming collection on this edited by Jon Sprouse and yours truly: out sometime in the fall). As for the former, well, if islands are grammatical effects (and like I said, the evidence seems to me overwhelming) then if PT does not extend to cover these then it is less empirically viable than ST.  This does not mean that it is wrong to divorce the two, but it does burden the revisionist with a pretty big theoretical note payable.
[6] MSO is effectively a theory of the PIC. Curiously, from what I gather, current versions of PT have began mitigating the view that SO removes structure by sending it to the interfaces. The problem is that such early shipping makes linearization problematic.  It is also necessitates processes by which spelled out material is “reassembled” so that the interfaces can work their interpretive magic (think binding which is across interfaces, or clausal intonation, which is also defined over the entire sentence, not just a phase).
[7] Nor is the assumption that lexical access is SC (i.e. the numeration is accessed in phase sized chunks). This is roughly motivated on (IMO view weak) conceptual reasons concerning SC arrays reducing computational complexity and empirical facts about Merge over Move (btw: does anyone except me still think that Merge over Move regulates derivations?).

Wednesday, May 1, 2013

Formal and Substantive Universals


This post will be pretty free form, involving more than a little thinking out loud (aka rambling). It will maunder a bit and end pretty inconclusively. If this sort of thing is not to your liking, here would be a good place to stop.

I’ve recently read an interesting paper on a question that I’ve been thinking about off and on for about a decade (sounds longer than 10 years eh?) by Epstein, Kitihara and Seely (EKS) (here).  The question: to what degree are licit formal dependencies of interacting expressions functions of the substantive characteristics of the dependent elements? This is a mouthful of a sentence, but the idea is pretty simple: we have lots of grammatical dependencies, how much do they depend on the specific properties of specific lexical/functional items involved?[1]  Let me give a couple of illustrations to clarify what I’m trying to get at.

Take the original subjacency condition. It prohibited two expressions from interacting if one is within an island and the other is outside that island. So in (1) Y cannot move to X:

(1)  […X…[island…Y…]…]

Now, we can list islands by name (e.g. CNPC, WH-island, Subject Islands etc.) or we can try to unify them in some way. The first unification (due to Chomsky) involved two parts; the first a specification of how far is too far (at most one bounding node between X and Y), the second an identification of the bounding nodes (BN) (DP and CP, optionally TP and PP etc.). Now, the way I always understood things is that the first part of the “definition” was formal (i.e. the same principle holds regardless of the BN inventory), the second substantive (i.e. the attested dependencies depend on the actual choice of BNs). Indeed, Rizzi’s famous paper (actually the one limned in the footnotes, rather than the one in the text) was all about how to model typological differences via small changes in the inventory of BNs for a given grammar.  So, the classical theory of subjacency comprises a formal part that does not care about the actual categories involved and a substantive part, that cares a lot.

Later theories of islands cut things up a little differently. So, for example, one intriguing feature of Barriers was its ambition to eliminate the substantive part of subjacency theory.  Rather than actually listing the BNs, Barriers tried to deduce the class of BNs to general formal properties of the phrase marker.  Roughly speaking, complements are porous, while non-complements are barriers.[2] Complementation is itself an abstract formal dependency, largely independent of the contents of the interacting expressions.  I say “largely independent” for in Barriers it was critical that there be some form of L-marking that was itself dependent on theta marking. However, the L-marking relation was very generic and applied widely to many different kinds of expressions.

Cut to the present and phases: phases have returned to the original conception of BNs. Of course we now call them phase heads rather than BNs, and we include v as well as CP (an inheritance from Barriers) but what is important is that we list them.[3] The grammar functions as it does because v and C are points of transfer and they are points of transfer because they are phase heads. Thus, if you are not a phase head you are not a point of transfer. However, theoretically, you are a phase head because you have been so listed. BTW, as you all know, unless D is included in this inventory, we cannot code island effects in terms of phases.  And as you also all know, the phase-based account of islands is no more principled than the older subjacency account.[4]  However, this is not my topic here.  All I want to observe is how substantive assumptions interact with formal ones to determine the class of licit dependencies and how some accounts have a “larger” substantive component than others. I also want to register a minimalist observation (by no means original) that the substantive assumption about the inventory of Phases/BNs raises non-trivial minimalist queries: “why these?” being the obvious one. [5]

Let’s contrast this case with Minimality.  This, so far as I can tell, is a purely formal restriction, even in its relativized form. It states that in a configuration like (2), X,Y,Z being of the same type (i.e. sharing the same relevant features) Y cannot interact with X over an intervening Z. For (2) the actual feature specifications do not matter. Whatever they are, minimality will block interaction in these cases. This is what I mean by treating it as a purely formal condition.

            (2) …X…Z…Y…

So, we now have two different examples, let’s get back to the question posed in EKS: we all assume a universal base hypothesis with the rough structure C-T-v-V, to what degree does this base hierarchical order follow from formal principles?  Note, the base theory names the relevant heads in their relevant hierarchical order, the question is to what degree do formal principles force this order. EKS discuss this and argue that given certain current assumptions about phases, we can derive the fact that theta domains are nested within case domains and, suggest, that the same reasoning can apply to the upper C-T part of the base structure. Like I said the paper is interesting and I recommend it.  However, I would like to ask EKSs question in a slightly different way, stealing a trick from our friends in physics (recall, I am deep green with physics envy). 

Among the symmetries physicists study is one in which different elements are swapped for one another. Thus, as Carrol (here) noted concerning nuclear structure: “In 1954, Chen Ning Yang and Robert Mills came up with the idea that this symmetry should be promoted to a local symmetry – i.e., that we should be allowed to “rotate” neutrons and protons into each other at every point in space (154).” They did this to consider whether the strong force really cared about the obvious differences between protons and neutrons. Let’s try a similar trick within the C-T-v-V domain, this time “rotating” theta and case markers into each other, to see whether the ordering of elements in the base really affects what kinds of formal dependencies we find.

More specifically: consider the basic form of the sentence, and let’s consider only the dependencies within TP:

            (3) [CP C [TP …T…[vP Subj v [VP V Obj]]]]

In (3) Subj gets theta from v and case from T. Object gets theta from V and case from v. So there is a T-Subj relation, a Subj-v relation a v-Obj relation and a V-Obj relation.  Does it matter to these relations and further derivations that in fact the specific features noted are checked by the indicated heads. To get a handle on this, imagine if we systematically changed case for theta assignments above (i.e. rotated case and theta into each other so that T assigns theta to Subj and v assigns case, v assigns theta to Obj and V assigns case etc.) what would go wrong? If nothing goes wrong, then the actual labels here make no difference. If nothing goes wrong then the formal properties do not determine the actual substantive order

To sound puffed up and super scientific we might say that the formal properties are symmetric wrt the substantive features of case and theta assignment. Note, btw, we already think this way for theta and case values. The grammatical operations are symmetric with respect to these (i.e. they don’t care what the actual theta role or case value is).  We are just extending this reasoning one step further by asking about assignment as well as values.

Observe that things can go “wrong” in various ways: we could get lots of decent looking derivations honoring the formal restrictions but the derivations either under or over generate.  For example, If T assigns the external theta role then transitive small clauses might be impossible if small clauses have no structure higher than v.  This seems false. Or, if this is right, then we might expect expletives to always sit under neg in English as they cannot move to Spec t this being a theta position. Again, this seems wrong. So, there seem to be, at least at first blush, empirical consequences of making this rotation.  However, the “look” of the system is not that different if this is the only kind of problem, i.e. the resulting system is language like if not exactly identical to what we actually find. In other words, it’s a possible UG, just not our UG. There is clearly a minimalist question lurking here.

A second way things could go wrong is that we do not in general get convergent derivations. EKS argue that certain phase-based accounts have this more expansive consequence. The problem is not a little over/under generation, the problem is that we can barely get a decent derivation at all. In our little though experiment this means that rotating the case and theta values results in impossible UGs. This would be a fascinating result, with obvious minimalist intepretations.

Both kinds of “problems” are interesting. This first showing that our UG deeply cares about the substantive heads having the specific properties they do. The second suggests that there is a very strong tie between the basic structure of the clause and the formal universals we have.  Both kinds of results would be interesting.

I have no worked out answer to the ‘what goes wrong?’ question (though if you get one I would love to hear about it). Note that I have abstracted away from everything but what is assumed to be syntactically relevant; case and theta “features.” I have also assumed that how these features are assigned is symmetrical: that both are assigned in the same way. If this is false then this might be the source of the substantive base order noted (e.g. if theta were only under merge and case could be under agree). However, right now I am satisfied to leave my version of the EKS question open.

Let me end with two further random observations:

First, whatever answer we find, I find the question to be really cool and novel and exciting! We may need to stipulate all sorts of things but until we start asking the kinds of question EKS pose, we won’t know what is principled and what not.

Second, as noted, the answer to this question will have significant implications for the Minimalist Program (MP).  To date, many minimalists (e.g. me) have concentrated on trying the unify the various dependencies as prelude to explaining their properties in generic cognitive/computational terms.  However, if C-T-v-V base structure is part of FL/UG then it presents a different kind of challenge to MP, given the apparent linguistic specificity of the universals.  Few notions are more linguistically parochial than BN or phase head or C-T-v-V.  It would be nice if some of this followed from architectural features of the grammar as EKS suggest, or from the demands of the interface (e.g. think Heim’s tri-partite semantic structures), or something else still. The real challenge to MP arises if these kinds of substantive universals are brut. At any rate, it seems to me that substantive universals present a different kind of challenge to MP and so they are worth thinking about very carefully. 

That’s enough rambling.

   


[1] I know, all the rage lately has been to pack all interesting grammatical properties into functional heads, specific lexical content being restricted to roots. The question still arises: do we need to know the special properties of the involved heads, be they functional or not, to get the right class of grammatical dependencies?
[2] This idea was not original to Barriers.  Cattell and Cinque, I believe, had a similar theoretical intuition earlier.
[3] I think that we can all agree that Chomsky’s attempt to relate the selction of C and v to semantic properties has not been a singular success. Moreover, his rationalization leaves out D, without which extending phases to cover island effects is impossible. If phases do not explain island phenomena, then their utility is somewhat circumscribed, given the view over the last 30 years that cyclicity and islandhood are tightly connected. Indeed, one might say that the central empirical prediction of the subjacency account was successive cyclic movement. All the other stuff was there to code Ross’s observations. Successive cyclicity was a novel (and verified) consequence.
[4] Boeckx and Grohmann (2007) have gone through this in detail and so far as I can tell, the theoretical landscape has stayed pretty much the same.
[5] There are attempts to give a “Barriers” version of phases (e.g. Den Dikken) and attempts to argue that virtually every “Max P” is a phase in order to finesse this minimalist problem.