Comments

Showing posts with label Minimal labeling algorithm. Show all posts
Showing posts with label Minimal labeling algorithm. Show all posts

Monday, August 4, 2014

Final comments on lecture 4

This ends the comments (here and here) on lecture 4.

The logic used to account for the EPP also covers Fixed Subject Condition effects (FSC).[1] Consider (2’) again:

(2’) *Who1 did John say that t1 saw Mary

The T needs to be labeled. If who moves there will be nothing in the Spec-T to agree with and so labeling will fail.  That’s Chomsky’s story. The obvious problem with this account is the absence of FSCs if there is no overt complementizer (I pointed to this in the comments to lecture 3). Chomsky addresses this problem here. He proposes that a deleted C is no longer a phase. More exactly, to delete a C you must transfer the feature that says that C is a phase to T. In effect, C deletion makes T the phase head. So not only do we lower phi and tense features form C to T, but phase-headedness as well.  This now ties FSCs to the presence of absence of an overt C.[2]

Observation 1: This story requires that that is deleted rather than not present at all. Were it never present, C could not transfer its features to T, and T has not features of its own (more below). Thus, to make this work, we need deletion operations in the syntax. A question that arises is how similar the operation deleting that is to more run of the mill ellipsis operations. The latter are generally treated as simply dephoneticization processes. This will not suffice here. It must be that getting rid of phonetic content requires that all features of C lower to T. For those with long enough memories, this smells a little of the old notion of “L-contains.”  At any rate, it’s worth observing that C deletion is not simply quieting the phonetics.

Observation 2: there are well known variations regarding that-t effects dialectally in English. This suggests that deletion might be sufficient for transferring all of Cs features to T but it is not necessary. So, contrary to what Chomsky suggests, the explanation requires that we say something “special” about these FSCs in English. IMO, things are a little worse than this. As I mentioned in earlier comments, many speakers seem to allow violations of the FSC in English even with if/whether in the C position (or at least so report a third of my syntax undergrads). There is no problem accommodating this by allowing feature lowering as Chomsky suggests. But this is now decoupled from phonetic articulation. Unfortunately, in the relevant dialects, deleting a that does not license null subjects, which one might have expected (4).  Or more exactly, why doesn’t lowering all the features of C to T serve to strengthen T? Note that Cs can label just fine without anything in Spec-C helping them along. Given this, why shouldn’t lowering all the features of C onto T (including the “phase head feature” see below) not make T as independent as C?  Dunno, but it doesn’t.
           
(1)  *John thinks (that) is a man here

In other words, the EPP and FSC do not really swing together, though they should if they were truly unified, one might suppose. 

Let’s put these questions to one side and continue. Chomsky then asks how we can get (2):

(2)  Who do you think t is kissing Sue

How do we label the lower “TP” if who moves. Chomsky says that it is labeled before C deletes and labels cannot be deleted. In other words, labels are indelible (think Lasnik and Saito).  Now, Chomsky really doesn’t like this way of putting things. What he wants to say is that CSs have phase sized memories (i.e. CS can recall what operations have taken place within a phase). In other words, there is phase level memory for all syntactic operations.  By lowering the phase-hood to T from C, the next higher v* phase can “remember” that the lower T was given a label via agree and so movement of the DP in Spec T is ok.  So, it is not that the labeling is indelible, but that all operations that happen within a phase are recollected in that phase. So, once labeled, memory tells us that it is always labeled. 

This emphasizes the computational aspects of phase theory.  What’s important about phases is that they reduce memory demands of a computation. The reverse of this is that it allows some memory of previous operations to be retained.  This is quite definitely not a conceptual argument. There is no conceptual motivation for these assumptions. The motivations are computational. The question becomes how we bring information forward in a derivation, how forgetting can be computationally efficient etc.  Phases and the properties Chomsky relies on here are entirely of this variety.

Comment: Two things: this is very Barrerish in spirit. Phases are no longer fixed, but change in the course of the derivation (as Den Dikken was the first to propose). And call it what you want, indelibility is back.  Moreover, just as in Barriers, T has a strange role in this system. It is not an inherent phase but can become one by inheritance. Sound familiar?[3] To me, this all has the feeling of a Rube Goldberg device, but this is partly a matter of taste. Some might think Barriers elegant. Go figure.

Let me make my unease clearer. It seems that Chomsky is not that happy with T and its various special properties within his system (again, just like T in Barriers).  He, in fact, proposes, that T has no properties of its own. It’s just there to receive properties from C. This makes T very similar to Agr in older MP, and recall that Chomsky argued that grammar internal formatives like Agr are to be eschewed. They cause DP headaches, as do any grammar internal formatives (viz. we need to explain how they and their properties got into FL).  Worse, IMO, the special properties of T are critical in Chomsky’s explanations of the EPP and FSC effects. But, this strikes me as a non-trivial problem for his account. Why doesn’t explaining these special properties of FL in terms of special properties of T not amount to re-description rather than explanation? One of the salutary effects of MP has been to warn us about confusing the two. However, the more T is special, the more accounts of Spec-T effects (EPP and FSCs) are weakened. And from what I can tell, Ts special properties are critical in deriving the results Chomsky obtains.

There are other Barrier like resonances here. Recall that in the Barriers framework, VP was never a barrier. Why? Because we could always adjoin to it and thereby void its barrierhood.  In the present story, there is also a big asymmetry between C and v*. The latter never displays FSC or ECP effects. Why not? Because v* always looses its phase property. How? Because Chomsky assumes that when the root raises to v*, v* gets buried inside the raised root and thereby looses its phase properties.  Though Chomsky does not discuss this, it raises the question of the role of v* in a phase based theory if it always looses its phase property. Do v*s no longer induce PIC effects?  It would be very odd to assume that the lower copy of the root inherited the phase property, like T inherits it from C. After all, this is the tail of a head chain and tails are generally grammatically inert.  So, one conclusion could be that v* never induces PIC effects on this revised account. Of course, once again, we can add technical fixes to obviate this conclusion, but, at least for me, their motivations will be empirical not conceptual.  This is not bad, but it does do some damage to the SMT line of reasoning Chomsky cherishes.

So, v* gets special treatment because of the properties of roots and T gets special treatment because its just, well, odd. The story may hang together, but it is hardly conceptually pretty, at least from where I sit.

Question: When R(V) raises to v* what happens to the phase property of v*? I assume that it is eliminated and this is why there are no EPP/FSC effects. Ok, does R(V) inherit the phase-hood property? This seems unlikely, as it is the tail of the head chain, but maybe.  If not, is the phase-hood of v* simply voided and what we are left with is one big C-phase?  If so, how does this enhance computational efficiency?[4]  

This post is already way too long. Let me end with three more highlights and maybe a remark or two.

Lecture 4 drops the idea that all grammatical action takes place at the phase head. In particular he allows I-merge to apply completely independently of any relationship to C. He does this to get rid of the counter-cyclic movement he needed in PoP. Recall counter-cyclic movement violates the NTC, which is a part of the conceptually best version of Merge. Chomsky really didn’t want to allow it and he eliminates here at the cost of abandoning the assumption that all grammatical operations are with the phase head. 

A consequence of this is that Chomsky has to abandon his explanation from PoP as to why Gs raise T to C but don’t raise DPs to C instead given that they are equally close if SLOs are without labels. You may recall, that he accounted for this by moving the subject to Spec T counter-cyclically in derivations that applied “all at once” at the phase level. 

I’m glad that Chomsky drops this now for two reasons. First, I never liked his earlier explanation (indeed I was part of a trio arguing that it didn’t work and was the wrong way to proceed (here)). Second, I never understood what “all at once” derivations meant. Ever.  I pretended to and taught it, but never got it.  Now I don’t have to, it seems. Curiously, Chomsky not only drops this analysis but states that it was always “artificial.” Yup.

Chomsky also finally abandons the last residues of Greed based MP theories. Merge is free. If you follow him here, you can stop worrying about what motivates this or that movement. Note, this is conceptually the right move (and I have thought this for a long time and even said so publically). Chomsky notes that E-merge is not subject to greed considerations and so I-merge shouldn’t be either, given that they are two instances of the very same operation.  Again, yup.  So much for probing and agreeing being a pre-condition for I-merge.

Does this mean that movement is never “for a reason”? Well it does mean that it is never for a local CS reason. Rather, we return once again to a “generate and filter” theory of computation, similar to what we had in GB. The main difference is that this time the filters are provided by Bare Output Conditions, in particular the requirement that all SLOs be labeled prior to Transfer and that MLA sometimes requires what is effectively a Spec-X configuration to provide a viable label. So, Probe/goal seems out and Spec-X is back. Plus ca change (and I know I’m missing a thingy under ‘c’).

This is enough. At least I’ve had enough. Let me once again end on a positive note. As you might have noticed, I am not (yet) convinced by Chomsky’s story here. I believe that he has mis-analyzed the role of labels in G.  His approach rests on the assumption that labels play no role in CS. They are only relevant to interface interpretation.  I currently believe that this is likely wrong. At the very least, I don’t really see how labels are required for CI (or SM) interpretive rules to apply. For CI, at least, all we need, so far as I can make out, is the branching structure and the contents of the individual atoms.  So the premise that motivates the MLA as a BOC seems (at least to me) very shaky.

However, though I am not that moved by Chomsky’s proposal, I am moved by his method and overall conception of the enterprise. I wholeheartedly agree that we should take Galileo’s Maxim as a strong boundary condition on theory.  I completely agree that we should respect GGs history and treat its generalizations as the targets for more principled explanation.[5] I also agree that the critics of GB that he mentions have contributed virtually nothing to our understanding of FL. I even believe that Chomsky’s efforts are an excellent illustration of what we should be aiming to do.  I just am not moved by the details of his effort.  But, really, that’s a small disagreement, among friends.






[1] Chomsky calls these ECP effects. I changed this to FSC effects to distinguish the subject/object asymmetry effects in the ECP from the argument/adjunct ones. Historically, the latter have proven more recalcitrant and were the ones that called forth the heavy ECP assumptions. Indeed, for some analyses, a good part of the subject/object asymmetries were relegated to head government effects (Rizzi, Aoun et al, Saito) more on the SM side of things than the CI.
[2] Transfer of the “phase feature” (PF) from C to T seems like a clear violation of the NTC. As stated, T that has no inherent PF receives one and thereby changes its grammatical powers. One can play around with definitions so that this violation of the NTC doesn’t count, but the definitions will not enjoy the conceptual naturalness that the conceptual versions of the SMT rely on.
[3] It seems that Chomsky is not that happy with T and its various properties.  He, in fact, suggests, that T has no properties of its own. It’s just there to receive properties from C. This makes T very similar to Agr in older MP, and recall that Chomsky argued that such grammar internal formatives are to be eschewed. They cause DP headaches, as do any grammar internal formatives.  The special properties of T are critical in Chomsky’s explanations of the EPP and FSC effects. But, this strikes me as a problem for his account. Why doesn’t explaining these special properties of FL in terms of special properties of T not amount to re-description rather than explanation? One of the salutary effects of MP has been to warn us about confusing the two. However, the more T is special, the more accounts of Spec-T effects (EPP and FSCs) are weakened. 
[4] And talking about memory load: what about the phase hood of D?  Chomsky really does not like to talk about D as a phase. He reluctantly mentions that it might be one from time to time, but it is reluctantly.  But if it is not a phase, then the memory demands on phases can be arbitrarily big given the recursive proclivities of DPs.  So, it would seem that on computational grounds, if v* is a phase then D should be one as well. Indeed, given that “clauses” are already bounded at C, why do we need to also computationally bound them at v*?  Remember, what constitutes a phase needs to be coded into FL and this causes problems for DP. So, all in all, we want the fewest phases we can get away with.
[5] I would go further: I think that these lectures constitute a bit of a departure in Chomsky’s minimalist style. The targets of explanation are now very broad generalizations that GG has established, the aim being to explain their properties.  Earlier efforts were far more local, the targets of explanation being Existential Constructions in Icelandic.  I think that this more general project gets the explanatory grain more correct.

Sunday, July 27, 2014

Comments on lecture 4-II

This is the second installment of comments on lecture 4. The first part is here.

The Minimal Labeling Algorithm (MLA)
           
Chomsky starts the discussion by reviewing the basics of his approach to labels.  Here are the key assumptions:
1.     Labels are necessary for interpreting structured linguistic objects (SLO) at the interfaces. I.e. Labels are not required for the operations of the computational system (CS).
2.     Labels are assigned by the MLA as follows
a.     All constructed SLOs must be labeled (see 1)
b.     Labels are assigned within phases
c.     MLA minimally searches a set (i.e SLO) to identify the most
prominent “element.” This element is the label of the set.[1]
                                                        i.     In cases like {X,YP} (i.e. in which X is atomic and Y complex), the MLA chooses X: the shallowest element capable of serving as a label
                                                      ii.     In {XP, YP} (i.e in cases where both members of the set are complex), There is no unique choice to serve as label. In such cases, if XP and YP “agree,” (e.g. phi feature agreement, or WH feature agreement) then MLA chooses the common agreement features as the label.
3.     Ancillary assumptions:
a.     Only XPs that are heads of their chains are “visible” within a given set. Thus, in {XP,YP} if XP is not head of its XP chain, then it is invisible within the set {XP,YP}, (i.e. movement removes X within XP as a potential label).
b.     Roots are inherently incapable of labeling.
c.     T is parametrically a label. The capacity of T to serve as a label is related to “richness of agreement.” It is “rich” in Italian so Italian T can serve as a label. It is “poor” in English so English T cannot serve as a label.
d.     If a “weak” head (root or T) agrees with some X in XP then the agreement features can serve as a label.

Assumptions in 3 and 4 suffice to explain several interesting features of FL: the Fixed Subject Condition (FSC) (the subject/object asymmetries in “ECP effects”), EPP effects, and the presence of displacement. Let’s see how.

Consider first the EPP.  {T, YP} requires a label. In Italian this is not a problem for rich agreement endows T with labeling prowess.[2] English finesses this problem by raising a DP to “Spec” T. In a finite clause, this induces agreement between the DP and the TP (well T’ in the “old” system, but whatever) and the shared phi features can serve as the label.  If, however, DP fails to raise to Spec T or if DP in Spec T I-merges into some higher position, then it will not be available for agreement and the set that contains T will not receive a label and so will not be interpretable at CI or SM.  This is the account for the unacceptability of the sentences in (1) and (2) (traces used for convenience):

                 (1)  * left John
                       (2) *Who1 did John say that t1 saw Mary

Comments: Note that for this account to go through, we must assume that in {T, YP} that the “head” of Y is not a potential label. The fact that T cannot serve as a label does not yet imply that minimal search cannot find one.  Thus, say the complement of T were vP (here a weak v). Then the structure is {T, vP}. If T is not a potential label, then the shallowest possible label is v. Thus, we should be able to label the whole set ‘v’ and not move anything up. Why then is movement required?

One possible reason is that John cannot stay in its base position for some reason. One reason that it might have to move is that John cannot get case here, forcing John to move. The problem however, is that on a Probe-Goal system with non transitive v as a weak phase, T can probe John and assign it case (perhaps as a by-product of phi agreement, though I believe that this is empirically dubious).  Thus, given Chomsky’s standard assumptions, John can discharge whatever checking obligations it has without moving a jot.

So maybe, it needs to move for some other reason. One consistent with the assumptions above is that it needs to move so that {R(left), John} can be labeled.[3]   Recall, however, that Chomsky assumes that roots are universally incapable of labeling (3-c above) (Question: is 3-c a stipulation of UG or does it follow from more general minimalist assumptions? If the former then it exacerbates DP and so is an unwelcome stipulation (which is not to say that it is incorrect, but given GM something we should be suspicious of)). The structure of the set {R(left), John} is actually {R(left), {n, R(John)}.  In the latter there is a highest potential label, namely ‘n.’ So, it the MLA is charged with finding the most prominent potential label, then it would appear that even without movement of {n, R(John)}, the MLA could unambiguously apply.  Once again, it is not clear why I-merge is required.

Indeed, things are more obscure yet.  In this lecture Chomsky suggests that roots raise and combine with higher functional heads.  This implies that in {v, {[R(left) {n, R(John)}}} that R(left) vacates the lowest set and unites with ‘v.’ But this movement will make ‘R(left)’ invisible in the lowest set, again allowing ‘n’ to label it.  So, once again, it is not clear why John needs to raise to Spec T and why ‘v’ cannot serve to label {T, vP}.

Here’s another possibility: Maybe Chomsky is assuming a theoretical analogue of defective intervention. Here’s what I mean.  The MLA looks not for the highest potential labeler, but for the highest lexical atom, whether it can serve as a label or not. So in {T, vP}, T is the highest atom, it’s just that it cannot label. So, unless something moves to its spec to agree with it, we will not be able to label {T, vP} and we will face interpretive problems at the interfaces.  On this interpretation, then, the MLA does not look for the highest possible labeling atomic element, but simply the most prominent element regardless of its labeling capacities.  This will have the effect of forcing I-merge of John to Spec-T.

So, let’s so interpret the MLA.  Chomsky suggests that the same logic will force raising of an object to Spec-R(V) in a standard transitive clause.[4] Thus in something like (3a) the structure of the complement of v* is (3b) and movement of the object to Spec-R(kiss) (3c), could allow for the set to be labeled.

(    (3)  a. Mary kissed John
     b. {v*, {R(kiss), John}}
     c. {John, {R(kiss), John}}

However, once again, this presupposes that raising the root to v*, which Chomsky assumes to be universally required, will not suffice to disambiguate the structure for labeling.[5]

So, the EPP follows given this way of reading the proposal: MLA searches not for the first potential label, but for the closest lexical atoms. If it finds one, it is the label if it can be. If it cannot be, then tough luck, we need to label some other way. One way would be for its Spec to be occupied with an agreeing element, then the agreement can serve as the label. So even atoms that cannot label (even cannot label inherently) can serve to interfere with other elements serving as labels. This forces EPP movement in English so that agreement can resolve the ambiguity that stifles the MLA.

Question: how are non-finite TPs labeled? If they are strong, then there is no problem. But clearly they generally don’t display any agreement (at least any overt morphology). Thus, for infinitives, the link to “rich” morphology is questionable. If non-finite Ts are weak, however, then how can we get successive cyclic movement? Recall, the DP must remain in Spec T to allow for labeling. Thus, successive cyclic movement should make labeling impossible and we should get EPP problems?  In short, if a finite T requires a present subject in order to label the “T’”, why not a non-finite T?  A puzzle.

Question: I am still not clear how this analysis applies to Existential Constructions (EC). In an earlier lecture, Chomsky insisted (to reply to David P) that the EPP is mainly “about” the existence of null expletives (note: I didn’t understand this as I said in comments to lecture 3).  Ok: so English needs an expletive to label “T’” but Italiina doesn’t. It’s not that Italian has a null expletive, but that it has nothing in Spec-T as nothing is required.  So, what happens in English?  How exactly does there help label the “TP”?  Recall, the idea is that in {XP,YP} configurations there isn’t an unambiguous most prominent atom and so the common agreement features serve as the label. Does this mean that in ECs there and the T share agreement features?[6]  I am happy enough with this conclusion, but I thought that Chomsky took the agreement in ECs to be between the associate and T. Thus the agreement would not be between T and there but between T and the associate. How then does there figure in all of this?  What’s it doing?  Let me put this another way: the idea that Chomsky is pursuing is that agreement provides a way for the MLA to find a label when the structure is ambiguous. The label resolves the problem by identifying a common feature(s) of XP and YP and using it to label the whole. But in ECs it is standardly assumed that the agreement is not with the XP in Spec T but an associate in the complement domain of T.  So, either it is not agreement that resolves the labeling problem in ECs, or there has agreement features, or ECs are, despite appearances, not {XP,YP} SLOs.  At any rate, I am not sure what Chomsky would say about these, and they seem central to his proposal. Inquiring minds want to know.



[1] Chomsky seems to say that phase heads (PH) determine the label. I am not at all sure why we need assume that it’s the PH that via the MLA determines the label. It seems to me that MLA can function without a PH head being involved at all to, as it were, “choose” the label. What is needed is that the MLA be a phase level operation that applies, I assume, at Transfer. However, I may be wrong about how Chomsky thinks of the role of PHs in labeling, though I think this is what Chomsky actually says.
From what I can tell, MLA requires that the choice of label respect minimal search and that it be deterministic. I have interpreted this to mean, that the MLA applies to a given set of elements to unambiguously choose the label. It is important that the MLA does not tolerate labeling ambiguity (i.e. in a given structure exactly one element can serve as the label and it will be chosen as the label) for this is what forces movement and agreement, as we shall see. However, I do not see that the MLA requires that PHs actually do the labeling (i.e. choose the label). What is needed is that in every set, there be a uniquely shallowest potential label and that the MLA choose it.
I am not clear why PHs are required (if they are) to get the MLA to operate correctly. This may be a residue of an idea that Chomsky later abandons, namely that all rules are products of properties of the phase head. Chomsky, as I noted, dumps this assumption at the end of lecture 4 and this might suffice to liberate the MLA from PHs.  Note it would also allow for matrix clauses to be labeled. Minimal search is generally taken to be restricted to sister’s of probes. PHs then do not “see” their specifiers, making labeling of a matrix clause impossible. Freeing the MLA from PHs would eliminate this problem.
[2] Reminder: this is not an explanation. “Rich agreement” is just the name we give to the fact that T can do this in some languages and not in others. How it manages to do this is unclear. I mention this to avoid confusing a diacritical statement for an explanation. It is well known, that rich morphological agreement is neither necessary nor sufficient to license EPP effects, though there is a long tradition that identifies morphological richness as the relevant parameter.  There is an equally long tradition that realizes that this carries very little explanatory force.
[3] ‘R(X)’ means root of X.
[4] Chomsky uses this to derive the Postal effects discussed by Koizumi and Lasnik and Saito that indicate that accusative case marking endows an expression with scope higher than its apparent base position. Chomsky thinks that this movement really counterintuitive and so thinks that the system he develops is really terrific precisely because it gets this as a consequence.  It is worth pointing out that if case were assigned in Spec-X configurations as in the earliest MP proposals, the movement would be required as well. I am sure that Chomsky would not like this, though it is pretty clear that Spec-X structures are not nearly as unpleasant to his sensibilities as they once were given that this are the configurations where agreement licenses labeling. That said, it is not configurations that license agreement per se.
[5] Chomsky does suggest that in head raising he raised head “labels” the derived {X,Y} structure. If so, then maybe one needs to label the VP before the V raises. I am frankly unclear about all of this.
[6] A paper I wrote with Jacek Witkos on ECs suggested that in ECs there inherited features from the nominal it was related to (it started out as a kind of dummy determiner and moved) and that the agreement one sees is not then directly with the associate. This would suffice here, though I strongly doubt that Chomsky would be delighted with this fix.

Wednesday, July 9, 2014

Guest Post: The [Spec,TP]-agreement fallacy

Omer Preminger sent me this interesting post commenting on Chomsky's lecture 3 proposal that Spec-TP agreement might circumvent problems for the MLA. The gist is that Chomsky's proposal faces some well-known empirical challenges, especially evident in languages like Icelandic (what Gert Webelhuth once called the super conducting super collider of linguistics).  I hope that this generates some useful discussion, especially among those partial to Chomsky's take on the labeling issues he raises. Given that Sp-X agreement lies at the chart of Chomsky's analysis of successive cyclic movement, EPP and Fixed Subject Constraint effects, Omer's challenge needs addressing if this proposal is to fly. So, let the games begin!

*******

DISCLAIMER: None of what I am about to write draws on my own research. These are results that, in one form or another, have been around for decades.

In Part 3 of Norbert’s comments on Chomsky’s third lecture, he discussed Chomsky’s suggestion for why it is that movement can stop at (what we ruffians call) the [Spec,TP] position. Why is this a question? Because Chomsky is assuming the Minimal Labeling Algorithm (henceforth, MLA), which normally cannot assign a label to a structure {X, Y} if both X and Y are internally complex; and under the MLA, an unlabelable structure needs to be broken up via movement of one of its terms. One loophole for this (see Norbert’s discussion for some others) is if X and Y enter into some agreement relation; then, the feature (or perhaps set of features) F that has undergone agreement can serve as the label of {X, Y}. What is the F, then, that allows – and often times, forces – subjects to remain in [Spec,TP]? Chomsky’s answer: phi (i.e., the familiar set of person, number, gender/noun-class).

The point of this post is to show that this is not, and cannot be, the answer (for reasons that have been known for quite a while now). It starts with Icelandic, but as I will note at the very end, we could have perhaps made the point even based on English alone (though perhaps in a somewhat more tenuous fashion).

Icelandic, as is well known, has non-nominative subjects. These are not merely noun phrases bearing non-nominative case that have come to c-command the other noun phrases in their clause (cf. German); everything that Chomsky wants to say about subjects in English holds of these non-nominative subjects as well, save for two properties: their case (obviously), and the fact that they don’t control agreement (crucially).

So you get, e.g., sentences of the form in (1), where the finite verb agrees with the nominative object (which also passes a series of direct-object diagnostics), not with the dative subject:

(1)  SUBJ(dative)  FINITE-VERB(agr-with-obj)  OBJ(nominative)

[There are other complications, as there are bound to be – in this case, concerning what happens when OBJ is 1st/2nd person. But if everything is 3rd person, things work as shown in (1). And, importantly, even if the OBJ is 1st/2nd person, agreement is not with the person features of SUBJ (i.e., choosing a 1st/2nd person SUBJ does not make possible 1st/2nd person agreement on the verb).]

So, the short version of the story: there are subjects, that show all the subjecthood properties (e.g. landing and staying in subject position), and yet they are not what enters into agreement in phi-features with T. Not only that, but T in fact enters into overt phi agreement with something else (in this case, the nominative direct object). Tying subjecthood properties (e.g. the ability to move to and stay in [Spec,TP]) to agreement in phi-features is wrong. Fin.

But there is a slightly longer version of this story. Norbert, for one, is partial to the idea that what someone like me would call “probe-goal agreement” (as in, for example, the relation between T and the direct object in (1)) is really a movement relation, one where both LF and PF privilege the lower copy for interpretation/pronunciation, and the consequences of this movement can only be seen via the effects it has on the formal features of the landing site (TP). I have suggested we refer to this kind of movement as “interface-vacuous” movement, since the interfaces ignore its having occurred.

Suppose, then, that the OBJ in (1) has a second merge position in [Spec,TP], but is pronounced and interpreted in its lower position within the verb phrase. This second position of OBJ enters into agreement in phi-features with T, allowing all of (what we would call) TP to be labeled by these phi-features, as discussed above. Does this salvage Chomsky’s story?

The answer is “no.” That is because Icelandic is not a null-subject language; Icelandic clauses need subjects, in a way that this “interface-vacuous” movement (if it actually exists) does not seem to satisfy. To put it another way, even if OBJ has a second unpronounced and uninterpreted merge position in [Spec,TP], the facts are that this does not absolve the clause of its need to have a(nother) subject. To see why that’s a problem for Chomsky, let’s consider how the need to have a subject arises in his system. In the proposed system, the difference between a null-subject language (say, Italian) and a non-null-subject language (say, English), is in the capacity of T to serve as the label of a {T, XP} structure (say, for XP=vP). In a non-null-subject language, it cannot (“T is weak”) – and so in fact the only way to assign TP a label is to move something to [Spec,TP] (as a sister of the {T, XP} node), have it agree with {T, XP} in phi-features, and have those phi-features label the resulting complex object (what we would call “TP”). In a null-subject language, T can serve as the label of {T, XP}, and thus movement to [Spec,TP] is not required (if such movement were to nevertheless occur, the English-style story just described could still kick-in).

Continuing to adopt (for the time being) the “interface-vacuous” movement wrinkle, the OBJ in (1) has moved to [Spec,TP], agreed with {T, vP} in phi features, and thus labeled the resulting object; why does this clause still need an overt subject? Or more to the point, why is the equivalent of “arrived.PL some people.NOM” (a VS-order unaccusative with no expletive) not grammatical in Icelandic? After all, the nominative will have “interface-vacuously” moved to [Spec,TP], satisfying all apparent labeling needs.


So what has gone wrong here? My answer would be: the [Spec,TP]-agreement connection is a red herring, and this is what happens when you build your edifice on a red herring (fish are slippery!). Yes, in many languages many of things that end up in [Spec,TP] were also the things that T agreed with (or as Norbert would have it: many of the things that end up in [Spec,TP] overtly, turn out to obviate the need for another, separate thing to undergo “interface-vacuous” movement to [Spec,TP]). But taking that to be a fundamental fact about the computational system is just wrong, for there are languages where that’s just not how it works. Icelandic is one such language – but depending on your analysis of expletive-associate constructions and of Locative Inversion, English may very well be such a language, as well.