Comments

Monday, July 14, 2014

What's in a Category? [Part 1]


Norbert's most recent comments on Chomsky's lecture implicitly touched on an issue that I've been pondering ever since I realized how categories can be linked to constraints. Norbert's primary concern is the role of labels and how labeling may drive the syntactic machinery. Here's what caught my attention in his description of these ideas:
In effect, labels are how we create equivalence classes of expressions based on the basic atomic inventory. Another way of saying this is that Labeling maps a "complex" set {a,b} to either a or b, thereby putting it in the equivalence class of 'a' or 'b'. If Labels allow Select to apply to anything in the equivalence class of 'a' (and not just to 'a' alone), we can derive [structured linguistic objects] via Iteration.
Unless I'm vastly miscontruing Norbert's proposal, this is a generalization of the idea of labels as distribution classes. Linguists classify slept and killed Mary as VPs because they are interchangeable in all grammatical sentences of English. Now in Norbert's case the labels presumably aren't XPs but just lexical items, following standard ideas of Bare Phrase Structure. Let's ignore this complication for now (we'll come back to it later, I promise) and just focus on the issue that causes my panties to require some serious untwisting:
  • Are syntactic categories tied to distribution classes in some way?
  • If not, what is their contribution to the formalism?
  • What does it mean for a lexical item to be, say, a verb rather than a noun?
  • And why should we even care?

A question about feature valuation

I've been working in a "whig history" (WH) of generative grammar. A WH is a kind of rational reconstruction which, if doable, serves to reconstruct the logical development of a field of inquiry. WHs, then, are not “real” histories. Rather, they present the past as “an inevitable progression towards ever greater…enlightenment.” Real history is filled with dead ends, lucky breaks, misunderstandings, confusions, petty rivalries, and more.  WHs are not. They focus on the “successful chain of theories and experiments that led to the present-day science, while ignoring failed theories and dead ends” (see here). The value of WHs is that they expose the cumulative nature of a given trajectory of inquiry. As one sign of a real science is that it has a cumulative structure and given that many think that the history of Generative Grammar (GG) fails to have a cumulative structure, many think that this tells against the GG enterprise. However, the "many" are wrong: GG has a perfectly respectable WH and both empirically and theoretically the development has been cumulative. In a word, we've made loads of progress.  But this is not the topic for this post. What is?

As I went about reconstructing the relation between current minimalist theory and earlier GB theory, I came to appreciate just how powerful the No Tampering Condition (NTC) really is (I know I know, I should have understood this before, but dim bulb that I am, I didn't). I understand the NTC as follows: the inputs to a given grammatical operation must be preserved in the outputs of that operation. In effect, the NTC is a conservation principle that says that structure can be created but not destroyed. Replacing an expression with a trace of that expression destroys (i.e. fails to preserve) the input structure in the output and so the GB conception of traces is theoretically inadmissible in a minimalist theory that assumes the NTC (which, let me remind you is a very nice computational principle and part of most (all?) current minimalist proposals).

The NTC has many other virtues as well. For example, it derives the fact that movement rules cannot "lower" and that movement (at least within a single rooted sub-"tree") is always to a "c-commanding" position. Those of you who have listened to any of the Chomsky lectures I posted earlier will understand why I have used scare quotes above. If you don't know why and don't want to listen to the lectures, as David Pesetsky. He can tell you.  

At any rate, the NTC also suffices to derive the GB Projection Principle and the MP Extension Condition. In addition, it suffices to eliminate trace theory as a theoretical option (viz. co-indexed empty categories that are residues of movement: [e]1). Why? because traces cannot exist in the input to the derivation and so they cannot exist in the output given the NTC. Thus, given the NTC, the only way to implement the Projection Principle is via the Copy Theory. This is all very satisfying theoretically for the usual minimalist reasons. However, it also raises a question in my mind, which I would like to ask here.

Why doesn't the NTC rule out feature valuation?  One of the current grammatical operations within MP grammars is AGREE. What it does is relate two expressions (heads actually) in a Probe/Goal configuration and the goal "values" the features of the probe.  Now, the way I've understood this is that the Probe is akin to a property, something like P(x) (maybe with a lambdaish binder, but who really cares) and the goal serves to turn that 'x' into some value, so turns P(x) into P(phi) for example (if you want, via something like lambda conversion, but again who really cares). At any rate, and here's my question: doesn't this violate the NTC? After all, the input to AGREE is P(x) and the output is, e.g. P(phi). Doesn't this violate a strict version of the NTC?

Note, interestingly, feature checking per se is consistent with the NTC, as no feature changing/valuing need go on to "check" if sets of features are "compatible."  However, if I understand the valuation idea, then it is thought to go beyond mere bookkeeping. It is intended to change the feature composition of a probe based on the feature composition of the goal.  Indeed, it is precisely for this reason that phases are required to strip off the valued yet uninterpretable features before Transfer. But if AGREE changes feature matrices then it seems incompatible with the NTC.

The same line of reasoning suggests that feature lowering is also incompatible with the NTC. To wit: if features really transfer from C to T or from v to V (either by being copied from the former to the latter or actually copied from the higher to the lower and deleted from the higher) then again the NTC in its strongest form seems to be violated. 

So, my question: are theories that adopt feature valuation and feature lowering inconsistent with the NTC or not? Note, we can massage the NTC so that it does not apply to such feature "checking" operations. But then we could massage the NTC so that it does not prohibit traces. We can, after all, do anything we wish. For example, current theory stipulates that pair merge, unlike set merge, is not subject to Extension, viz. the NTC (though I think that Chomsky is not happy with this given some oblique remarks he made in lecture 3). However, if  the NTC is strictly speaking incompatible with these two operations, then it is worth knowing, as it would seem to be theoretically very consequential. For example, a good chunk of phase theory, as currently understood, depends on these operations and would we discover that they are incompatible with the NTC then this might (IMO, likely does) have consequences for Darwin's Problem.

So, all you thoroughly modern minimalists out there: what say you?

Friday, July 11, 2014

Comments on lecture 3: the finale

This is the final part of my comments on lecture 3.  The first three parts are (here, here and here). I depart from explication mode in these last comments and turn instead to a critical evaluation of what I take to be Chomsky’s main line of argument (and it is NOT empirical).  His approach to labels emerges directly from his conception of the basic operation Merge. How so? Well, there are only two “places” that MPish approaches can look to in order to ground linguistic processes, the computational system (CS) or the interface conditions (Bare Output Conditions (BOC)).  Given Chomsky’s conceptually spare understanding of Merge, it is not surprising that labeling must be understood as a BOC. I here endorse this logic and conclude that Chomsky’s modus ponens is my modus tolens. If correct, this requires us to rethink the basic operation. Here’s what I believe we should be aiming for: a conception that traces the kind of recursion we find in FL to labeling. In other words, labeling is not a BOC but intrinsic to CS; indeed the very operation that allows for the construction of SLOs. Thus, just as Merge now (though not in MPs early days) includes both phrase building and movement, the basic operation, when properly conceptualized, should also include labeling.

To motivate you to aim for such a conception, it’s worth recalling that in early MP it was considered conceptually obvious that Move and Merge were different kinds of things and that the latter was more basic and that the former was an “imperfection.” Some (including me and Chris Collins) did not buy this dichotomy suggesting that whatever process produced phrase structure (now E-Merge) should also suffice to give one move (now I-merge). In other words, that FL when it arose came fully equipped with both merge and move neither being more basic than the other. On this view, Move is not an “imperfection” at all. Chomsky’s later work endorsed this conception. He derived the same conclusion form other (arguably (though I’m not sure I would so argue) simpler) premises. I derive one moral from this little history: what looks to be conceptually obvious is a lot clearer after the fact than ex ante. Here’s a good place to mention the owl of Minerva, but I will refrain. Thus, here’s a project: rethink the basic operation in CS so that labels are intrinsic consequences. I will suggest one way of doing this below, but it is only a suggestion. What I think there are arguments for is that Chomsky’s way of including labels in FL is very problematic (this is as close as I come to saying that it’s wrong!) and misdiagnoses the relevant issues. The logic is terrific, it just starts from the wrong place. Here goes.

1.     The logic revisited and another perspective on “merge”

There are probably other issues to address if one wants to pursue Chomsky’s proposal. IMO, right now his basic idea, though suggestive, is not that well articulated. There are many technical and empirical issues that need to be ironed out. However, I doubt that this will deter those convinced by Chomsky’s conceptual arguments. So before ending I want to discuss them. And I want to make two points: first that I think that there is something right about his argument. What I mean is that if you buy Chomsky’s conception of Merge, then adding something like a labeling algorithm in CS is conceptually inelegant, if not worse. In other words, Chomsky is right in thinking that adding labeling to his conception of Merge is not a good theoretical move conceptually. And second, I want to suggest that Chomsky’s idea that projection is effectively an interface requirement, a BOC in older terminology, has things backwards.  The interfaces do not require labeled structures to do what they do. At least CI doesn’t, so far as I can tell. The syntax needs them. The interfaces do not. The two points together point to the conclusion that we need to re-think Merge. I will very briefly suggest how we might do this.

Let’s start.  First, Chomsky is making exactly the right kind of argument. As noted at the outset, Chomsky is right to question labeling as part of CS given his view that Merge is the minimal syntactic operation. His version of Merge provides unboundedly many SLOs (plus movement) all by itself. One can add projection (i.e. labeling) considerations to the rule but this addition will necessarily go beyond the conceptual minimum. Thus, Merge cannot have a labeling sub-part (as earlier versions of Merge did).  In fact, the only theoretical place for labels is the interface as the only place for anything in an MP-style account is as an interface BOC or the CS. But as labels cannot be part of CS, they must be traced to properties of the CI/SM interface. And given Chomsky’s view that the CI interface is really where all the action is, this means that labeling is primarily required for CI interpretation.  That’s the logic and it strikes me as a very very nice argument. 

Let me also add, before I pick at some of the premises of Chomsky’s argument, that lecture 3 once again illustrates what minimalist theorizing should aim for: the derivation of deep properties of FL from simple assumptions.  Lecture 3 continues the agenda from lecture 2 by aiming to explain three prominent effects: successive cyclicity, FSCs and EPP effects.  As I have stressed before in other places, these discovered effects are the glory of GG and we should evaluate any theoretical proposal by how well and how many it can explain. Indeed, that’s what theory does in any scientific discipline. In linguistics theory should explain the myriad effects we have discovered over the last 60 years of GG research. In sum, not surprisingly, and even though I am going to disagree with Chomsky’s proposal, I think that lecture 3 offers an excellent model of what theorists should be doing.

So which premise don’t I like?  I am very unconvinced that labels reflect BOCs. I do not see why CI, for example, needs labeled structures to interpret SLOs.  What is needed is structured objects (to provide compositional structure) but I don’t see that it needs labeled SLOs.  The primitives in standard accounts of semantic interpretation are things like arguments, predicates, events, proposition, operator, variable, scope, etc. Not agreeing phrases, VPs or vPs or Question Ps etc.  Thus, for example, though we need to identify the Q operator in questions to give the structure a question “meaning” and we need to determine the scope of this operator (something like its CC domain), it is not clear to me that we also need to identify a question phrase or an agreement phrase.  At least in the standard semantic accounts I am familiar with, be it Heim and Kratzer or Neo-Davidsonian, we don’t really need to know anything about the labels to interpret SLOs at CI. It’s the branching that matters, not what labels sit on the nodes.[1]

I know little about SM (I grew up in a philo dept and have never taken a phonology course (though some of my best friends are phonologists)), but from what I can gather the same seems true on the SM side. There are difference between stress in some Ns and Vs but at the higher levels, the relevant units are XPs not DPs vs VPs vs TP vs CPs etc. Indeed the general procedure in getting to phrasal phonology involves erasing headedness information. In other words, the phonology does not seem to care about labels beyond the N vs V level (i.e. the level of phonological atoms).

If this impression is accurate (and Chomsky asserts but does not illustrate why he thinks that the interfaces should care about labeled SLOs) then we can treat Chomsky’s proposal as a reductio: He is right about how the pieces must fit together given his starting assumptions, but they imply something clearly false (that labels are necessary for interface legibility) therefore there must be something wrong with Chomsky’s starting point, viz. that Merge as he understands it is the right basic operation.

I would go further. If labeling is largely irrelevant for interface interpretation (and so cannot be traced to BOCs) then labeling must be part of CS and this means that Chomsky’s conception of Merge needs reconsideration.[2] So let’s do that.[3]

What follows relies on some work I did (here). I apologize for the self-referential nature of what follows, but hey it’s the end of a very long post.

Here’s the idea: the basic CS operation consists of two parts, only one of which is language specific. The “unbounded” part is the product of a capacity for producing unboundedly big flat structures that is not peculiarly linguistic or unique to humans. Call this operation Iteration. Birds (and mice and bats and whales) do it with songs. Ants do it with path integration.  Iteration allows for the production of “beads on a string” kinds of structures and there is no limit in principle to how long/big these structures can be.

The distinctive feature of Iteration is that it could care less about bracketing. Consider an example: Addition can iterate. ((a+b)+c)+d) is the same as (a+(b+c+d)) which is the same as (a+b+c+d) etc. Brackets in iterative structures make no difference.  The same is true in path integration. What the ant does is add up all the information but the adding up needs no particular bracketing to succeed. So if the ant goes 2 ft N and then 3 ft W and then 6 feet south and then 4 ft E, it makes no difference to calculation how these bits of directional information are added together. However you do this provides the same result. Bracketing does not matter. The same is true for simple conjunction: ((a&b)&c)&d) is equivalent to (a & (b & (c&d))) which is the same as (a&b&c&d). Again brackets don’t matter. Let’s assume then that iterative procedures do not bracket. So there are two basic features of Iteration: (i) there is no upper bound to the objects it can produce (i.e. there is no upper bound on the length of the beaded string), and (ii) bracketing is irrelevant, viz. Iteration does not bracket. It’s just like beads on a string.

Here’s a little model. Assume that we treat the basic Iterative operation as the set union operation. And assume that the capacity to iterate involves being able to map atoms (but only atoms to their unit sets (e.g. a--> {a}). Let’s call this Select. Select is an operation whose domain is the lexical atoms and whose range is the unit set of that atom. Then given a lexicon we can get arbitrarily big sets using U and Select.[4] For example: If ‘a’, ‘b’ and ‘c’ are atoms, then we can form {a} U {b} U {c} to give us {a,b,c}. And so forth. Arbitrarily big unstructured sets.

Clearly, what we have in FL cannot just be Iteration (ie. U plus Select). After all we get SLOs. Question: what if added to Iteration would yield SLOs? I suggest the capacity to Select the outputs of Iteration. More particularly, let’s assume the little model above. How might we get structured sets? By allowing the output to Iteration to be the input to Select. So, if {a,b} has been formed (viz. {a} U {b}-> {a,b}) and Select applies to {a,b} then out comes the structured SLO {{a,b}, c} (viz. {{a,b}} U {c} -> {{a,b},c}. One can also get an analogue of I-merge: select {{a,b},c} (i.e. {{{a,b},c}}, select c (i.e. {c}), Union the sets (i.e. {{{a,b},c}}} U {c}) and out comes {c, {{a,b},c}}.  So if we can extend the domain of Select to include outputs of the union operation then we can get use Iteration to deliver unboundedly many SLOs.

The important question then is what licenses extending the domain of Select to the outputs of Union?  Labeling. Labeling is just the name we give for closing Iteration in the domain of the lexical atoms.[5]  In effect, labels are how we create equivalence classes of expressions based on the basic atomic inventory. Another way of saying this is that Labeling maps a “complex” set {a,b}to either a or b, thereby putting it in the equivalence class of ‘a’ or ‘b’. If Labels allow Select to apply to anything in the equivalence class of ‘a’ (and not just to ‘a’ alone), we can derive SLOs via Iteration.[6]

Ok, on this view, what’s the “miracle”? For Chomsky, the miracle is Merge. On the view above, the miracle is Label, the operation that closes Iteration in the domain of the lexical atoms. Label effectively maps any complex set into the equivalence class of one of its members (creating a modular structure) and then treats these as syntactically indistinguishable from the elements that head them (as happens in modular arithmetic (i.e. ‘1’ and ‘13’ and ‘12’ and ‘24’ are computationally identical in clock arithmetic). Effectively the lexicon serves as the modulus with labels mapping complexes of atoms to single atoms bringing them within the purview of Select.[7]

Note that this “story” presupposes that Iteration pre-exists the capacity to generate SLOs. The U(nion) operation is cognitively general as is Select which allows U to form arbitrarily large unstructured objects. Thus, Iteration is not species specific (which is why birds, ants and whales can do it). What is species specific is Label, the operation that closes U in the domain of the lexical atoms and this what leads to a modular combinatoric system (viz. allows an operation defined over lexical atoms to also operate over non-atomic structures). Note that if this is right, then labels are intrinsic to CS; without it there are no SLOs for without it U, the sole combination operation, cannot derive sets embedded within sets (i.e. hierarchy).

The toy account above has other pleasant features. For example, the operation that combines things is the very general U operation. There are few conceivably simpler operations.  The products of U produce objects that necessarily obey NTC, Inclusiveness and produce copies under “I-merge.” Indeed, this proposal treats U as the main combinatoric operation (the operation that constructs sets containing more than one member). And if combination is effectively U, then phrases must be sets (i.e. U is a set theoretic operation so the objects it applies to must be sets). And that’s why the products of this combinatoric operation respect the NTC, Inclusiveness and produce “copies.”[8]

Let’s now get back to the main point: on this reconstruction, hierarchical recursion is the product of Iterate plus Label. To be “mergeable” you need a label for only then are you in the range of Select and U.  So, labels are a big deal and intrinsic to CS. Moreover, this makes labeling facts CS facts, not BOC facts.

This is not the place to argue that this conception is superior to Chomsky’s. My only point is that if my reservations above about treating Labels as BOCs is correct, then we need to find a way of understanding labels as intrinsic to the syntax, which in turn requires reanalyzing the minimal basic operation (i.e. rethinking the “miracle”).

IMO, the situation regarding projection is not unlike what took place when minimalists rethought the Merge/Move distinction central to early Minimalism (see the “Black Book”).  Movement was taken to be an “imperfection.” Rethinking the basic operation allowed for the unification of E and I-merge (i.e. Gs with SLOs would also have displacement). I think we need to do the same thing for Labeling. We need to find a way to make labels intrinsic features of SLOs, labels being necessary for building structure and displacing them.  Chomsky’s views on projection don’t do this. They start from the assumption that Labels are BOCs. If this strikes you as unconvincing as it does me, then we need to rethink the basic minimal operation.

That’s it. These comments are way too long. But that’s what happens when you try and think about what Chomsky is up to. Agree or not, it’s endlessly fascinating.





[1] Edwin Williams once noted that syntactic categories cross cut semantic ones. Predicative nominals have the same syntactic structure as argument nominal, though they differ a lot semantically. I think Edwin’s point is more generally correct. And if it is, then syntactic labels contribute very little (if anything) to CI interpretation.
[2] Though I won’t go into this here, there is plenty of apparent evidence that Gs care about labeled SLOs. So languages target different categories for movement and deletion. Moreover there are structure preservation principles that need explaining: XPs move to Max P positions, X’s don’t move and heads target heads.  In a non-labeling theory, it is still unclear why phrases move at all. And the Pied Piping mantra is getting a bit thin after 20 years.  So, not only is there little evidence that the interfaces care about labels, there is non-negligible evidence that CS does. If correct, this strengthens the argument against Chomsky’s approach to projection.
[3] One more aside: I am always wary of explanations that concentrate in interface requirements. We know next to nothing about the interfaces, especially CI, so stories that build on these requirements always seem to me to have a “just so” character.  So, though the logic Chomsky deploys is fine, the premise he needs about BOCs will have little independent motivation. This does not make the claims wrong, but it does make the arguments weak.
[4] If we distinguish selections from the “lexicon” so that two selections of a are distinguished (a vs a’), we can get unboundedly big sets. Bags can be substituted for sets if you don’t like distinguishing different selections of atoms.
[5] Chomsky flirted with this idea in his earlier discussion of “edge features” (EF). As yourself where EFs came from? They were taken as endemic to lexical atoms. It is natural to assume that complexes of such atoms inherited EFs from their atomic parts. Sound familiar?  EFs, Labels? Hmm.  The cognoscenti know that Chomsky abandoned this way of looking at things. This is an attempt to revive this idea by putting it on what might be a more principled basis.
[6] For those who care, this is a grammatical analogue of “clock/modular arithmetic” (see here).
[7] Here’s how Wikipedia describes the process:
In mathematics, modular arithmetic is a system of arithmetic for integers, where numbers "wrap around" upon reaching a certain value—the modulus. The modern approach to modular arithmetic was developed by Carl Friedrich Gauss in his book Disquisitiones Arithmeticae, published in 1801.
A familiar use of modular arithmetic is in the 12-hour clock, in which the day is divided into two 12-hour periods. If the time is 7:00 now, then 8 hours later it will be 3:00. Usual addition would suggest that the later time should be 7 + 8 = 15, but this is not the answer because clock time "wraps around" every 12 hours; in 12-hour time, there is no "15 o'clock". Likewise, if the clock starts at 12:00 (noon) and 21 hours elapse, then the time will be 9:00 the next day, rather than 33:00. Since the hour number starts over after it reaches 12, this is arithmetic modulo 12. 12 is congruent not only to 12 itself, but also to 0, so the time called "12:00" could also be called "0:00", since 12 is congruent to 0 modulo 12.
[8] Note, that Labels allow one to dispense with Probe/Goal architectures as heads are now visible in “Spec-head” configurations. Not that there is anything “special” about Specs (as opposed to complements or anything else). It’s just that given labels, XPs can combine with YPs even after “first” merge and still allow their heads to “see” each other. This, in fact, is what endocentricity was made to do: put expressions that are not simple heads “next to” each other. And they will be adjacent whether the elements combined are complements or specifiers. Chomsky is right that there is nothing “special” about specifiers. But that’s just as true of complements.

Thursday, July 10, 2014

A crapification of academic life

I have several bees in my bonnet. One is the systematic degradation of academic life. This can be traced to two different  causes: (i) the diminution of resources to the university system and (ii) the progressive restriction on funding for fundamental research. Here are two papers that say something about each.

This one reports on a recent PNAS study on funding of biomedical research. It observes that things are bad and getting worse. More people chasing fewer dollars has led to short-termism. One doesn't have to be a vulgar economist to think that incentives matter and that this kind of situation can lead to cutting corners and focus on boring hygienic problems rather than risky fundamental ones.

Now, linguistics is not as well endowed as bio-med research, but I believe that we face analogous problems. See in particular the discussion of PhD bloat. I would be the last to suggest that grad programs start cutting down on admissions because the chances of landing a job are getting thinner. However, I think that we might owe new admits a stern warning (akin to what one finds on cigarette packages) that there is no guarantee of a career in ling after graduation. At any rate, the sociology in ling is similar if not quite identical to what currently finds in bio-science if this PNAS report is anything to go by.

This second paper discusses the first problem alluded to; the decline of the US university and it traces many the problems to the explosive growth of managerialism. It seems that lots and lots of money is going to administrative salaries. Put bluntly, the people at the top find that increasing their salaries and perks is easy to do.  In the US, this growth has coincided with the idea that everything is, and should be run as, a business. Students are customers, faculty are workers, and the real brains are the administrators that keep everything growing. I have even heard of administrators congratulating themselves on "herding" the faculty "cats" and that this guidance was critical to managing good research. Just another assembly line, albeit one with quirky workers.  At any rate, lots of money is being redirected to administration and this has had a baleful effect on university life. I personally do not see it getting any better very soon. Once in place, bureaucracies are self-sustaining and tend to grow. Oh well.

Wednesday, July 9, 2014

Guest Post: The [Spec,TP]-agreement fallacy

Omer Preminger sent me this interesting post commenting on Chomsky's lecture 3 proposal that Spec-TP agreement might circumvent problems for the MLA. The gist is that Chomsky's proposal faces some well-known empirical challenges, especially evident in languages like Icelandic (what Gert Webelhuth once called the super conducting super collider of linguistics).  I hope that this generates some useful discussion, especially among those partial to Chomsky's take on the labeling issues he raises. Given that Sp-X agreement lies at the chart of Chomsky's analysis of successive cyclic movement, EPP and Fixed Subject Constraint effects, Omer's challenge needs addressing if this proposal is to fly. So, let the games begin!

*******

DISCLAIMER: None of what I am about to write draws on my own research. These are results that, in one form or another, have been around for decades.

In Part 3 of Norbert’s comments on Chomsky’s third lecture, he discussed Chomsky’s suggestion for why it is that movement can stop at (what we ruffians call) the [Spec,TP] position. Why is this a question? Because Chomsky is assuming the Minimal Labeling Algorithm (henceforth, MLA), which normally cannot assign a label to a structure {X, Y} if both X and Y are internally complex; and under the MLA, an unlabelable structure needs to be broken up via movement of one of its terms. One loophole for this (see Norbert’s discussion for some others) is if X and Y enter into some agreement relation; then, the feature (or perhaps set of features) F that has undergone agreement can serve as the label of {X, Y}. What is the F, then, that allows – and often times, forces – subjects to remain in [Spec,TP]? Chomsky’s answer: phi (i.e., the familiar set of person, number, gender/noun-class).

The point of this post is to show that this is not, and cannot be, the answer (for reasons that have been known for quite a while now). It starts with Icelandic, but as I will note at the very end, we could have perhaps made the point even based on English alone (though perhaps in a somewhat more tenuous fashion).

Icelandic, as is well known, has non-nominative subjects. These are not merely noun phrases bearing non-nominative case that have come to c-command the other noun phrases in their clause (cf. German); everything that Chomsky wants to say about subjects in English holds of these non-nominative subjects as well, save for two properties: their case (obviously), and the fact that they don’t control agreement (crucially).

So you get, e.g., sentences of the form in (1), where the finite verb agrees with the nominative object (which also passes a series of direct-object diagnostics), not with the dative subject:

(1)  SUBJ(dative)  FINITE-VERB(agr-with-obj)  OBJ(nominative)

[There are other complications, as there are bound to be – in this case, concerning what happens when OBJ is 1st/2nd person. But if everything is 3rd person, things work as shown in (1). And, importantly, even if the OBJ is 1st/2nd person, agreement is not with the person features of SUBJ (i.e., choosing a 1st/2nd person SUBJ does not make possible 1st/2nd person agreement on the verb).]

So, the short version of the story: there are subjects, that show all the subjecthood properties (e.g. landing and staying in subject position), and yet they are not what enters into agreement in phi-features with T. Not only that, but T in fact enters into overt phi agreement with something else (in this case, the nominative direct object). Tying subjecthood properties (e.g. the ability to move to and stay in [Spec,TP]) to agreement in phi-features is wrong. Fin.

But there is a slightly longer version of this story. Norbert, for one, is partial to the idea that what someone like me would call “probe-goal agreement” (as in, for example, the relation between T and the direct object in (1)) is really a movement relation, one where both LF and PF privilege the lower copy for interpretation/pronunciation, and the consequences of this movement can only be seen via the effects it has on the formal features of the landing site (TP). I have suggested we refer to this kind of movement as “interface-vacuous” movement, since the interfaces ignore its having occurred.

Suppose, then, that the OBJ in (1) has a second merge position in [Spec,TP], but is pronounced and interpreted in its lower position within the verb phrase. This second position of OBJ enters into agreement in phi-features with T, allowing all of (what we would call) TP to be labeled by these phi-features, as discussed above. Does this salvage Chomsky’s story?

The answer is “no.” That is because Icelandic is not a null-subject language; Icelandic clauses need subjects, in a way that this “interface-vacuous” movement (if it actually exists) does not seem to satisfy. To put it another way, even if OBJ has a second unpronounced and uninterpreted merge position in [Spec,TP], the facts are that this does not absolve the clause of its need to have a(nother) subject. To see why that’s a problem for Chomsky, let’s consider how the need to have a subject arises in his system. In the proposed system, the difference between a null-subject language (say, Italian) and a non-null-subject language (say, English), is in the capacity of T to serve as the label of a {T, XP} structure (say, for XP=vP). In a non-null-subject language, it cannot (“T is weak”) – and so in fact the only way to assign TP a label is to move something to [Spec,TP] (as a sister of the {T, XP} node), have it agree with {T, XP} in phi-features, and have those phi-features label the resulting complex object (what we would call “TP”). In a null-subject language, T can serve as the label of {T, XP}, and thus movement to [Spec,TP] is not required (if such movement were to nevertheless occur, the English-style story just described could still kick-in).

Continuing to adopt (for the time being) the “interface-vacuous” movement wrinkle, the OBJ in (1) has moved to [Spec,TP], agreed with {T, vP} in phi features, and thus labeled the resulting object; why does this clause still need an overt subject? Or more to the point, why is the equivalent of “arrived.PL some people.NOM” (a VS-order unaccusative with no expletive) not grammatical in Icelandic? After all, the nominative will have “interface-vacuously” moved to [Spec,TP], satisfying all apparent labeling needs.


So what has gone wrong here? My answer would be: the [Spec,TP]-agreement connection is a red herring, and this is what happens when you build your edifice on a red herring (fish are slippery!). Yes, in many languages many of things that end up in [Spec,TP] were also the things that T agreed with (or as Norbert would have it: many of the things that end up in [Spec,TP] overtly, turn out to obviate the need for another, separate thing to undergo “interface-vacuous” movement to [Spec,TP]). But taking that to be a fundamental fact about the computational system is just wrong, for there are languages where that’s just not how it works. Icelandic is one such language – but depending on your analysis of expletive-associate constructions and of Locative Inversion, English may very well be such a language, as well.



Monday, July 7, 2014

Comments on lecture 3, part III

Here’s part 3. First two are here and here. Chomsky’s lectures are here.

1.     The Halting Problem

As usual, let’s assume that the earlier objections don’t derail the project (which clearly they don’t) and let’s keep following Chomsky’s logic.  Here is one more problem that Chomsky addresses. We know why XPs move and why they can stop. The next question is why they must stop.  Rizzi called this the “halting problem” (no, it’s not related to the real halting problem, though it does sound like it might be eh?). The issue is why a WH (or a DP in an agreeing Spec) does not move any further once there. Chomsky attributes this to uninterpretability of the resulting structure at the CI interface. Let’s look at the details.

The relevant structure is illustrated by (4). This illustrates that criterial agreement “freezes” the DP preventing further movement.  Why?
                        (4)  *What does John wonder [what [C+Q [ Bill ate]]]
Chomsky suggests that (4) is not syntactically illicit but is illegible at CI. Why?  Chomsky does not distinguish between the +Q-C in Wh questions and the one in Yes/No questions.  Empirically, this is a necessary assumption given the observation that a verb can take an embedded WH question as complement iff it can also take a Y/N question. This only makes sense if the Qs in both are the same. If this is so, we can ask why (4) cannot be interpreted like (5):
                        (5) What does John wonder if/whether Bill ate
This has the same structure as (4) but for if/whether. Note that (4) cannot be interpreted as a degraded version of (5). What is less clear to me is why not? There is a +Q-C there, just as in (4). So what’s the problem? We know from matrix clauses that Y/N questions do not require an overt if/whether to license the Y/N interpretation. So, the embedded Q should not need an overt WH morpheme to license the interpretation. Like I said, Chomsky asserts that this derivation has CI problems, but I really don’t see why.[1]

Why does Chomsky attribute the problem with (4) to CI interpretation? He has few other options. Though it is true that agreement suffices to disambiguate the application of MLA, movement will do so as well (i.e. movement will not cause problems for the MLA). But I don’t think that Chomsky wants to say that what blocks further movement should be traced to the details of uF valuation.  Though he could offer the following story: The WH can have its features valued in CP (or via Agree before moving there) and when features can be valued they must be. If so, the WH in the embedded position must have its features valued. If we further assume that feature values cannot be over-written, then if the WH moves further it cannot Agree with the matrix WH and so the matrix {XP, YP} cannot be labeled (recall, we need more than mere feature identity, we need feature agreement).  Note that this relies on some substantive details about feature valuation (e.g. it’s not optional, it’s indelible, uFs cannot “stack”). Perhaps these details follow from the minimal theory feature checking. I leave this to those with a better sense of what is minimal here. However, I suspect that Chomsky does not want to tie his theories to the specifics of feature checking algorithms (I don’t think that I would) and if not, he needs a CI interface story.[2] Whatever the upshot of all of this, we should recall that examples like (4) are orthogonal to the MLA. Chomsky’s CI interface account is there to plug a hole in his story. It does not follow from it and it should be fine if something else explained (4).

So, the MLA has a story for successive cyclic movement, though it is unclear that it is much different from the standard story in the literature, IMO. When all the details are considered the relevant accounts all assume that without agreement in “Spec XP positions,” Gs require movement and with agreement in Spec XP positions Gs forbid movement. The virtues of the MLA are not empirical, but theoretical (viz. that it purportedly follows from the minimal assumptions regarding the algorithm that provides labels in endocentric configurations that the interfaces can read). Just how simple these assumptions are I leave for you to decide. IMO, they ain’t that minimal or simple. But this is partly a matter of taste, I think. I return to a broader consideration of Chomsky’s system towards the end of this exegesis.



            5. Fixed Subject Condition Effects (ECP) and the EPP

Chomsky argues that the MLA is able to unify two further well-known effects given reasonable ancillary assumptions. The two are the subject/object asymmetry in the ECP and the EPP illustrations in (6):
                        (6)       a. Which man did you wonder if Bill liked
                                    b. *Which man did you wonder if liked Bill
                                    c. (*there) arrived a beautiful woman
                                    d. arrivato una bella ragazza
                                    e. (*I) ate a pizza
                                    f.  (I) ho mangiato una pizza

As Chomsky notes (under prodding from David P), the EPP fact that interests Chomsky is the one involving (null) expletives (6c,d), not null subjects together with referential null pronouns (6e,f). It’s not actually clear to me why Chomsky makes this distinction as the account he gives seems to cover both cases. I suspect that the issue is empirical, as I will explain below.

Here’s the account of the EPP.  Chomsky assumes that English T0 is weak (careful!). He treats them as having the same labeling powers as lexical roots (viz. they can’t do it).  This is true even after the tense and agreement features lower from C0 onto T0.  So how is a label assigned by the MLA? Via agreement with a subject in Spec TP. The resultant phi features provide the label that weak T cannot provide on its own. So in (6c), without there the past tense T0 head by itself cannot provide a label. If there is inserted, however, agreement between there and the “T’” suffices to license a phi-label. 

This story raises two questions. First, why is agreement necessary at all? Why can’t the expletive alone label the structure, just like v,n,a suffice to label the structures in e.g. {n, Root} configurations?  This would be empirically awkward, but it is not clear what prevents this theoretically, given Chomsky’s assumptions.  Second, it’s odd that the expletive suffices to license agreement given the standard assumption (Chomsky assumes this too) that in expletive constructions it’s the associate that determines the agreement features on T. But if T agrees not with there but with a beautiful woman then how does the MLA explain these EPP effects? It must be that there agrees with T0 in a way that the associate cannot. What way is that? Unless we are told, it looks like we are again assuming that the EPP is basically the “I-need-a-specifier” condition.

What on Chomsky’s account explains the English contrast with Italian? It assumes that in Italian T0 is strong and thus suffices to license a label on T0 without a DP in its Spec. Thus, Italian T0 is like n,v,a in English.  This account eschews null expletives. Rather the Italian TP in (6d) is a simple {Y, XP} structure with Y being the strong T0.[3] 

The explanation Chomsky gives easily extends to explain the unacceptability of (6e). T0 is weak and a lexical subject is needed to license the labeling via the MLA.  The problem arises with the Italian examples. Here’s what I mean. There is good reason to think that in cases like (6f) there is a pronominal like element in the Spec TP position. The reason is that such sentences are understood as having thematic subjects and these null pronouns care bindable.  If there is nothing there at all, this is hard to understand. Ok, so say there is a null pronoun (aka pro) there.  This is ok for Chomsky so long as we assume that this pro can agree with T0 and this agreement is what affords the label. If we assume this, then pro has phi-features.

Now to English: if Italian pro has phi-features then the minimal assumption is that English pro does too. But then why is (6e) unacceptable?  It seems that what English needs is not merely an agreeing element in Spec XP, but an overt agreeing element. But this seems to obviate the need for Chomsky’s more elaborate assumptions concerning the weak status of T0 in English.  At the least, it suggests that Weak/Strong here has more to do with the SM system than with the CI system. The problem is that it is not clear what this has to do with labeling? 

Are phrase labels required for SM interpretation? Maybe, though specific values seem not to be. What I mean is that identifying that something is an XP might be important for phrasal phonology. But distinguishing VPs from TPs from CPs is not obviously relevant. The question then is whether one needs a labeling algorithm to determine whether something is an XP. Can one identify a structure as XP in the absence of identifying a particular head. It would seem that one can do so easily at least in the relevant cases: any {XP,YP} configuration will be an MaxP for the purposes of SM (as will any {X, YP}. So it would seem that for these purposes MLA is not required, at least conceptually.  But if this is so, then the English/Italian contrast becomes not a property of the MLA but a fact about TPs in English requiring overt subjects in contrast to those in Italian. Why? Well because.  Does the MLA provide a better story? Not so far as I can tell. It all comes down to an idiosyncratic difference between Italian and English and embedding this difference in MLA technology does not appear to do much work.

However, this might be unfair. Chomsky argues that the same account can explain Fixed Subject Constraint (FSC) effects (originally discovered by Perlmutter in his thesis, I believe, and tied together with the availability of pro in the relevant G). Chomsky ties them together as well. Recall, that in English T0 is weak. If so, we need something in “Spec TP” to allow the MLA to label the structure.  That serves to block A’-movement (and DP movement as well, one should note[4]). Note the copy left behind will not suffice as it is part of a chain with links outside the “TP” domain. So the subject WH cannot move at pains of not labeling the TP and causing interface interpretation problems.

Chomsky observes that this predicts that in Italian, where T0 is strong, we should not find FSCs.  And this is plausibly correct (but see below).

In sum, Chomsky argues that the MLA serves to unify EPP and ECP (more accurately FSC) effects and that is another good argument in its favor, in addition to the conceptual ones he uses to motivate the elimination of labeling in the CS.[5]

Here are some potential problems with this analysis: First, though unacceptable, sentences like (6b) are hardly uninterpretable.[6] Indeed, these cases are a bit like the student seems sleeping. These have perfectly obvious CI interpretations and so do examples like (6b). It’s not clear why if they cannot be interpreted at the CI interface. 

Second, we know that FSCs appear in non-interrogatives as well. It’s known as the that-t effect. However, it is also well known that the unacceptability of these seems to vary across speakers. This would predict that for such speakers T0 is strong. But this further predicts that they should find sentences like arrived a nice boy/entered a well dressed dog perfectly acceptable. I have my doubts, but it’s worth looking to find out. 

Third, it’s not clear to me why the indicated reasoning doesn’t block movement from subjects altogether. Why is (7) fine?
                        (7) Which man do you believe t saw Mary
How does the MLA label the embedded TP if the Wh moves?  In other words, why is moving the subject bad if there is something overtly in C but fine if there isn’t?  Chomsky’s proposal does not obviously distinguish these cases. In fact, where Chomsky’s story differs from the traditional one is in tracing FSCs to something about “SpecT-T” relations. The standard accounts tie them to “C-Spec T” relations.[7] By taking the “Spec T-T” relation as central, it’s quite unclear how to account for the obviation of FSCs once the C is deleted.[8]

Fourth, as David Pesetsky noted during the lecture, Chomsky really misdescribes the Italian data.  On his story, it should be possible to extract a WH from “Spec T” position in Italian because Italian T0 is strong.  However, this seems to be incorrect.  In cases where the morphology indicates where the subject has moved from, we find that we cannot WH move from Spec T (this is discussed in a great paper by Brandi and Cordin (here)) Now, one might say that in these dialects T0 is weak, and that might be true. But, one would then also expect no analogues of (6d) in these dialects, and this seems to be incorrect (see Brandi and Cordin 115: (13)/(14)). Of course, appearances here may be deceiving. So, let’s chalk this up as another puzzle.

In sum, Chomsky offers some intriguing connections between the MLA and two more well known FL effects, the FSC and the EPP.  IMO, the analyses are at best suggestive.  There are many loose ends. Chomsky is aware of this and seems not to really be bothered for in his opinion the strength of the proposal lies in its conceptual simplicity. The empirical benefits are a bonus (maybe even a big one) but the problems are tolerable for the proposal depends less on their viability than on the fact that in Chomsky’s opinion, the MLA is the conceptually optimal account of labeling given a the minimal basic operation Merge.  In the last part, I return to consider the conceptual lay of the land once again.





[1] Note that a similar problem does not occur with DPs in “Spec TP” positions. In such cases, TP will be in the complement domain of the C phase head. Thus, before it can move out of the CP, Transfer will remove it from the purview of the computational system. Note that this requires that Cs be strong phases and that T must “inherit” its features from C, otherwise we could generate a TP either with a weak C or no C and then Transfer would not serve to prevent a hyper-raising derivation. Note that such derivations appear to exist (i.e. some Gs allow hyper-raising). Consider this another puzzle for the analysis.
[2] Observe that for this kind of story to work, we need to assume that it is moving WH that has uFs, contrary to the assumption that uFs are limited to phase heads and are checked in Probe/goal configurations.
Note too that this is effectively a greed based account: no movement without feature checking. This makes a lot of sense if agreement takes place in {XP,YP} configurations. Once the features of an XP are valued in {XP, YP} no feature checking is possible and so things stay put. Thus greedy movement suffices to block this. The problem, of course, is that it is not clear how conceptually necessary it is for I-merge to be greedy and why I-merge would be greedy but E-merge would not be (for the cognoscenti, I am trying to develop a slippery slope argument for Q-feature checking: if that were so, then both E and I merge would be greedy).
[3] Is it worth noting that Chomsky’s assumptions regarding T0 make it hard to see why it exists at all. In English, it is indistinguishable from AGR heads: it has virtually no properties of its own and the properties it does have are extremely idiosyncratic.  T has enjoyed a weird position in GG for a very long time. It’s odd within the Barrier’s framework and is no less odd within this one. Again, the odder its properties the greater the challenge for DP concerns.
[4] This seems to suggest that either non-finite T is strong (otherwise successive cyclic DP movement will be prohibited) or that non-finite “TP” doesn’t require a label. Either assumption strikes me as strange. Another puzzle? Another possibility is that non-finite TP is not subject to the EPP as Castillo, Drury and Grohmann as well as Epstein and Seeley have proposed a while back.
[5] For the record, the analysis does not address ECP’s argument/adjunct asymmetries.
[6] Indeed, anyone who has taught undergrad syntax 2 will have encountered speakers that find these kinds of sentences to be only marginally unacceptable. Bleive me, they do exist.
[7] Actually this is true of Rizzi’s proposal and the one in Aoun et al. The idea was that C was able to license the trace in subject position in some cases but not in others. Pesetsky and Torrego develops a version of this account: with movement of Nominative WH to Spec C obviating the need for T to C and that being the morphological reflex of T to C in embedded contexts. In both however the relation of interest in FSCs is that between the C and the expression in Spec TP.
[8] The obviation effects go beyond the removal of an overt C. We find them attenuated in cases where an adverb has been fronted:
                        (i) Who do you think that sooner or later will solve the problem
However, why this should be so on any of these accounts is not particularly clear.