Comments

Showing posts with label ECP/Fixed Subject Constraints. Show all posts
Showing posts with label ECP/Fixed Subject Constraints. Show all posts

Monday, August 4, 2014

Final comments on lecture 4

This ends the comments (here and here) on lecture 4.

The logic used to account for the EPP also covers Fixed Subject Condition effects (FSC).[1] Consider (2’) again:

(2’) *Who1 did John say that t1 saw Mary

The T needs to be labeled. If who moves there will be nothing in the Spec-T to agree with and so labeling will fail.  That’s Chomsky’s story. The obvious problem with this account is the absence of FSCs if there is no overt complementizer (I pointed to this in the comments to lecture 3). Chomsky addresses this problem here. He proposes that a deleted C is no longer a phase. More exactly, to delete a C you must transfer the feature that says that C is a phase to T. In effect, C deletion makes T the phase head. So not only do we lower phi and tense features form C to T, but phase-headedness as well.  This now ties FSCs to the presence of absence of an overt C.[2]

Observation 1: This story requires that that is deleted rather than not present at all. Were it never present, C could not transfer its features to T, and T has not features of its own (more below). Thus, to make this work, we need deletion operations in the syntax. A question that arises is how similar the operation deleting that is to more run of the mill ellipsis operations. The latter are generally treated as simply dephoneticization processes. This will not suffice here. It must be that getting rid of phonetic content requires that all features of C lower to T. For those with long enough memories, this smells a little of the old notion of “L-contains.”  At any rate, it’s worth observing that C deletion is not simply quieting the phonetics.

Observation 2: there are well known variations regarding that-t effects dialectally in English. This suggests that deletion might be sufficient for transferring all of Cs features to T but it is not necessary. So, contrary to what Chomsky suggests, the explanation requires that we say something “special” about these FSCs in English. IMO, things are a little worse than this. As I mentioned in earlier comments, many speakers seem to allow violations of the FSC in English even with if/whether in the C position (or at least so report a third of my syntax undergrads). There is no problem accommodating this by allowing feature lowering as Chomsky suggests. But this is now decoupled from phonetic articulation. Unfortunately, in the relevant dialects, deleting a that does not license null subjects, which one might have expected (4).  Or more exactly, why doesn’t lowering all the features of C to T serve to strengthen T? Note that Cs can label just fine without anything in Spec-C helping them along. Given this, why shouldn’t lowering all the features of C onto T (including the “phase head feature” see below) not make T as independent as C?  Dunno, but it doesn’t.
           
(1)  *John thinks (that) is a man here

In other words, the EPP and FSC do not really swing together, though they should if they were truly unified, one might suppose. 

Let’s put these questions to one side and continue. Chomsky then asks how we can get (2):

(2)  Who do you think t is kissing Sue

How do we label the lower “TP” if who moves. Chomsky says that it is labeled before C deletes and labels cannot be deleted. In other words, labels are indelible (think Lasnik and Saito).  Now, Chomsky really doesn’t like this way of putting things. What he wants to say is that CSs have phase sized memories (i.e. CS can recall what operations have taken place within a phase). In other words, there is phase level memory for all syntactic operations.  By lowering the phase-hood to T from C, the next higher v* phase can “remember” that the lower T was given a label via agree and so movement of the DP in Spec T is ok.  So, it is not that the labeling is indelible, but that all operations that happen within a phase are recollected in that phase. So, once labeled, memory tells us that it is always labeled. 

This emphasizes the computational aspects of phase theory.  What’s important about phases is that they reduce memory demands of a computation. The reverse of this is that it allows some memory of previous operations to be retained.  This is quite definitely not a conceptual argument. There is no conceptual motivation for these assumptions. The motivations are computational. The question becomes how we bring information forward in a derivation, how forgetting can be computationally efficient etc.  Phases and the properties Chomsky relies on here are entirely of this variety.

Comment: Two things: this is very Barrerish in spirit. Phases are no longer fixed, but change in the course of the derivation (as Den Dikken was the first to propose). And call it what you want, indelibility is back.  Moreover, just as in Barriers, T has a strange role in this system. It is not an inherent phase but can become one by inheritance. Sound familiar?[3] To me, this all has the feeling of a Rube Goldberg device, but this is partly a matter of taste. Some might think Barriers elegant. Go figure.

Let me make my unease clearer. It seems that Chomsky is not that happy with T and its various special properties within his system (again, just like T in Barriers).  He, in fact, proposes, that T has no properties of its own. It’s just there to receive properties from C. This makes T very similar to Agr in older MP, and recall that Chomsky argued that grammar internal formatives like Agr are to be eschewed. They cause DP headaches, as do any grammar internal formatives (viz. we need to explain how they and their properties got into FL).  Worse, IMO, the special properties of T are critical in Chomsky’s explanations of the EPP and FSC effects. But, this strikes me as a non-trivial problem for his account. Why doesn’t explaining these special properties of FL in terms of special properties of T not amount to re-description rather than explanation? One of the salutary effects of MP has been to warn us about confusing the two. However, the more T is special, the more accounts of Spec-T effects (EPP and FSCs) are weakened. And from what I can tell, Ts special properties are critical in deriving the results Chomsky obtains.

There are other Barrier like resonances here. Recall that in the Barriers framework, VP was never a barrier. Why? Because we could always adjoin to it and thereby void its barrierhood.  In the present story, there is also a big asymmetry between C and v*. The latter never displays FSC or ECP effects. Why not? Because v* always looses its phase property. How? Because Chomsky assumes that when the root raises to v*, v* gets buried inside the raised root and thereby looses its phase properties.  Though Chomsky does not discuss this, it raises the question of the role of v* in a phase based theory if it always looses its phase property. Do v*s no longer induce PIC effects?  It would be very odd to assume that the lower copy of the root inherited the phase property, like T inherits it from C. After all, this is the tail of a head chain and tails are generally grammatically inert.  So, one conclusion could be that v* never induces PIC effects on this revised account. Of course, once again, we can add technical fixes to obviate this conclusion, but, at least for me, their motivations will be empirical not conceptual.  This is not bad, but it does do some damage to the SMT line of reasoning Chomsky cherishes.

So, v* gets special treatment because of the properties of roots and T gets special treatment because its just, well, odd. The story may hang together, but it is hardly conceptually pretty, at least from where I sit.

Question: When R(V) raises to v* what happens to the phase property of v*? I assume that it is eliminated and this is why there are no EPP/FSC effects. Ok, does R(V) inherit the phase-hood property? This seems unlikely, as it is the tail of the head chain, but maybe.  If not, is the phase-hood of v* simply voided and what we are left with is one big C-phase?  If so, how does this enhance computational efficiency?[4]  

This post is already way too long. Let me end with three more highlights and maybe a remark or two.

Lecture 4 drops the idea that all grammatical action takes place at the phase head. In particular he allows I-merge to apply completely independently of any relationship to C. He does this to get rid of the counter-cyclic movement he needed in PoP. Recall counter-cyclic movement violates the NTC, which is a part of the conceptually best version of Merge. Chomsky really didn’t want to allow it and he eliminates here at the cost of abandoning the assumption that all grammatical operations are with the phase head. 

A consequence of this is that Chomsky has to abandon his explanation from PoP as to why Gs raise T to C but don’t raise DPs to C instead given that they are equally close if SLOs are without labels. You may recall, that he accounted for this by moving the subject to Spec T counter-cyclically in derivations that applied “all at once” at the phase level. 

I’m glad that Chomsky drops this now for two reasons. First, I never liked his earlier explanation (indeed I was part of a trio arguing that it didn’t work and was the wrong way to proceed (here)). Second, I never understood what “all at once” derivations meant. Ever.  I pretended to and taught it, but never got it.  Now I don’t have to, it seems. Curiously, Chomsky not only drops this analysis but states that it was always “artificial.” Yup.

Chomsky also finally abandons the last residues of Greed based MP theories. Merge is free. If you follow him here, you can stop worrying about what motivates this or that movement. Note, this is conceptually the right move (and I have thought this for a long time and even said so publically). Chomsky notes that E-merge is not subject to greed considerations and so I-merge shouldn’t be either, given that they are two instances of the very same operation.  Again, yup.  So much for probing and agreeing being a pre-condition for I-merge.

Does this mean that movement is never “for a reason”? Well it does mean that it is never for a local CS reason. Rather, we return once again to a “generate and filter” theory of computation, similar to what we had in GB. The main difference is that this time the filters are provided by Bare Output Conditions, in particular the requirement that all SLOs be labeled prior to Transfer and that MLA sometimes requires what is effectively a Spec-X configuration to provide a viable label. So, Probe/goal seems out and Spec-X is back. Plus ca change (and I know I’m missing a thingy under ‘c’).

This is enough. At least I’ve had enough. Let me once again end on a positive note. As you might have noticed, I am not (yet) convinced by Chomsky’s story here. I believe that he has mis-analyzed the role of labels in G.  His approach rests on the assumption that labels play no role in CS. They are only relevant to interface interpretation.  I currently believe that this is likely wrong. At the very least, I don’t really see how labels are required for CI (or SM) interpretive rules to apply. For CI, at least, all we need, so far as I can make out, is the branching structure and the contents of the individual atoms.  So the premise that motivates the MLA as a BOC seems (at least to me) very shaky.

However, though I am not that moved by Chomsky’s proposal, I am moved by his method and overall conception of the enterprise. I wholeheartedly agree that we should take Galileo’s Maxim as a strong boundary condition on theory.  I completely agree that we should respect GGs history and treat its generalizations as the targets for more principled explanation.[5] I also agree that the critics of GB that he mentions have contributed virtually nothing to our understanding of FL. I even believe that Chomsky’s efforts are an excellent illustration of what we should be aiming to do.  I just am not moved by the details of his effort.  But, really, that’s a small disagreement, among friends.






[1] Chomsky calls these ECP effects. I changed this to FSC effects to distinguish the subject/object asymmetry effects in the ECP from the argument/adjunct ones. Historically, the latter have proven more recalcitrant and were the ones that called forth the heavy ECP assumptions. Indeed, for some analyses, a good part of the subject/object asymmetries were relegated to head government effects (Rizzi, Aoun et al, Saito) more on the SM side of things than the CI.
[2] Transfer of the “phase feature” (PF) from C to T seems like a clear violation of the NTC. As stated, T that has no inherent PF receives one and thereby changes its grammatical powers. One can play around with definitions so that this violation of the NTC doesn’t count, but the definitions will not enjoy the conceptual naturalness that the conceptual versions of the SMT rely on.
[3] It seems that Chomsky is not that happy with T and its various properties.  He, in fact, suggests, that T has no properties of its own. It’s just there to receive properties from C. This makes T very similar to Agr in older MP, and recall that Chomsky argued that such grammar internal formatives are to be eschewed. They cause DP headaches, as do any grammar internal formatives.  The special properties of T are critical in Chomsky’s explanations of the EPP and FSC effects. But, this strikes me as a problem for his account. Why doesn’t explaining these special properties of FL in terms of special properties of T not amount to re-description rather than explanation? One of the salutary effects of MP has been to warn us about confusing the two. However, the more T is special, the more accounts of Spec-T effects (EPP and FSCs) are weakened. 
[4] And talking about memory load: what about the phase hood of D?  Chomsky really does not like to talk about D as a phase. He reluctantly mentions that it might be one from time to time, but it is reluctantly.  But if it is not a phase, then the memory demands on phases can be arbitrarily big given the recursive proclivities of DPs.  So, it would seem that on computational grounds, if v* is a phase then D should be one as well. Indeed, given that “clauses” are already bounded at C, why do we need to also computationally bound them at v*?  Remember, what constitutes a phase needs to be coded into FL and this causes problems for DP. So, all in all, we want the fewest phases we can get away with.
[5] I would go further: I think that these lectures constitute a bit of a departure in Chomsky’s minimalist style. The targets of explanation are now very broad generalizations that GG has established, the aim being to explain their properties.  Earlier efforts were far more local, the targets of explanation being Existential Constructions in Icelandic.  I think that this more general project gets the explanatory grain more correct.

Monday, July 7, 2014

Comments on lecture 3, part III

Here’s part 3. First two are here and here. Chomsky’s lectures are here.

1.     The Halting Problem

As usual, let’s assume that the earlier objections don’t derail the project (which clearly they don’t) and let’s keep following Chomsky’s logic.  Here is one more problem that Chomsky addresses. We know why XPs move and why they can stop. The next question is why they must stop.  Rizzi called this the “halting problem” (no, it’s not related to the real halting problem, though it does sound like it might be eh?). The issue is why a WH (or a DP in an agreeing Spec) does not move any further once there. Chomsky attributes this to uninterpretability of the resulting structure at the CI interface. Let’s look at the details.

The relevant structure is illustrated by (4). This illustrates that criterial agreement “freezes” the DP preventing further movement.  Why?
                        (4)  *What does John wonder [what [C+Q [ Bill ate]]]
Chomsky suggests that (4) is not syntactically illicit but is illegible at CI. Why?  Chomsky does not distinguish between the +Q-C in Wh questions and the one in Yes/No questions.  Empirically, this is a necessary assumption given the observation that a verb can take an embedded WH question as complement iff it can also take a Y/N question. This only makes sense if the Qs in both are the same. If this is so, we can ask why (4) cannot be interpreted like (5):
                        (5) What does John wonder if/whether Bill ate
This has the same structure as (4) but for if/whether. Note that (4) cannot be interpreted as a degraded version of (5). What is less clear to me is why not? There is a +Q-C there, just as in (4). So what’s the problem? We know from matrix clauses that Y/N questions do not require an overt if/whether to license the Y/N interpretation. So, the embedded Q should not need an overt WH morpheme to license the interpretation. Like I said, Chomsky asserts that this derivation has CI problems, but I really don’t see why.[1]

Why does Chomsky attribute the problem with (4) to CI interpretation? He has few other options. Though it is true that agreement suffices to disambiguate the application of MLA, movement will do so as well (i.e. movement will not cause problems for the MLA). But I don’t think that Chomsky wants to say that what blocks further movement should be traced to the details of uF valuation.  Though he could offer the following story: The WH can have its features valued in CP (or via Agree before moving there) and when features can be valued they must be. If so, the WH in the embedded position must have its features valued. If we further assume that feature values cannot be over-written, then if the WH moves further it cannot Agree with the matrix WH and so the matrix {XP, YP} cannot be labeled (recall, we need more than mere feature identity, we need feature agreement).  Note that this relies on some substantive details about feature valuation (e.g. it’s not optional, it’s indelible, uFs cannot “stack”). Perhaps these details follow from the minimal theory feature checking. I leave this to those with a better sense of what is minimal here. However, I suspect that Chomsky does not want to tie his theories to the specifics of feature checking algorithms (I don’t think that I would) and if not, he needs a CI interface story.[2] Whatever the upshot of all of this, we should recall that examples like (4) are orthogonal to the MLA. Chomsky’s CI interface account is there to plug a hole in his story. It does not follow from it and it should be fine if something else explained (4).

So, the MLA has a story for successive cyclic movement, though it is unclear that it is much different from the standard story in the literature, IMO. When all the details are considered the relevant accounts all assume that without agreement in “Spec XP positions,” Gs require movement and with agreement in Spec XP positions Gs forbid movement. The virtues of the MLA are not empirical, but theoretical (viz. that it purportedly follows from the minimal assumptions regarding the algorithm that provides labels in endocentric configurations that the interfaces can read). Just how simple these assumptions are I leave for you to decide. IMO, they ain’t that minimal or simple. But this is partly a matter of taste, I think. I return to a broader consideration of Chomsky’s system towards the end of this exegesis.



            5. Fixed Subject Condition Effects (ECP) and the EPP

Chomsky argues that the MLA is able to unify two further well-known effects given reasonable ancillary assumptions. The two are the subject/object asymmetry in the ECP and the EPP illustrations in (6):
                        (6)       a. Which man did you wonder if Bill liked
                                    b. *Which man did you wonder if liked Bill
                                    c. (*there) arrived a beautiful woman
                                    d. arrivato una bella ragazza
                                    e. (*I) ate a pizza
                                    f.  (I) ho mangiato una pizza

As Chomsky notes (under prodding from David P), the EPP fact that interests Chomsky is the one involving (null) expletives (6c,d), not null subjects together with referential null pronouns (6e,f). It’s not actually clear to me why Chomsky makes this distinction as the account he gives seems to cover both cases. I suspect that the issue is empirical, as I will explain below.

Here’s the account of the EPP.  Chomsky assumes that English T0 is weak (careful!). He treats them as having the same labeling powers as lexical roots (viz. they can’t do it).  This is true even after the tense and agreement features lower from C0 onto T0.  So how is a label assigned by the MLA? Via agreement with a subject in Spec TP. The resultant phi features provide the label that weak T cannot provide on its own. So in (6c), without there the past tense T0 head by itself cannot provide a label. If there is inserted, however, agreement between there and the “T’” suffices to license a phi-label. 

This story raises two questions. First, why is agreement necessary at all? Why can’t the expletive alone label the structure, just like v,n,a suffice to label the structures in e.g. {n, Root} configurations?  This would be empirically awkward, but it is not clear what prevents this theoretically, given Chomsky’s assumptions.  Second, it’s odd that the expletive suffices to license agreement given the standard assumption (Chomsky assumes this too) that in expletive constructions it’s the associate that determines the agreement features on T. But if T agrees not with there but with a beautiful woman then how does the MLA explain these EPP effects? It must be that there agrees with T0 in a way that the associate cannot. What way is that? Unless we are told, it looks like we are again assuming that the EPP is basically the “I-need-a-specifier” condition.

What on Chomsky’s account explains the English contrast with Italian? It assumes that in Italian T0 is strong and thus suffices to license a label on T0 without a DP in its Spec. Thus, Italian T0 is like n,v,a in English.  This account eschews null expletives. Rather the Italian TP in (6d) is a simple {Y, XP} structure with Y being the strong T0.[3] 

The explanation Chomsky gives easily extends to explain the unacceptability of (6e). T0 is weak and a lexical subject is needed to license the labeling via the MLA.  The problem arises with the Italian examples. Here’s what I mean. There is good reason to think that in cases like (6f) there is a pronominal like element in the Spec TP position. The reason is that such sentences are understood as having thematic subjects and these null pronouns care bindable.  If there is nothing there at all, this is hard to understand. Ok, so say there is a null pronoun (aka pro) there.  This is ok for Chomsky so long as we assume that this pro can agree with T0 and this agreement is what affords the label. If we assume this, then pro has phi-features.

Now to English: if Italian pro has phi-features then the minimal assumption is that English pro does too. But then why is (6e) unacceptable?  It seems that what English needs is not merely an agreeing element in Spec XP, but an overt agreeing element. But this seems to obviate the need for Chomsky’s more elaborate assumptions concerning the weak status of T0 in English.  At the least, it suggests that Weak/Strong here has more to do with the SM system than with the CI system. The problem is that it is not clear what this has to do with labeling? 

Are phrase labels required for SM interpretation? Maybe, though specific values seem not to be. What I mean is that identifying that something is an XP might be important for phrasal phonology. But distinguishing VPs from TPs from CPs is not obviously relevant. The question then is whether one needs a labeling algorithm to determine whether something is an XP. Can one identify a structure as XP in the absence of identifying a particular head. It would seem that one can do so easily at least in the relevant cases: any {XP,YP} configuration will be an MaxP for the purposes of SM (as will any {X, YP}. So it would seem that for these purposes MLA is not required, at least conceptually.  But if this is so, then the English/Italian contrast becomes not a property of the MLA but a fact about TPs in English requiring overt subjects in contrast to those in Italian. Why? Well because.  Does the MLA provide a better story? Not so far as I can tell. It all comes down to an idiosyncratic difference between Italian and English and embedding this difference in MLA technology does not appear to do much work.

However, this might be unfair. Chomsky argues that the same account can explain Fixed Subject Constraint (FSC) effects (originally discovered by Perlmutter in his thesis, I believe, and tied together with the availability of pro in the relevant G). Chomsky ties them together as well. Recall, that in English T0 is weak. If so, we need something in “Spec TP” to allow the MLA to label the structure.  That serves to block A’-movement (and DP movement as well, one should note[4]). Note the copy left behind will not suffice as it is part of a chain with links outside the “TP” domain. So the subject WH cannot move at pains of not labeling the TP and causing interface interpretation problems.

Chomsky observes that this predicts that in Italian, where T0 is strong, we should not find FSCs.  And this is plausibly correct (but see below).

In sum, Chomsky argues that the MLA serves to unify EPP and ECP (more accurately FSC) effects and that is another good argument in its favor, in addition to the conceptual ones he uses to motivate the elimination of labeling in the CS.[5]

Here are some potential problems with this analysis: First, though unacceptable, sentences like (6b) are hardly uninterpretable.[6] Indeed, these cases are a bit like the student seems sleeping. These have perfectly obvious CI interpretations and so do examples like (6b). It’s not clear why if they cannot be interpreted at the CI interface. 

Second, we know that FSCs appear in non-interrogatives as well. It’s known as the that-t effect. However, it is also well known that the unacceptability of these seems to vary across speakers. This would predict that for such speakers T0 is strong. But this further predicts that they should find sentences like arrived a nice boy/entered a well dressed dog perfectly acceptable. I have my doubts, but it’s worth looking to find out. 

Third, it’s not clear to me why the indicated reasoning doesn’t block movement from subjects altogether. Why is (7) fine?
                        (7) Which man do you believe t saw Mary
How does the MLA label the embedded TP if the Wh moves?  In other words, why is moving the subject bad if there is something overtly in C but fine if there isn’t?  Chomsky’s proposal does not obviously distinguish these cases. In fact, where Chomsky’s story differs from the traditional one is in tracing FSCs to something about “SpecT-T” relations. The standard accounts tie them to “C-Spec T” relations.[7] By taking the “Spec T-T” relation as central, it’s quite unclear how to account for the obviation of FSCs once the C is deleted.[8]

Fourth, as David Pesetsky noted during the lecture, Chomsky really misdescribes the Italian data.  On his story, it should be possible to extract a WH from “Spec T” position in Italian because Italian T0 is strong.  However, this seems to be incorrect.  In cases where the morphology indicates where the subject has moved from, we find that we cannot WH move from Spec T (this is discussed in a great paper by Brandi and Cordin (here)) Now, one might say that in these dialects T0 is weak, and that might be true. But, one would then also expect no analogues of (6d) in these dialects, and this seems to be incorrect (see Brandi and Cordin 115: (13)/(14)). Of course, appearances here may be deceiving. So, let’s chalk this up as another puzzle.

In sum, Chomsky offers some intriguing connections between the MLA and two more well known FL effects, the FSC and the EPP.  IMO, the analyses are at best suggestive.  There are many loose ends. Chomsky is aware of this and seems not to really be bothered for in his opinion the strength of the proposal lies in its conceptual simplicity. The empirical benefits are a bonus (maybe even a big one) but the problems are tolerable for the proposal depends less on their viability than on the fact that in Chomsky’s opinion, the MLA is the conceptually optimal account of labeling given a the minimal basic operation Merge.  In the last part, I return to consider the conceptual lay of the land once again.





[1] Note that a similar problem does not occur with DPs in “Spec TP” positions. In such cases, TP will be in the complement domain of the C phase head. Thus, before it can move out of the CP, Transfer will remove it from the purview of the computational system. Note that this requires that Cs be strong phases and that T must “inherit” its features from C, otherwise we could generate a TP either with a weak C or no C and then Transfer would not serve to prevent a hyper-raising derivation. Note that such derivations appear to exist (i.e. some Gs allow hyper-raising). Consider this another puzzle for the analysis.
[2] Observe that for this kind of story to work, we need to assume that it is moving WH that has uFs, contrary to the assumption that uFs are limited to phase heads and are checked in Probe/goal configurations.
Note too that this is effectively a greed based account: no movement without feature checking. This makes a lot of sense if agreement takes place in {XP,YP} configurations. Once the features of an XP are valued in {XP, YP} no feature checking is possible and so things stay put. Thus greedy movement suffices to block this. The problem, of course, is that it is not clear how conceptually necessary it is for I-merge to be greedy and why I-merge would be greedy but E-merge would not be (for the cognoscenti, I am trying to develop a slippery slope argument for Q-feature checking: if that were so, then both E and I merge would be greedy).
[3] Is it worth noting that Chomsky’s assumptions regarding T0 make it hard to see why it exists at all. In English, it is indistinguishable from AGR heads: it has virtually no properties of its own and the properties it does have are extremely idiosyncratic.  T has enjoyed a weird position in GG for a very long time. It’s odd within the Barrier’s framework and is no less odd within this one. Again, the odder its properties the greater the challenge for DP concerns.
[4] This seems to suggest that either non-finite T is strong (otherwise successive cyclic DP movement will be prohibited) or that non-finite “TP” doesn’t require a label. Either assumption strikes me as strange. Another puzzle? Another possibility is that non-finite TP is not subject to the EPP as Castillo, Drury and Grohmann as well as Epstein and Seeley have proposed a while back.
[5] For the record, the analysis does not address ECP’s argument/adjunct asymmetries.
[6] Indeed, anyone who has taught undergrad syntax 2 will have encountered speakers that find these kinds of sentences to be only marginally unacceptable. Bleive me, they do exist.
[7] Actually this is true of Rizzi’s proposal and the one in Aoun et al. The idea was that C was able to license the trace in subject position in some cases but not in others. Pesetsky and Torrego develops a version of this account: with movement of Nominative WH to Spec C obviating the need for T to C and that being the morphological reflex of T to C in embedded contexts. In both however the relation of interest in FSCs is that between the C and the expression in Spec TP.
[8] The obviation effects go beyond the removal of an overt C. We find them attenuated in cases where an adverb has been fronted:
                        (i) Who do you think that sooner or later will solve the problem
However, why this should be so on any of these accounts is not particularly clear.