Comments

Showing posts with label Minimalism. Show all posts
Showing posts with label Minimalism. Show all posts

Friday, March 24, 2017

Lexical items and "mere" morphology-2

This continues the saga from the previous post.

The LK theory had a good run. However, two particularly thorny problems strongly urged revision.  Here we briefly outline the main problems with LK approaches. We then discuss the properties of the GB Binding Theory (BT) that came to replace it with particular emphasis on how it addressed these problems.

There are two problems with LK approaches. First, how to analyze sentences like (5a). On analogy with examples like John kissed himself (derived form underlying John kissed John), they should have underlying structures like (5b). However, (5a) does not mean what (5b) does, and this serves to gut the basic LK approach to reflexivization.[1] Analogous examples in (6) argue against the LK analysis of pronominalization.

(5)       a. Everyone kissed himself                                                                                                               
            b. Everyone kissed everyone
(6)       a. Everyone thinks that he is tall
            b. Everyone thinks that everyone is tall

It is worth considering how these examples pose a problem for the LK analysis. First, given LK’s background assumptions, it appears that the anaphoric morphemes are semantically significant.  Specifically, it is plausible that (5a)/(6a) differ in meaning from (5b)/(6b) precisely because the semantic contributions of himself and he are different from that of everyone.  Adverting to the semantic contributions of the specific anaphors requires reneging on the assumption that binding dependencies are morpheme blind, and suggests that a morpheme centered analysis of binding is the right way to proceed. In other words, rather than being grammatical afterthoughts, the licensing requirements of anaphoric morphemes drive the grammar. Reflexives and bound pronouns are elements that require grammatical validation and grammatical processes cater to their licensing requirements.

Second, once binding theory becomes morpheme centered, economy is no longer required (or desired). Reflexivization and Pronominalization must be ordered with (1) before (2) because Pronominalization applies to the same inputs (aka Structural Descriptions (SD)) as Reflexivization. Were they unordered, the grammar would generate illicit bound pronoun constructions.  However, ordering (1) before (2) is nugatory if they have different SDs. Requiring reference to specific reflexive and a pronoun morphemes in the statement of the operations (as the data in (5) and (6) suggests is necessary) automatically results in rules with different SDs and this then obviates strictly ordering reflexivization before pronominalization. In other words, once binding theory aims to license reflexive morphemes and bound pronoun morphemes, the principles that do so can apply without reference to the applicability of other rules, i.e. the binding principles can be stated unconditionally rather than conditions whose applicability is relative to that of others.

In addition to the problems posed by (5) and (6), there is a second problem with the LK account, forcefully noted in Lasnik 1976. Say it is correct that rules like (1) and (2) prevent bound pronouns from appearing where reflexives do.  This still leaves open the possibility that referential pronouns might appear with the requisite interpretation. Recall, that LK does not regulate (co-)reference, merely binding. Thus, what is wrong with (7a) where the pronoun is not the product of any grammatical process.  This kind of base generated deictic pronoun is what we find in sentences like (7b). So what prevents generating such a pronoun in (7a) with the same reference as John? As Lasnik 1976 notes, (7a) with the co-referential interpretation of the pronoun is rather unacceptable and this is left unexplained so long as the co-reference is not the result of binding but is “accidental.” We can patch up the LK account by adding another disjointness rule, something to the effect that pronouns cannot have local binders (along the lines of Principle B) but, and this is the critical point, this patch seems to render the Pronominalization rule superfluous as the anti-binding Principle B alone suffices to account for the distribution of both bound and unbound pronouns. In effect, they can appear anywhere they are not prohibited.

(7)       a. John loves him
            b. Mary saw him

Lasnik’s (1976) proposal has an interesting theoretical effect: it reinterprets what the grammar tracks. For LK, the grammar licenses antecedence and establishes antecedent-anaphor dependencies. However, the disjointness rule does not establish antecedence but anti-antecedence. This results in a somewhat schizophrenic grammatical theory of bound anaphora.[2] Reflexive are treated more or less as in LK with the grammar requiring that a reflexive have an appropriate local antecedent. In effect, all reflexives are anaphoric and require antecedents and the grammar functions to ensure this result.  Pronouns, in contrast, do not require antecedents and what the grammar does is ensure that if there is an antecedent it is not too local.  Importantly, this collapses bound and non-bound pronouns into one class. Indeed, it’s this move that allows Lasnik (1976) to address the puzzle he identified. For LK, grammars code for anaphoric dependency. After Lasnik (1976) grammars regulate the distribution of classes of morphemes regardless of interpretation, with some later extra grammatical process (we might now say interface processes) determining which pronouns are understood as bound and which not.
So with this as background, let’s consider the general features of the GB Binding Theory (BT).  It consists of three principles:

(8)       A. An anaphor (e.g. a reflexive) must be bound in its domain.
            B. A Pronoun cannot be bound (must be free) in its domain.
            C. An R-expression cannot be bound.

Domains have been variously characterized, but for the nonce we can assume that it is the minimal (finite) clause containing the anaphor/pronoun/R-expression. There are several things to observe about these principles. [3]

First, they are morpheme centered. The inputs to the rules are expressions that fall into one of three classes, anaphors (+a,-p), pronominals (+p,-a) and R-expressions (-a,-p).[4] Furthermore, the rules state licensing requirements for these expressions. In other words, differing binding requirements restrict the distribution (and interpretation) of lexical items with differing feature constitutions. Note that this effectively rejects the LK distinction between lexical and grammatical formatives.  Reflexives and bound pronouns are just as lexical as cat and run, though their features call forth grammatical licensing in ways that more “ordinary” lexical items do not.

Second, the principles are unconditional in the sense that neither A nor B is ordered or needs to be ordered with respect to the other.  What accounts for the complementarity of reflexives and pronouns is the fact that both elements meet inverse licensing requirements in the same local domain. Where anaphors must be bound, pronouns must not be. There is no sense within the standard GB version of BT that pronouns and anaphors or their licensing requirements are in an economy relation or that one kind of dependency is preferred to the other.  Both apply where they can and the presence or absence of either has no grammatical effect on the other. 

A thought experiment makes the contrast with LK evident. I imagine that an epidemic struck wiping out the rule of reflexivization. In an LK grammar, pronominalization would apply to yield sentences like John likes him with a bound interpretation.  Given the same epidemic, this time wiping out A, such sentences would still be prohibited for pronominals would still be subject to B.

Third, the principles apply regardless of how the morphemes are semantically interpreted. This is clearest in the case of B. The grammar does not single out any particular antecedence relation. B states that pronouns cannot appear in certain configurations, it does not specify or distinguish and DP as antecedent.  In fact, the co-indexing that is part of the system has no univocal semantic interpretation. This is what allows B to accommodate Lasnik’s (1976) worries. The indexing in (8) violates B regardless of whether it is interpreted as binding or co-reference.

(8) *John1 likes him1 

This has the effect of enriching the interpretive systems, as some indexations are interpreted as binding dependencies and some are not. Which are which falls beyond the purview of the grammar despite the fact that they have consequences for acceptability that do not seem particularly “semantic.” Consider one example. Weak Crossover Effects (WCO) are restricted to bound pronouns, e.g. ones where the antecedent is a quantified DP. Thus there is a contrast in (9a,b) where him can be interpreted as John but not as bound by everyone.

(9)       a. his1 photos distressed John1
            b. *his1 photos distressed everyone1

The standard description of this is that WCO constrains binding relations but not (co-) reference. The indexing in (9a) is interpreted as co-reference while the one in (9b) cannot be so interpreted as everyone is not a referring expression.  However, as binding is illicit in WCO configurations, it cannot receive a binding interpretation so the indexing yields no good interpretation and the structure yields unacceptability.[5] If this is correct, then BT requires supplementation to distinguish those indexings that will be interpreted as bindings from those that will not. Why? Because the indexing above is purely syntactic and it tracks a mixed bag of interpretive possibilities.  This is just another way of saying that BT per se accounts for the distribution of certain morphemes, not the dependencies that they enter into.

Another consequence of the GB re-invention of the binding theory is Principle C. As noted above, LK had no corresponding principle. Rather, for LK Principle C effects are by products of how Reflexivization and Pronominalization are stated.  Once the idea that reflexives and pronouns are real inputs to semantic interpretation and not mere morphological dress-up, the LK strategy for dealing with Principle C effects is not longer viable and an explicit statement is required. 

As has long been noted, Principle C is somewhat odd. First, it is not bounded in any way, unlike A and B, which apply within circumscribed domains.  Second, at least initially, it was taken to apply to what on the surface appear to be entirely different kinds of cases. The examples in (10) do not violate Principles A or B.  Consequently, if these are to be excluded an additional binding principle is required. Furthermore, whereas (10c,d) plausibly pertain to the structural conditions imposed on antecedents and their dependent anaphors, (10a,b) do not appear to involve anaphoric elements at all.  To collect all these cases under the same principle it is critical to add another category of expressions, R-expressions, to the inventory of elements regulated by the Binding Theory.[6]  Like Principle B, Principle C does not specify licit anaphoric dependencies but blocks illict ones.  It is negative, rather than a positive rule of grammar.[7] 

(10)     a. *John1 likes John1
            b. *John thinks that John is tall
            c. *John1 expect himself1 to like John1
            d. *He1 thinks that John1 is tall

Ok, let’s take a step back now and consider these two approaches. As should be evident, the theoretical intuitions behind BT are very different than those behind LK.  They differ not only technologically (LK uses transformations, is derivational in spirit and has morpheme rewrite rules while BT has indexing algorithms and is representational in spirit and uses filters stated at various grammatical levels) but in what they take the subject matter of binding to be.  They differ along three critical dimensions:

1.     Do binding principles have a natural semantic interpretation?
2.     Are binding principles economy principles (or absolute)?
3.     Are binding principles morpheme (or dependency) centered?
LK answers yes to (1) and (2) and no to (3). GB answers no to (1) and (2) and yes to (3). What of minimalist approaches? I don’t know, but my hunch is that LK approaches have some features that are worth reconsidering in in MP context. How so?

Well, first, the main problem with LK accounts noted above disappear in the context of theories that distinguish copies due to I-merge and those that result from multiple selections of the same expression from the lexicon. The LK theory did not (and could not) distinguish these two possibilities and so the fact that everyone loves himself does not mean the same as everyone loves everyone, was sufficient to sink the LK approach. However, as you know, this does not hold for MP accounts so long as binding tracks movement. If it does, then we can distinguish the two case above syntactically.

As noted, this requires endorsing a movement theory of binding. The main obstacle to this in earlier GG theories was D(eep) Structure. Given DS, there could be no movement between “theta” positions and so the technology MP provides to distinguish everyone1 loves everyone2 from everyone1 loves everyone1 could not be applied. However, if we eschew DS (as MP stories do) then it is in principle possible to move between “theta” positions and so distinguish these two kinds of chains. We can then restrict LK reasoning to the second kind of chain without empirical hazard. So, the elimination of DS restrictions is a pre-requisite for revivifying the LK approach, and this is precisely what MP theories allow.

Last, we must allow Gs that differentially spell out copies. The LK theory, recall, treats the morphological differences between reflexives and bound pronouns as syntactically very superficial. What counts are the underlying chains/dependencies, not the morphemes that express them. Imo, this is likely a good thing. Let me explain why.

The aim of MP is to explain why we have the UG principles we do. In other words, the aim is to explain the structure of FL/UG. Morpheme centered Gs don’t cannot support this kind of project. Why not? Because they effectively stipulate G requirements by packing them into the idiosyncratic feature make-ups of specific lexical items. Why must anaphors be locally bound? Because they have features that require that they be locally bound. Why do they have these features? Well, because they are reflexives and reflexives inherently have such features by stipulation. This explanation is circular in the worst sense: the circle is very very tight.

Note, that such stipulations are particularly problematic for those with minimalist fish to fry. They are not, for example, worrisome for Plato Problem kinds of issues. So, if A-anaphors are innately part of any lexicon then their binding requirements need not be learned.[8] However, if one wants to go beyond explanatory adequacy, then such stipulations stink. They defeat the MP project from the get-go, as they stipulate what we want to have explained.[9] In other words, from an MP perspective the problem with classical binding theory is that it is based on a series of morphological stipulations concerning specific lexemes, and stipulations like these prevent explanation. So, if your goal is to explain why the binding theory looks the way it does, then you don’t want morpheme centered accounts of binding like the ones we find in GB. Of course, this goal might be unattainable and morpheme based accounts might be the best we can do, but…

Ok, basta! This has gone on far too long. Let me just suggest that earlier GG theories had some properties that are worth re-examining, most particularly the idea that some morphemes are just by-products of grammatical processes rather than being the causal engines behind them. This is not a new idea, but MP has given them, imo, a new lease on life and part of this lease implies rethinking the idea that all formatives are created equal and have an equal purchase in interpretation at the interfaces.



[1] Recall in LK accounts the pre-transformational phrase marker was sole input to semantic interpretation. However, even later approaches which allowed information from several grammatical levels to contribute to semantic interpretation would not have been able to incorporate an LK analyses which sensibly captured the meaning of examples like (5a) and (6a).
[2] This is curiously mirrored in Lasnik’s paper where the appendix deals with bound pronouns and the body of the paper with co-reference.
[3] There is a third reason for moving from an LK approach to a morpheme centered one. In the mid 1970s there was a well motivated theoretical move to dramatically simplify structural descriptions (SDs) and Structural Changes (SCs). This made rules that included the insertion of specific morphemes less natural/desirable. The main problem was that such morpho-phonological intruders complicated the move to simple rules like Move alpha anywhere. Flash forward 40 years and the analogue of LK rules finds a natural home: it results from processes the spell out copies/occurrences. See below.
[4] PRO was taken to be (+a,+p).  We ignore PRO here.
[5] Strong Crossover effects yield a similar problem. Variables are the semantically quintessential anaphors.  They are no referring expressions and require binders.  As such, one might think that they would qualify as anaphoric (+a,-p) elements. In fact the earliest versions of trace theory categorized wh-traces/variables as anaphors. However, so categorizing variables leads to a big empirical problem. Sentences like (i) are wrongly expected to be interpretable as (ii).
(i)            Who1 does he1 think t1 is intelligent
(ii)          Who1 t1 thinks he1 is intelligent
This conclusion can be finessed by cataloging residues of movement, t1, as an R-expression subject to principle C.  This works, but it also highlights the fact that ‘R’ does not mean “referential” in the naïve sense of the term.
[6] One of the authors suspects that the attractiveness of the GB theory of PRO was in part the result of filling an available cell in the required feature matrix. If anaphors are +a,-p and pronouns are –a,+p, and R-expression are –a,-p then there should be something that is +a,+p. PRO was the proposed missing link. 
[7] There are some problems with this way of dealing with the data in (10). First, it is not clear that (10a,b) are really as unacceptable as (10c,d). Indeed there are languages where analogous sentences seem perfectly well formed and express anaphoric dependencies (c.f. Boeckx, Hornstein and Nunes 200x for some discussion). Second, it is not particularly clear what an R-expression is. Among the elements that fell under the category are traces (interpreted as bound variables), names, definite descriptions, demonstratives etc.  What all these expressions have in common besides is unclear. Variables are prototypical anaphors. Names are prototypical non-anaphors.  Definite descriptions can be used anaphorically or not and semantically and grammatically they share many properties with pronouns. Nonetheless, they too are categorized as R-expressions.  The category seems to be a catch-all with entry requirements to the fraternity driven entirely by empirical necessity. This results in negligible explanatory force.
[8] These kinds of theories, however, do require some non-trivial explications of how morpho-phonoligical agreement implicates semantic dependency. Just because two expressions have the same features need imply nothing about whether/how these features have semantic significance.
[9] Incidentally, this is why PRO based accounts of control should also be MP suspect. Assuming PRO with its special licensing requirements allows to track control facts but not explain why control exists.

Wednesday, June 15, 2016

Case & agreement: beware of prevailing wisdom

Someone recently told me a (possibly apocryphal) story about the inimitable Mark Baker. The story involves Mark giving a plenary lecture somewhere on the topic of case. To open the lecture, possibly-apocryphal-Mark says something along the following lines:
Those of you who don't work on case probably have in your heads some rough sketch of how case works. (e.g. Agree in person/number/gender between a designated head and a noun phrase, resulting in that noun phrase being case-marked.) What you need to realize is that basically nobody who actually works on case believes that this is how case works.
Now, whether or not this is really how it all went down, possibly-apocryphal-Mark has a point. In fact, I'm here to tell you that his point holds not only of case, but of agreement, too.

In one sense, this situation is probably not all that unique to case & agreement. I'm sure presuppositions and focus alternatives don't actually work the way that I (whose education on these matters stopped at the introductory stage) think they work, either. The thing is, no less than the entire feature calculus of minimalist syntax is built on this purported model of case & agreement. [If you don't believe me, go read "The Minimalist Program" again; you'll find that things like the interpretable-uninterpretable distinction are founded on the (supposed) behavior of person/number/gender and case (277ff.).] And it is a model of case & agreement that – to repeat – simply doesn't work.

So what model am I talking about? I'm really talking about a pair of intertwined theories of case and of agreement, which work roughly as follows:
  1. there is a Case Filter, and it is implemented through feature-checking: each noun phrase is born with a case feature that, were it to reach the interfaces (PF/LF) unchecked, would cause ungrammaticality (a.k.a., a "crash"); this feature is checked when the noun phrase enters into an agreement relation with an appropriate functional head (T0, v0, etc.), and only if this agreement relation involves the full set of nominal phi features (person, number, gender)
  2. agreement is also based on feature-checking: the aforementioned functional heads (T0, v0, etc.) carry "uninterpretable person/number/gender features"; if these reach the interfaces (PF/LF) unchecked, the result is – you guessed it – ungrammaticality (a.k.a., a "crash"); these uninterpretable features get checked when they are overwritten with the valued person/number/gender features found on the noun phrase
Thus, on this view, case & agreement live in something of a happy symbiosis: agreement between a functional head and a noun phrase serves to check what would otherwise be ungrammaticality-causing features on both elements.

From the vantage point of 2016, however, I think it is quite safe to say that none of this is right. And, in fact, even the Abstractness Gambit (the idea that (1) and (2) are operative in the syntax, but morphology obscures their effects) cannot save this theory.

What follows builds heavily on some of my own work (though far from exclusively so; some of the giants whose shoulders I am standing on include Marantz, Rezac, Bobaljik, and definitely-not-apocryphal Mark Baker) – and so I apologize in advance if some of this comes across as self-promoting.

––––––––––––––––––––

Let's start with (1). Absolutive(=ABS) is a structural case, but there are ABS noun phrases that could not possibly have been agreed with, living happily in grammatical Basque sentences. How do we know they could not possibly have been agreed with (not even "abstractly")? Because we know that (non-clitic-doubled) dative arguments in Basque block agreement with a lower ABS noun phrase, and we can look specifically at ABS arguments that have a dative coargument. (Indeed, when the dative coargument is removed or clitic-doubled, morphologically overt agreement with the ABS – impossible in the presence of the dative coargument – becomes possible.)

So if an ABS noun phrase in Basque has a dative coargument, we know that this ABS noun phrase could not have been targeted for agreement by a head like v0 or T0 (because they are higher than the dative coargument). Notice that this rules out agreement with these heads regardless of whether that supposed agreement is overt or not; it is a matter of structural height, coupled with minimality. The distribution of overt agreement here serves only to confirm what our structural analysis already leads us to expect.

And yet despite the fact that it could not have been targeted for agreement, there is our ABS noun phrase, living its life, Case Filter be damned. [For the curious, note that this is crucially different from seemingly similar Icelandic facts, which Bobaljik (2008) suggests might be handled in terms of restructuring. That is because whether the embedded predicate is ditransitive (=has a dative argument) or monotransitive (=lacks one) cannot, to the best of my knowledge, affect the restructuring possibilities of the embedding predicate one bit.]

If you would like to read more about this, see my 2011 paper in NLLT, in particular pp. 929 onward. (That paper builds on the analysis of the relevant Basque constructions that was in my 2009 LI paper, so if you have questions about the analysis itself, that's the place to look.)

––––––––––––––––––––

Moving to (2), this is demonstrably false, as well. This can be shown using data from the K'ichean languages (a branch of Mayan). These languages have a construction in which the verb agrees either with the subject or with the object, depending on which of the two bears marked features. So, for example, Subj:3sg+Obj:3pl will yield the same agreement marking (3pl) as Subj:3pl+Obj:3sg will. It is relatively straightforward to show that this is not an instance of Multiple Agree (i.e., the verb does not "agree with both arguments"), but rather an instance of the agreeing head looking only for marked features, and skipping constituents that don't bear the features it is looking for. Just like an interrogative C0 will skip a non-[wh] subject to target a [wh] object, so will the verb in this construction skip a [sg] (i.e., non-[pl]) subject to target a [pl] object.

This teaches us that 3sg noun phrases are not viable targets for the relevant head in K'ichean. Ah, but now you might ask: "What if both the subject and the object are 3sg?" The facts are that such a configuration is (unsurprisingly) fine, and an agreement form which is glossed as "3sg" shows up in this case (so to speak; it is actually phonologically null). That's all well and good; but what happened to the unchecked uninterpretable person/number/gender features on the head? Remember, they couldn't have been checked, because everything is now 3sg. And if 3sg things were viable targets for this head, then you could get "3sg" agreement in a Subj:3sg+Obj:3pl configuration, too – by simply targeting the subject – but in actuality, you can't. [This line of reasoning is resistant even to the "but what about null expletives?" gambit: if the uninterpretable phi features on the head were checked by a null expletive, then either the expletive is formally plural or formally singular. If it is singular, then we already know it could not have been a viable target for this head; if it is plural, and it has been targeted for agreement, then we predict plural agreement morphology, contrary to fact. Thus, alternatives based on a null expletive do not work here.]

What about Last Resort? It is entirely possible that grammar has an operation that swoops in should any "uninterpretable features" have made it to the interface unchecked, and deletes the offending features. But now ask yourself this: what prevents this operation from swooping in and deleting the features on the head even when there was a viable agreement target there for the taking (e.g. a 3pl nominal)? i.e., why can't you just gratuitously fail to agree with an available target, and just have the Last Resort operation take care of your unchecked features later? The only possible answer is that the grammar "knows that this would be cheating"; the grammar makes sure the Last Resort is just that – a last resort – it keeps track of whether you could have agreed with a nominal, and only if you couldn't have are you then eligible for the deletion of offending features. Put another way, the compulsion to agree with an available target is not reducible to just the state of the relevant features once they reach the interfaces; it is obligatory independently of such considerations. You see where this is going: if this bookkeeping / independent obligatoriness is going on anyway, uninterpretable features become 100% redundant. They bear exactly none of the empirical burden (i.e., there is no single derivation in the entire grammar that would be ruled out by unchecked features, only by illicit application of the Last Resort operation).

Bottom line: there is no grammatical device of any efficacy that corresponds to this notion of "uninterpretable person/number/gender feature."

––––––––––––––––––––

At this juncture, you might wonder what, exactly, I'm proposing in lieu of (1-2). The really, really short version is this: agreement and case are transformations, in the sense that they are obligatory when their structural description is met, and irrelevant otherwise. (Retro, ain't it?) To see what I mean, and how this solves the problems associated with (1) and (2), I'm afraid you'll have to read some of my published work. In particular, chapters 5, 8, and 9 of my 2014 book. Again, sorry for the self-promotional nature of this.

––––––––––––––––––––

Epilogue:

Every practicing linguist has, in their head, a "toy theory" of various phenomena that are not that linguist's primary focus. This is natural and probably necessary, because no one can be an expert in everything. The difference, when it comes to case and especially when it comes to agreement, is that these phenomena have been (implicitly or explicitly, rightly or wrongly) taken as the exemplar of feature interaction in grammar. And so other members of the field have (implicitly or explicitly) taken this toy theory of case & agreement as a model of how their own feature systems should work.

And lest you think I have constructed a straw-man, let me end with an example. If you follow my own work, you know that I have been involved in a debate or two recently where my position has amounted to "such and such phenomenon X is not reducible to the same mechanism that underlies agreement in person/number/gender." What strikes me about these debates is the following: if A is the mechanism that underlies agreement, these (attempted) reductions are not reductions-to-A at all; they are reductions-to-the-LING-101-version-of-A (e.g. Chomsky's Agree), which – to paraphrase possibly-apocryphal-Mark – nobody who works on agreement thinks (or, at least, nobody who works on agreement should think) is a viable theory of agreement.

Now, it is logically possible that a feature calculus that was invented to capture agreement in person/number/gender (e.g. Agree), and turns out to be ill-suited for that purpose, is nevertheless – by sheer coincidence – the right theory for some other phenomenon (or set of phenomena) X. But even if that turns out to be the case, because the mechanism in question doesn't account for agreement in the first place, there is no "reduction" here at all.


Monday, June 6, 2016

Theory, again

It’s the start of the summer so it’s time to return to some pet peeves. Here’s the fortune cookie version of the history of Generative Grammar (GG): we have moved from the study of Gs to the study of possible Gs to the study of possible FL/UGs. The (bulk of the) earliest work in GG (e.g. Syntactic Structures, LSLT, the Standard Theory) aimed to adumbrate the kinds of rules that Gs contain by studying the actual recursive mechanisms that specific Gs embody. The next stage aimed to adumbrate not only the rules that Gs actually contain but also the principles restricting the kinds of operations a G could contain (this is what UG in GB was all about). Minimalism builds on the results of all of this earlier research and aims to limn the contours of a possible human Faculty of Language (FL). It, in effect, addresses the question: why do we have the FL/UG we in fact have rather than some conceivable others?

As is obvious (but this won’t stop me form pressing the point) these research questions are closely inter-related with connections in two directions.  First, each later question starts from answers provided by the earlier one. It’s pointless to wonder about possible rules without some candidate actual ones and it is futile to investigate the limits of FL/UG without some candidate principles of FL/UG. Second, answers to later questions limit the range of answers to earlier ones. If a rule is not FL/UG possible then a particular G cannot contain such a rule and if a principle is not a possible principle of FL/UG then no FL/UG can contain that kind of principle.

So, two observations: first, the three kinds of questions above are importantly different even if closely related (as such, they must be kept logically and conceptually distinct). Second, the dialectic from answer to answer moves in both directions from “lower” level to “higher” and back again. “Lower” and “higher” are not intended as evaluative. They are just used to mark the conceptual flow noted above.

Here’s a third observation: despite their interconnections, the methods used to study each of these questions are partially autonomous from each other. People who study particular Gs can do useful work without resort to the accepted/proposed principles of FL/UG and those interested in the universal properties of Gs (i.e. the structure of FL/UG) can get a good way into this problem without bothering too much with minimalist concerns. The methods used to investigate all three questions partially overlap, but the criteria for success are not the same and even some of the detailed kinds of arguments advanced can have somewhat different flavors. So not only are the questions different, but progress on addressing them is somewhat independent of progress in addressing the others. Just as there is no discovery procedure for Gs (no reduction of later levels to earlier ones), there is none for theories of GGG (no requirement that later questions uncritically respect the answers provided to earlier ones). The questions are related to one another in roughly the way that levels in a G are: they take in one another’s washing in complicated ways.

Why do I mention this? Because I believe that some of the unease in current syntax stems from misunderstanding what question is being addressed by a particular proposal and thus what counts as evidence for or against it. Or to put this another way: if the above is a roughly correct characterization of the conceptual GG landscape, then it is important to understand that many proposals, especially “higher” ones, are hidden conditionals. For example, minimalist proposals are of the form: Given that such and such is a plausible (better still, actual) principle of FL/UG then so and so is why this kind of principle obtains rather than others.

If this is so, then there are two ways to reject a specific proposal: (i) argue against the conditional as a whole or (ii) argue only against the antecedent. The former denies that the deductive link between premise and conclusion holds. The latter denies the relevance of the deductive link even if it does hold. As I see it, most critiques of minimalist proposals are of the second kind. They deny that what is taken as given should be so taken because the premise is empirically suspect. In other words, many objections are actually objections to the underlying “GB” principle being “explained” (and hence assumed) in minimalist terms rather than the explanation itself.[1]  These critiques deny the utility of the explanation rather than question its deductive validity. Thus they conclude that showing how to deduce the principle from more general considerations is valueless because the premise is false. IMO, this conclusion is unfortunate and it reflects a general disdain for theory characteristic of much work in contemporary “theoretical” syntax. Let me vent a bit (again).

In the real sciences, a lot of time is spent trying to find ways of tying together seemingly disparate principles. It really isn’t easy to show that two principles that look different are nonetheless fundamentally the same. And the problem is in large part conceptual. And one way that conceptual problems are investigated is by (often radically) simplifying them. Of course, the hope is that the simplification will preserve many of the core features of interest and so the simplification can “scale up” as we make the premises more realistic. Such simplifications often rest on “stylized” facts that are acknowledged to be (ahem) “incomplete” (aka: false). However, investigating such empirically inadequate simple problems based on stylized facts is often a vital step in advancing understanding even though the premises might be false (as simplifications almost always are). The same should hold true in syntax.

Btw, this sort of investigation (largely pencil and paper kind of stuff) is what is commonly called ‘theoretical.’ Theoretical work consists in investigating how simple concepts can be related to produce theories with rich deductive structure. Theory places a premium on (i) the reasonableness (rather than the truth) of the basic simplification (i.e. the rough accuracy of the stylized facts), (ii) the naturalness of the assumed basic concepts and (iii) the depth of the deductive structure that results.

A good example of this in GG is Chomsky’s recent proposals concerning Merge. It runs roughly as follows: if you assume that Merge is a very simple binary operation that takes two syntactic objects (SOs) and combines them into a set of those SOs (i.e. If A is an So and B is an SO then {A,B} is an SO) then you can generate objects with unbounded hierarchical structure with the following “nice” properties: Merge must be structure dependent (linear order irrelevant to syntax) and cyclic (e.g. no lowering rules), phrase structure building and movement are two faces of the self same basic Merge operation (E and I-merge), movement (aka I-Merge) must target c-commanding positions (due to Extension), and the products of I-Merge necessarily produce copies (due to Inclusiveness and hence producing structures supporting operator-variable structures and allowing for reconstruction effects). So, from a simple idea concerning the recursive mechanism, Chomsky derives a bunch of plausible properties of Gs and UG that GGers have proposed over the last 50 years of research.

However, the generalizations deduced (cyclicity, c-command, copies etc.) are not perfect (e.g. tucking-in is not strictly speaking cyclic in the standard usage, there are many cases in which reconstruction is impossible, movement is not the only operation for which c-command is relevant). Does that mean that the Chomsky’s unification of these properties in terms of Merge is a bad one? Not necessarily. Conceptually it is an achievement for it shows how to link certain salient (stylized) features of Gs together. Empirically, it is a step forward for it links properties that have non-negligible empirical backing and that are plausibly descriptive of our FL. Is it “true”? Well, that depends on how we eventually handle the (apparent) problems for the (lower level) principles that it has unified. Should these prove to be false, then this unification will not be what we ultimately want. However, and this is important, Chomsky’s unification provides a strong (explanatory) incentive for going back and reanalyzing the (empirical) “problems” for the lower level principles, and it provides a nice example of the kind of theory we want. We really do want to have our cake and eat it too and this is what the dialectic between empirical “coverage” and theoretical “explanation” aims to provide. The problem is that for this dialectic to gain a foothold we need to appreciate both sides of the going-and-froing. We need to concretely understand the tension between explanatory force and empirical coverage and understand that the right theory needs both. Right now, IMO, our attitudes over-prize (apparent) empirical coverage. We very seldom count (or even address) the cost of lost explanation when we evaluate our proposals.

This is not a new complaint, at least from me. I make it again because in my experience GGers have a low tolerance for theoretical ambition. I suspect that this is so for several reasons. First, we tend to confuse formal work with theoretical work and this muddies our sensitivity to the explanatory oomph of different approaches. Second, linguistics is a data rich field and so supporting theory means tolerating some empirical slack at least for a while. But, last, I think that we don’t actually spend enough time teaching and touting the explanatory virtues of our best accounts. We seldom go back and ask what we have lost or try to theoretically motivate the new principles we adopt to “capture” the data. Indeed, the whole idea that data is something that needs capturing (rather than explaining) is, to my mind, quite odd.

Does this mean that theory does not need empirical support? Nope. Theories need to be justified by facts. But, facts also need to be justified by theories. One of the original hopes of the minimalist program was that it would sensitize us to what a good explanation was. It would make us aware that our “explanations” (and these are scare quotes) are often as complex as the data they address. And this is not good. IMO, this appreciation is less vivid today than it was in the earliest days of the minimalist program. And part of the problem is un-interest in theory and a misplaced belief that lots of data signifies empirical progress. In this regard, GG work has been disimproving.



[1] “GB” is in quotes because I do not mean to invidiously distinguish between GB proper and its many theoretical twins (many of them identical IMO for most of the questions I am interested in). These include LFG, RG, GPSG, HPSG a.o. From where I sit, most of these theories are intertranslatable and make effectively the same distinctions in the same theoretical places. They are more notationally than notionally distinct.