Comments

Showing posts with label effective vs fundamental theory. Show all posts
Showing posts with label effective vs fundamental theory. Show all posts

Monday, March 18, 2013

How I See The Minimalist Project


On reading through the comments to various posts I have developed a feeling that my take on current minimalist research is different from the norm. Most likely mine is idiosyncratic. However, for that reason, I want to outline how I see things so that others can compare their views to this one (should they want to, of course). Along the way, I will discuss what I think makes Minimalism different from earlier programs.  Again, my suspicion is that I see things somewhat differently from others and this may provoke a useful discussion, Let’s see.

Here’s how I interpret the minimalist project. It starts with the acknowledged success of GB.  Some may want to stop reading here for they do not think that GB has been successful. Good. If you judge otherwise, then here’s a good place to stop reading.  At any rate, GB was a success.[1] How so? It identified a dozen or so relatively robust non-trivial generalizations about linguistic dependencies. These generalizations reflect the innate structure of FL. Indeed, they are (a good part of what we know about) UG. The Minimalist Project takes it as given that these generalizations are indeed roughly accurate and proceeds to ask an additional question: why these generalizations and not others?  In other words, it constitutes the project of trying to explain why FL has the GB properties it does. Hopefully, it is clear why the judged success of GB is a sine qua non with this as project.

So, what is it to explain why something has the properties it has?  Well, it amounts to asking whether a given description is fundamental.  My physics friends have names for this (of course). They distinguish two kinds of theories: effective theories vs fundamental ones. The distinction is not meant to be invidious. Effective theories are very highly prized, even if not fundamental.  Indeed, they are often seen to be important necessary steps to deeper understanding (and generally seen to be critical steps towards deeper theories, c.f. discussion by Mitra here on thermodynamics and Landau’s theories). Here are some famous ones: Newtonian mechanics, classical thermodynamics, Maxwell’s theory of electromagnetism, the Standard Theory in quantum mechanics.  Pretty fancy stuff, so saying that some theory is effective is hardly an insult. However, despite their grandeur, these theories are not believed to be fundamental.  What are examples of more fundamental accounts?  Well, Newton’s laws, as we all know, are a special case of Einstein’s when the relevant velocities were small when compared to the speed of light.  In particular, Newton’s laws are limit values to Einstein’s. Another useful example is the relation between the Ideal Gas Laws and statistical mechanics. To my mind, Minimalism is to GB as Statistical mechanics is to the Ideal Gas Laws. Just as the Ideal Gas Laws are a pretty good (though far from perfect) description of regularities that hold among pressure, temperature and volume, so too is GB a pretty good (though not perfect) description of movement/binding/case/etc. dependencies we find in Natural Language grammars. And just like statistical mechanics aims to explain these regularities in terms of the aggregate mechanics of very small very hard atom like particles banging into one another, so too the Minimalist Program aims to isolate the fundamental operations (e.g. Merge) and principles (e.g. minimality) in terms of which it is possible to derive the GB identified grammatical regularities. So, one minimalist project, in my mind the empirically most tractable one given standard linguistic techniques, is to show how the heterogeneous primitives, principles and operations of GB can be unified to a small fundamental core of basic operations and principles.[2]

In my humble (well, not so humble) opinion, this project has had some notable high points.  Let me describe some (and I do mean some). GB has various modules embodying distinctive relations, locality conditions and licensing conditions, viz. Case theory is different from binding theory which is different from control theory which is different from movement theory which is different from phrase structure theory. In my view, one minimalist achievement has been to outline ways of unifying these dependencies by showing how they are all underlyingly reflections of a small common core of operations, viz. Merge. For example:

·      Chomsky 1993 analyzed case as a special instance of movement (i.e. I-merge, case lives on A-chains).
·      Zwart, Idsardi and Lidz, Baltin, Reuland, Drummond, Hicks, and me have proposed that local anaphora and obligatory control live on A-chains and so anaphora and control are also special cases of an A-chain dependency (hence the product of I-merge).[3]
·      Kayne, and Drummond, Kush and Hornstein have proposed treating pronominal binding as a kind of A’-dependency (i.e. a product of I-merge).[4]
·      Kayne has analyzed Weak and Strong Crossover effects in terms of movement (aka, I-merge).
·      Boskovic and Richards have argued that Superiority effects reduce to a species of minimality (i.e. a condition on I-merge).
·      Pesetsky and Torego have analyzed fixed subject effects in terms of derivational economy.
·      Chomsky 2004 has proposed that movement (aka I-merge) and phrase structure (aka E-merge) are actually products of a singular Merge operation.

The achievement is theoretical in that this work unifies apparently disparate grammatical domains.[5]  However, there have been novel empirical by-products as well, e.g. the scope effects of case discussed by Lasnik and Saito (building on earlier observations by Postal), backwards control and raising phenomena as discussed by Polinsky and Potsdam, finite control phenomena discussed by Ferreira, Nunes, Rodrigues, Landau, Seely. The net effect of all of this has been to offer a coherent picture in which the core grammatical relations of GB come together as effects of Merge plus whatever general conditions regulate its application (e.g. extension, minimality, inclusiveness, phases, last resort, LCA, and economy). In other words, theta marking, case checking, local and non-local antecedence, raising, control, crossover, question formation, relativization, X’-theory, (etc.) are all special cases of merge.[6]  If on the right track, this provides fundamental Merge based explanations for effective GB. And, if even roughly correct, this is a hell of a story, a story with interesting biolinguistic implications for the evolution of language and its neurological realization.[7] What I personally find remarkable is that the last 15 years of syntactic research has made a pretty decent case that this kind of picture is not at all implausible. Please note: showing that a very ambitious proposal is not implausible is a very respectable scientific goal. Of course, it would be even better if we could demonstrate (i.e. provide more and more evidence) that this unification is correct. All in good time!

Let me make a few ancillary points.

First, this kind of project of unification, has antecedents in previous generative research. The unification of islands in terms of subjacency/barriers is a project of the very same sort.  The unification of island phenomena has, I believe, been more or less vindicated, most recently in the careful work of Jon Sprouse.  What Jon has shown is not only that islands are almost surely grammatical structure effects (and not performance complexity effects) but that the islands display a pretty unified acceptability profile when violated.  Thus we find what a unified account of islands would lead you to expect. Nice. What I want to emphasize here is that it also provides a nice paradigm within generative research for the kind of minimalist project limned above.

Second, there remain large parts of earlier GB that have proven to be minimalistically refractory. For example, we don’t really have a good alternative account of the argument adjunct asymmetries that the Lasnik-Saito theory was designed to track. There are some sketches, but they fall short, in my view, of being more than suggestive.  There are also a large number of questions that GB ignored, e.g. we have no theory of grammatical features or categories.  So, even were the unification proposed above entirely successful, there would remain many unasked and unanswered questions to investigate.

Third, the project of unification also prompts another question, the significance of which has, in my view, been misinterpreted.  There is a distinctively minimalist question: which features of FL are distinctively linguistic? This is the “beyond explanatory adequacy” (BEA) question. Note that it only gains traction in a post GB context, one in which the aforementioned project of unification has proven successful. Why? Because the fine structure of GB is obviously linguistically tuned.  The relations that the modules traffic in (government, c-command, specified subject condition, antecedent government) have no obvious (or non-obvious) analogues in other domains of cognition.  However, if the relevant modules can all be unified as (e.g. a species of) Merge, then the question of what remains of FL that is specifically linguistic gains traction. Are any of the basic operations (e.g. Merge) or conditions on its application (e.g. extension) specifically linguistic or are they simply the reflections in the domain of language of more general cognitive-biological-computational principles?

There are evolutionary reasons (weak reasons actually, but reasons nonetheless) to hope that there is not much that is cognitively distinctive about UG for the more parochial FL’s principles and primitives the more work we make for evolutionary accounts that must bridge the divide between the grammatically endowed, i.e. us, and the grammatically deprived, i.e. every other animal.  As it is a mitzvah to lighten Darwin’s load to the greatest degree possible, all minimalists (mensches all) hope and pray that FL is more cognitively cosmopolitan than GB envisages. But (you knew this was coming, right?), but this is a worthwhile project just in case a sufficient number of the peculiar properties we find in NL are actually explicable in these more general terms. Here’s what’s not interesting: declaring that Merge is the only distinctive linguistic operation and pronouncing that anything not reducible to Merge is thereby not a linguistic property, but the product of some as yet unidentified interface condition.  This is a game we can be assured of winning and, hence, not very interesting. The declaration that anything that Merge can deal with isn’t   properly part of FL but is rather an interface effect may, in fact, be right. However, relocating something to the interface is not in and of itself doing much of interest. To be interesting requires showing how the specific properties of the interface explain the “linguistic” properties of interest. Minus this last step, what we have is a kind of actuarial minimalism, which, in my view, is an obstacle to serious research. Minimalism is not an exercise in balancing the grammatical books, or at least, no version of minimalism that mainly worries about this is one that I am interested in.

Getting back to the "misinterpreation" alluded to a paragraph back, many have concluded that if minimalism is right then there is no distinctive FL module. However, this is to confuse distinctive with parochial.  As Chomsky has emphasized, and I agree, it’s obvious that there is something special about our linguistic capacities.  Nothing does language like we do.  The problem is to understand the fine structure of this capacity. One hunch has been that this capacity lives on very specialized linguistic “circuits.” Minimalism is questioning this view, and to the degree that unification works, we might be able to whittle down (possibly to zero) what is specifically linguistic in this sense.  However, even if all the basic operations and principles are cognitively-biologically-computationally generic, it is clear that they have been put together in humans (and it appears in humans alone) in a very special way.  One job of generative grammar is to discover the fine structure underlying linguistic competence, i.e. the structure of UG.  It is currently an open question how generic UG is. The minimalist bet is that it is, well, minimal. It is no mean feat to have shown that the project is both viable and plausible, but we are still far from being able to declare the bet ours. Furthermore, even if FL has no dedicated operations or basic principles there is little reason to question that it is cognitively unique and figuring out how its put together from recycled parts is still a very hard unanswered problem.

Let me end: From where I sit, Minimalism has been a raging success.  It has deepened our understanding of UG and has added new questions (or made old ones more prominent) to the research agenda.  However, like all progressive shifts in the larger Generative research program, seen correctly (i.e. from my perspective), minimalism has been far more conservative than it is often advertised to be. And this is a good thing! It’s the hallmark of a progressive research program that earlier results lay the foundations for novel deeper inquiries. Minimalism has built on and conserved and deepened our prior theories, just as one should expect from a thriving program. Respect for one’s elders is a mark of good science. So minimalists, when throwing out bathwater, look out for the babies.



[1] Again, let me make clear, that ‘GB’ here includes its kissing cousins LFG, HPSG, RG, etc.  To my mind there is very little that empirically or conceptually distinguishes these various “frameworks.” Indeed thinking of these as different “frameworks” is analogous to thinking that a ling paper written in French is done in a different “framework” from one written in English or Japanese.
[2] What I mean by “standard linguistic techniques” are the ones that syntacticians are wont to deploy in their daily work.  There are other minimalist questions (see below) that will require developing a different armamentarium, currently more common in biology, psychology, CS and who knows where else. 
[3] Or its functional equivalent, viz. the A-type Agree dependency.
[4] This and the treatment of local antecedence as an A-chain dependency has the effect of deriving the complimentary distribution of reflexivization and pronominalization.
[5] I hope it goes without saying that there are many empirical “puzzles” yet to resolve, as is to be expected from a project of unification.
[6] Jan Koster in the comments has made fun of the “app” like nature of merge. I am not sure why. However, what’s interesting is not whether merge itself is trivial (frankly from one point of view, one hopes that it is) but whether this trivial “addition” can in the right context serve to unify the disparate operations and principles of GB. Showing that this is possible is definitely not trivial. Showing that it is actual would be an amazing achievement.
[7] The interested can see some slight discussion of this by me here.

Friday, January 18, 2013

Effects, Phenomena and Unification


In the previous post, I mentioned that there is a general consensus that UG has roughly the features described in GB. In the comments, Alex, quotes Cederic Boeckx as follows and asks if Cederic is “a climate change denier.”

I think that minimalist guidelines suggest an architecture of grammar that is more plausible biologically speaking that a fully specified, highly specific UG – especially considering the very little time nature had to evolve this remarkable ability that defines our species. If syntax is at the heart of what had to evolve de novo, syntactic parameters would have to have been part of this very late evolutionary addition. Although I confess that our intuitions pertaining to what could have evolved very rapidly are not as robust as one would like, I think that Darwin’s Problem (the logical problem of language evolution) becomes very hard to approach if a GB-style architecture is assumed.

The answer is no, he is not (but thanks for asking). I’ll explain why but this will involve rehearsing material I’ve touched upon elsewhere so if you feel you already know the answer please feel free to go off and do something more worthwhile.

My friends in physics (remember, I am a card carrying hyper-envier) make a distinction between effective and fundamental theories.  Effective theories are those that are phenomenologically pretty accurate. They are also the explananda for fundamental theories.  Using this terminology, GB is an effective theory, and minimalism aspires to develop a fundamental theory to explain GB “phenomena.” Now, ‘phenomena’ is a technical term and I am using it in the sense articulated in Bogen and Woodward (here). Phenomena are well-grounded significant generalizations that form the real data for theoretical explanation. Phenomena are often also referred to as ‘effects.’ Examples in physics include the Gas Laws, the Bernoulli effect, black body radiation, Doppler effects, the photoelectric effect etc.  In linguistics these include island effects, principle A, B and C effects, weak and strong crossover effects, the PRO theorem, Superiority effects etc. GB theory can be seen as a fairly elaborate compendium of these. Thus, the various modules within GB elaborate a series of well-massaged generalizations that are largely accurate phenomenological descriptions of UG. I have at times termed these ‘Laws of Grammar,’ (said plangently you can sound serious, grown-up and self-important) to suggest that those with minimalist aspirations should take these as targets of explanation.  Thus, in the requisite sense, GB (and its cousins described in the last post) can serve as an effective theory, one whose generalizations a minimalist account, a fundamental theory, should aim to explain. 

I hope it is clear how this all relates to the Cedric quote above, but if not here’s the relevance.  Cedric rightly observes that if one is interested in evolutionary accounts then GB cannot be the fundamental theory of linguistic competence.  It’s jus appears as too complex, all that internal modularity (case and theta and control and movement and phrase structure), all those different kinds of locality conditions (binding domains and subjacency/phase and minimality and phrasal domains of a head and government) all those different primitives (case assigners, case receivers, theta markers, arguments, anaphors, bound pronouns, r-expressions, antecedents etc., etc., etc.).  Add to this that this thing popped out in such a short time and there really seems no hope for a semi-reasonable (even just-so) story.  So, GB cannot be fundamental.  BTW, I am pretty sure that I have interpreted Cedric correctly here for we have discussed this a lot over the last five to ten years on a pretty regular basis.

Given the distinction of GB as effective theory and MP as aiming to develop a fundamental theory, how should a thoroughly modern minimalist proceed? Well, as I mentioned before (here) one model is Chomsky’s unification of Ross’s islands via subjacency.  What Chomsky did was (i) treat Ross’s descriptions as effective and (ii) propose how to derive these on more empirically, theoretically and computationally more natural grounds. Go back and carefully read ‘On Wh-Movement’ and you’ll see that how these various strands combine in his (to my taste buds) rather beautiful account. Taking this as a model, a minimalist theory should aspire to the same kind of unification. However, this time it will be a lot harder. For two main reasons.

First, what MP aspires to unify have been thought to be fundamentally different from “the earliest days of generative grammar” (two points and a bonus questions to anyone who identifies the source of this quote). Unifying movement, binding and control goes against the distinction between movement and construal that has been a fundamental part of every generative approach to grammar since Aspects (and before, actually), as has been the distinction between phrase structure and movement. However, much minimalist work over the last 20 years can be seen as chipping away at the differences. Chomsky’s 1993 unification of case as a species of movement or Probe-Goal licensing (PGL), the assimilation of control to a species of movement (moi) or PGL (Landau), reflexive licensing as a species of movement (Idsardi and Lidz, moi) or PGL (Reuland), the collapsing of phrase structure and movement as species of E/I merge, the reduction of Superiority effects to movement via minimality. All of these are steps in reducing the internal modularity of GB and erasing the distinctions between the various kinds of relationships described so well in GB. This unification, if it can be pulled off (and showing that it might be has been, IMO, the distinctive contributions of MP), would do for GB what Chomsky did for islands and the resultant theory would have a decent claim to being fundamental.

The second hurdle will be articulating some notion of computational complexity that makes sense. In ‘On Wh-Movement,’ Chomsky tried to suggest some computational advantages of certain kinds of locality considerations.  Whatever, his success, the problem of finding reasonable third factor features with implications for linguistic coding is far more daunting, as I’ve discussed in other posts. The right notion, I have suggested elsewhere, will reflect the actual design features of the systems that FL interact with and use it. Sadly, we know relatively little about interface properties (especially CI) and we know relatively little about how FL would fit in with other cognitive modules. We know a bit more about the systems that use FL and there have been some non-trivial results concerning what kinds of considerations matter. As I have discussed this in other posts, I will not burden you with a rehash (see here and here). Consequently, whatever is proposed is very speculative, though speculation is to be encouraged for the problem is interesting and theoretically significant.  This said, it will be very hard and we should appreciate that.

So, is Cedric a denier? Nope. He accepts the “laws of grammar” as articulated in GB as more or less phenomenologically correct. Is his strategy rational? Yup. The aim should be to unify these diverse laws in terms of more fundamental constructs and principles. Are people who quote Cedric to “épater les Norberts” doing the same thing? Not if they are UG deniers and not if their work does not aim to explain the phenomena/effects that GB describes. These individuals are akin to climate change deniers for their work has all the virtues of any research that abstracts away from the central facts of the matter.