Comments

Showing posts with label languistics. Show all posts
Showing posts with label languistics. Show all posts

Friday, March 14, 2014

Hornstein's lament

Chomsky has long noted the tension between description and explanation. One methodological aim of the early Minimalist Program (MP) was to focus attention on this tension so as to nudge theories in a more explanatory direction.[1] The most important hygienic observation of early MP was that many proffered analyses within GB lacked explanatory power for they were as complex as the data they intended to “explain.” Early MP aimed to sharpen our appreciation of the difference between (re)description and explanation and to encourage us to question the utility of proposals that serve mainly to extend the descriptive range of our technical apparatus.

How well has this early MP goal been realized? IMO, not particularly well. I have had an opportunity of late to read more widely in the literature than I generally do and from this brief foray into the outside world it looks to me that the overriding imperative of the bulk of current syntax research is descriptive. Indeed, the explanatory urge, weak as it has been in the past, appears to me now largely non-existent. Here’s what I mean.

The papers I’ve read have roughly the following kinds of structure:

1.     Phenomenon X has been analyzed in two different ways; W1 and W2. In this paper I provide evidence that both W1 and W2 are required.
2.     Principle P forbids structures like S. This paper argues that phenomenon X shows that P must be weakened to P’ and/or that an additional principle P’’ is required to handle the concomitant over-generation.
3.     Principle P prohibits S. Language L exhibits S. To reconcile P with S in L we augment the features in L with F, which allows L to evade P wrt S.

In each instance, the imperative to cover the relevant data points has been paramount. The explanatory costs of doing so are largely unacknowledged. Let me put this a different way. All linguists agree that ceteris paribus simpler theories are preferable to more complex ones (i.e. Ockham shaves us all!). The question is: What makes things not equal? Here is where the trade-off between explanation and description plays out.

MP considerations urge us to live with a few recalcitrant data points rather than weaken our explanatory standards.[2]  The descriptive impulse urges us to weaken our theory to “capture” the strays. There is no right or wrong move in these circumstances. Which way to jump is a matter of judgment. The calculus requires balancing our explanatory and descriptive urges. My reading of the literature is that right now, the latter almost completely dominate. In other words, much (most?) current research is always ready to sacrifice explanation in service of data coverage. Why?

I can think of several reasons for this.

First, I think that the Chomsky program for linguistics has never been widely endorsed by the linguistic community. To be more precise, whereas Chomskian technology has generally garnered an enthusiastic following, the larger cognitive/bio-linguistic program that this technology was in service of, has been warily embraced, if at all. How many times have you seen someone question a technical innovation because it would serve to make a solution to Plato’s or Darwin’s problem more difficult? How many papers have you read lately that aim to reduce rather than expand the number of operative principles in FL/UG? More pointedly, how often do we insist that students be able to construct Poverty of Stimulus (POS) arguments? Indeed, if my experience is anything to go by, the ability to construct a POS is not considered a necessary part of a linguist’s education. Not surprisingly, asking whether a given phenomenon or proposal raises “learnability” issues at all is generally not part of evaluative calculus.  The main concerns revolve around how to deploy/reconfigure the given technology to “capture the facts.”[3]

Second, most linguists take their object of study to be language not the faculty of language. Sophisticates take the goal of linguistics to be the discovery of grammatical patterns. This contrasts with the view that the goal of linguistics is to uncover the basic architecture of FL. I have previously dubbed the first group languists and the second linguists. There is no question that languists and linguists can happily co-exist and that the work of each can be beneficial to the other.[4] However, it is important to appreciate that the two groups (the first being very much bigger than the second) are engaged in different enterprises. The grandfather of languistics is Greenberg, not Chomsky. The goal of languistics is essentially descriptive, the prize going to the most thorough descriptions of the most languages. Its slogan might be no language left behind.

Linguists are far more opportunistic. The description of different languages is not a goal in itself. It is valuable precisely to the degree that it sheds light on novel mechanisms and organizing principles of FL. The languistic impulse is that working on a new or understudied language is ipso facto valuable. Not so the linguistic one. Linguists note that biology has come a long way studying essentially three organisms: mice, e-coli and fruit flies. There is nothing wrong in studying other organisms, but there is nothing particularly right about it either. Why? Because the aim is to uncover the underlying biological mechanisms, and if these do not radically differ across organisms then myopically focusing on three is a fine way to proceed. Why assume that the study of FL will be any different? Don’t get me wrong here: studying other organisms/languages might be useful. What is rejected is the presupposition that this is inherently worthwhile. 

Together, the above two impulses (actually two faces of the same one) reflect a deeper belief that the Chomsky program is more philosophy than science. This reflects a more general view that theoretical speculation is largely flocculent, whereas description is hard and grounded. You must know what I am about to say next: this reflects the age-old difference between empiricist and rationalist approaches to naturalistic explanation. Empiricism has always suspected theoretical speculation, conceiving of theory in largely instrumental terms. For a rationalist what’s real are the underlying mechanisms, the data being complex manifestations of the interacting more primitive operations. For the empiricist, the data are real, the theoretical constructs being bleached averages of the data, with their main virtue being concision. Chomsky’s program really is Rationalist (both in its psychology and its scientific methodology). The languistic impulse is empiricist. Given that Empiricism is the default “scientific” position (Lila Gleitman once remarked that it is likely innate), it is not surprising that explanation almost always gives way to description, for the former has not independent existence from the latter.

So what to do? I don’t know. What I suspect, however, is that the rise of languistics is not good for the future of the field. To the degree that it dominates the discipline it cuts Linguistics off from the more vibrant parts of the wider intellectual scene, especially cognitive and neuro-science. Philology is not a 21st century growth industry. Linguistics won’t be either if it fails to make room for the explanatory impulse.



[1] MP had both a methodological and a substantive set of ambitions. A good part of the original 93 paper was concerned with the former in the context of GB style theories.
[2] Living with them does not entail ignoring or forgetting about them. Most theories, even very good ones, have a stable of problem cases. Collecting and cataloguing them is worthwhile even if theory is insulated from them.
[3] This is evident in the review process as well. One of the virtues of NELs and WCCFL papers is that they are short and their main ideas easy to discern. By the time a paper is expanded for publication in a “real” journal it has bloated so much that often the main point is obscured. Why the bloat? In part this is a response to the adept reviewer’s empirical observations and scrambling by the author to cover the unruly data point. This is often done by adding some ad hoc excrescence that it takes pages and pages to defend and elaborate. Wouldn’t the world be a better place if the author simply noted the problem and went on leaving the general theory/proposal untouched? I think it would. But this would require judging the interest of the general proposal, i.e. the theory would have to be evaluated independently of whether it covered every last data point.
[4] IMO, one of the reasons for GG’s success was that it allowed the two to live harmoniously without getting into disputes. MP sharpens the differences between languists and linguists and this may in part explain its more luke warm embrace.

Wednesday, July 24, 2013

LSA Summer Camp


The LSA summer institute just finished last week. Here are some impressions.

In many ways it was a wonderful experience and it brought back to me my life as a graduate student.  My apartment was “functional” (i.e. spare and tending towards the slovenly). As in my first grad student apartments, I had a mattress on the floor and an AC unit that I slept under. The main difference this time around was that that the AC unit I had at U Mich was considerably smaller than the earlier industrial strength machine that was able to turn my various abodes into a meat locker (I’m Canadian/Quebecois and ¯“mon pays ce n’est pas un pays c’est l’hiver…”¯ !). In fact, this time around the AC was more like ten flies flapping vigorously. It was ok if I slept directly under the fan (hence the floor mattress).  The downside, something that I do not remember from my experience 40 years ago, was that this time around, getting up out of bed was more demanding now than it was then.

I was at the LSA to teach intro to minimalist syntax.  It was a fun course to teach. There were between 80-90 people that attended regularly, about half taking the course for some kind of credit. To my delight, there was real enthusiasm for minimalist topics and the discussion in class was always lively.  The master narrative for the course was that the Minimalist Program (MP) aims to answer a “newish” question: what features of FL are peculiarly linguistic? The first lecture and a half consisted of a Whig history of Generative Grammar, which tried to locate the MP project historically. The main idea was that if one’s interest lies in distinguishing the cognitively general from the linguistically parochial within FL there have to be candidate theories of FL to investigate. GB (for the first time) provides an articulated version of such a theory, with the sub-modules, (i.e. Binding theory, control theory, movement, subjacency, the ECP, X’ theory etc.) providing candidate “laws of grammar.” The goal of MP is to repackage these “laws” in such a way as to factor out those features that are peculiar to FL from those that are part of general cognition/computation.  I then suggested that this project could be advanced by unifying the various principles in the different modules in terms of Merge, in effect eliminating the modular structure of FL. In this frame of mind, I showed how various proposals within MP could be seen as doing just this: Phrase Structure and Movement as instances of Merge (E and I respectively), case theory as an instance of I-merge, control, and anaphoric binding as instances of I-merge (A-chain variety) etc.  It was fun. The last lectures were by far the most speculative (it involved seeing if we could model pronominal binding as an instance of A-to-A’-to-A movement (don’t ask)) but there was a lot of interesting ongoing discussion as we examined various approaches for possible unification.  We went over a lot of the standard technology and I think we had a pretty good time going over the material. 

I also went on a personal crusade against AGREE.  I did this partly to be provocative (after all most current approaches to non-local dependencies rely on AGREE in a probe-goal configuration to mediate I-merge) and partly because I believe that AGREE introduces a lot of redundancy into the theory, not a good thing, so it allowed us to have a lively discussion of some of the more recondite evaluative considerations that MP elevates.[1]  At any rate, here the discussion was particularly lively (thanks Vicki) and fun. I would love to say that the class was a big hit, but this is an evaluation better left to the attendees than to me. Suffice it to say, I had a good time and the attrition rate seemed to be pretty low.

One of the perks of teaching at the institute is that one can sit in on one’s colleagues’ classes. I attended the class given by Sam Epstein, Hisa Kitihara and Dan Seely (EKS).  It was attended by about 60 people (like I said, minimalism did well at this LSA summer camp).  The material they covered required more background than the intro course I taught and EKS walked us through some of their recent research. It was very interesting. The aim was to develop of an account of why transfer applies when it does. The key idea was that cyclic transfer is forced in computations that result in in multi-peaked structures that themselves result from strict adherence to derivations that respect (an analogue of) Merge-Over-Move and feature lowering of the kind that Chomsky has recently proposed.  The technical details are non-trivial so those interested should hunt down some of their recent papers.[2]

A second important benefit of EKS’s course was the careful way that they went through some of Chomsky’s more demanding technical suggestions, sympathetically yet critically.  We had a great time discussing various conceptions of Merge and how/if labeling should be incorporated into core syntax. As many of you know, Chomsky has lately made noises that labeling should be dispensed with on simplicity grounds. Hisa (with kibbitzing from Sam and Dan) walked us though some of his arguments (especially those outlined in “Problems of Projection”). I was not convinced, but I was enlightened. 

Happily, in the third week, Chomsky himself came and discussed these issues in EKS’s class.  The idea he proposed therein was that phrases require labels at least when transferred to the CI interface. Indeed, Chomsky proposed a labeling algorithm that incorporated Spec-Head agreement as a core component (yes, it’s back folks!!).  It resolves labeling ambiguities.  To be slightly less opaque: in {X, YP} configurations the label is the most prominent (least embedded) lexical item (LI) (viz. X). In {XP, YP} configurations there are two least embedded LIs (viz. descriptively, the head of X and the head of Y). In these cases, agreement enters to resolve the ambiguity by identifying the two heads (i.e. thereby making them the same). Where agreement is possible, labeling is as well. Where it is not, one of the phrases must move to allow labeling to occur in transfer to CI.  Chomsky suggested that this requirement for unambiguous labeling (viz. the demand that labels be deterministically computed) underlies successive cyclic movement. 

To be honest, I am not sure that I yet fully understand the details enough to evaluate it (to be more honest, I think I get enough of it to be very skeptical). However, I can say that the class was a lot of fun and very thought provoking. As an added bonus, it brought me and Vicki Carstens together on a common squibbish project (currently under construction). For me it felt like being back in one of Chomsky’s Thursday lectures. It was great.

Chomsky gave two other less technical talks that were also very well attended. All in all, a great two days.

There were other highlights. I got to talk to Rick Lewis a lot. We “discussed” matters of great moment over excellent local beer and some very good single malt scotch. It was as part of one of these outings that I got him to allow me to post his two papers here. One particularly enlightening discussion involved the interpretation of the competence/performance distinction. He proposed that it be interpreted as analogous to the distinction between capacities and exercisings of capacities.  A performance is the exercise of a capacity. Capacities are never exhausted by their exercisings.  As he noted, on this version of the distinction one can have competence theories of grammars, of parsers, and of producers. On this view, it’s not that grammars are part of the theory of competence and parsers part of the theory of performance. Rather, the distinction marks the important point that the aim of cognitive theory is to understand capacities, not particular exercisings thereof. I’m not sure if this is exactly what Chomsky had in mind when he introduced the distinction, but I do think that it marks an important distinction that should be highlighted (one further discussed here).

Let me end with one last impression, maybe an inaccurate one, but one that I nonetheless left with.  Despite the evident interest in minimalist/biolinguistic themes at the institute, it struck me that this conception of linguistics is very much in the minority within the discipline at large. There really is a linguistics/languistics divide that is quite deep, with a very large part of the field focused on the proper description of language data in all of its vast complexity as the central object of study. Though, there is no a priori reason why this endeavor should clash with the biolinguistic one, in practice it does. 

The two pursuits are animated by very different aesthetics, and increasingly by different analytical techniques.  They endorse different conceptions of the role of idealization, and different attitudes towards variation and complexity. For biolinguists, the aim is to eliminate the variation, in effect to see through it and isolate the individual interacting sub-systems that combine to produce the surface complexity. The trick on this view is to find a way of ignoring a lot of the complex surface data and hone in on the simple underlying mechanisms. This contrasts with a second conception, one that embraces the complexity and thinks that it needs to be understood as a whole. On this second view, abstracting from the complex variety manifested in the surface forms is to abstract away from the key features of language.  On this second view, language IS variation, whereas from the biolinguistic perspective a good deal of variation is noise.

This, of course, is a vast over-simplification. But I sense that it reflects two different approaches to the study of language, approaches that won’t (and can’t) fit comfortably together. If so, linguistics will (has) split into two disciplines, one closer to philology (albeit with fancy new statistical techniques to bolster the descriptive enterprise) and one closer to Chomsky’s original biolinguistic conception whose central object of inquiry is FL.

Last point: One thing I also discovered is how much work running one of these Insitutes can be. The organizers at U Michigan did an outstanding job. I would like to thank Andries Coetze, Robin Queen, Jennifer Nguyen and all their student helpers for all their efforts.  I can be very cranky (and I was on some days) and when I was, instead of hitting me upside the head, they calmly and graciously settled me down, solved my “very pressing” problem and sent me on my merry way. Thanks for your efforts, forbearance and constant good cheer.



[1] I make this argument in chapter 6 here.
[2] See the three papers in 2010, 2011, and 2012 by EKS noted here