Comments

Showing posts with label Plato's Problem. Show all posts
Showing posts with label Plato's Problem. Show all posts

Friday, May 25, 2018

Three quintessentially minimalist projects

I was in Barcelona last week giving some lectures on the triumphant march forward of the Minimalist Program (MP). As readers may know, I believe that MP has been a success in its own terms in that it has gone a fair way towards answering the questions it first posed for itself, viz. Why do we have the FL we actually have and not some other? Others are more skeptical, but I believe that this is mainly because critics demand that MP address questions not central to its research mission. Of course, answering the MP question might leave many others untouched, but that hardly seems like a reason to disparage MP so much as a reason for pursuing other programs simultaneously. At any rate, this was what the lectures were about and I thank the gracious audience at the Autonomous University of Barcelona for letting me outline these views for their delectation. 

In getting these lectures into shape I started thinking about a question prompted by recent comments from Peter Svenonius (thx, see here). Peter thinks, if I understand him correctly, that the MP obsession (ok, my obsession) with Darwin’s Problem (DP) really adds nothing to the MP enterprise. Things could proceed more or less as we see them without this biological segue. This got me thinking about the following question: Which MP projects are largely motivated by DP concerns? I can think of three. They may well be motivated on other grounds as well. But they seem to me a direct consequence of taking the DP perspective on the emergence of FL seriously. This post is a first stab at enumerating these and explaining why I think they are intimately tied to DP (in much the same way that the P&P project was intimately tied to Plato’s Problem (PP)). So what are the three lines of inquiry? A warning: The order of discussion does not imply anything about their relative salience or importance to MP. I should note, that many of the points I make below, I have made elsewhere and before. So do not expect to be enlightened. This is probably more for my benefit than for yours.

First, Unification of the Modules (UoM). MP is based on the success of the P&P program, in particular the perceived success of GBish conceptions of FL/UG. Another way of saying this is that if you don’t think that GB managed to limn a fairly decent picture of the fine structure (i.e. the universals) of FL/UG then the MP project will seem to you to be, at best, premature and, at worst, hubristic. 

I believe that a lot of the problems that linguists have with MP has less to do with its failure to make progress in answering the MP question above, than with the belief that the whole project presupposes accepting as roughly right a deeply flawed conception of FL/UG (viz. roughly the GB conception). So, for example, if you don’t like classical case theory (and many syntacticians today do not) then you won’t like a project that takes it to be more or less accurate and tries to derive its properties from deeper principles. If you don’t think that classical binding theory is more or less correct then you won’t like a project that tries to reduce it to something else. The problem for many is that what MP presupposes (namely that GB was roughly correct, but not fundamental) is precisely what they believe ought to be very much up for grabs. 

I personally have a lot of sympathy for this attitude. However, I also think that it misses a useful feature of the presupposition. MP motivated unification/reduction doesn’t require that the GB description be correct so much that it be plausible enough (viz. that it be of roughly the right order of complexity) as to make deriving its properties a useful exercise (i.e. an exercise that if successful will provide a useful modelfor future projects armed with a more accurate conception of FL/UG). To put this another way: the GB conception of FL/UG has identified design features of FL/UG (aka, universals) that are theoretically plausible and empirically justifiable and so it is worth asking why principles such as theseshould govern the workings of our linguistic capacities. In Poeppel’s immortal phrase, these GB principles are the right grain size for analysis, have non-trivial empirical backing and so the exercise of showing why they should be part of FL would be doing something useful even should they fail to reflect the exact structure of FL.[1]

So, a central project animated by DP is unification of the disparate GB modules. And this is a very non-trivial project. As many of you know, GB attributes a rather high degree of internalmodularity to FL. There are diverse principles regulating binding vs control vs movement vs selection/subcategorization vs theta role assignment vs case assignment vs phrase structure. From the perspective of Plato’s Problem, the diversity of the modules does not much matter as their operations and principles are presumed to be innate (and hence not learned). In fact, the main impetus behind P&P architectures was to isolate the plausibly invariant features of Gs, and explain them by attributing them to the internal workings of FL thereby constraining the Gs FL produces to invariably respect these features. Thus the reason that Gs are always structure dependent is that FL has the property of only being able to construct Gs that are structure dependent. The reason that movement and binding require c-command is that FL imposes c-command as a condition on these diverse modular operations. The aim of GB was to identify and factor out the invariant properties of specific Gs and treat them as fixed features of FL/UG so that they did not have to be acquired on the basis of PLD (and a good thing too as there is not sufficient data in the PLD to fix them (aka PoS considerations apply to these)). The problem of G acquisition could then focus on the variable parts (where Gs differ) using the invariant parts as Archimedean fixed points for leveraging the PLD into specific Gs. That was the picture. And for this P&P project to succeed, it did not much matter how “complex” FL/UG was so long as it was innate. 

All of this changes, and changes dramatically, once one asks how this system couldhave arisen. Then the internal complexity matters, and matters a lot. Indeed, once one asks this question there is a great premium on simple FL architectures, with fewer modules and fewer disparate principles for the simpler the structure of FL, the easier it is to imagine how it mighthave arisen from the cognitive architecture of predecessors that did not have one. 

If this is correct, then one central MP project is to show that the diversity of the GB modules is only apparent and that they are only different reflections of the same underlying operations and principles. In other words, the project of unifying the modules is central to MP and it is central becauseof DP. A solution to DP requiresthat what appearsto be a very complex FL system (i.e. what GB depicts) is actually quite simple and what appearto be very different modules with different operations and regulative principles are really all reflections of the same underlying generative procedures. Why? Because short of this it will be impossible to explain how the system that GB describes could have arisen from a mind without it. 

This is entirely analogous, in its logic, to Plato’s Problem. How can kids acquire the Gs they do with the properties they have despite a poverty of the linguistic stimulus? Because much of what they know they do not have to learn. How could humans have evolved an FL from non-FL cognitive minds? Because FL minds are only a very small very simple step away from the minds that they emerged from and this requires that the modular complexity GB attributes to FL is only apparent. It’s what you get when you add to the contents of non-linguistic minds the small simple addition MP hypothesizes bridged the ling/non-ling gap.

Are there other plausible motives for such a project, the project of unifying the modules? Well perhaps. One might argue that an FL with unified modules are in some methodological sense better than one with non-unified ones. Something like a principle that says fewer modules are better than more. Again, I think that this is probably correct, but let’s face it, this kind of methodological Ockamist accounting is very weak (or at least perceived to be so). When push comes to shove data coverage (almost?) always trumps such niceties (remember the ceteris paribus clausethat always accompanies such dicta). So it is worth having a big empiricalfact of interest driving the agenda as well. And there are few facts bigger and heftier than the fact that FL arose from non-FL capable minds and it is easier to explain how this could havehappened if FL capable minds are only mildly different from non-FL capable minds and this means that the complex modularity that GB attributes to FL capable minds is almost certainly incorrect. That’s the line of argument. It rests on DPish assumptions and, to my mind, provides a powerful empirical motivation for module unification, which is what makes unification a central MP project.

It suggests a second related project: not only must the modules be unified, but the unification should makes use of the fewest possible linguistically proprietary operations and principles. In other words, linguistically capable minds, ones withFLs should be as minimally linguisticallyspecial as possible. Why? Because evolution proceeds most smoothly when there is minimal qualitative difference between the evolved states. If the aim is to explain how language ready minds appeared from non language ready minds than the fewer the differences between the two, the easier it will to be to account for the emergence of the former form the latter. If one assumes that what makes an FL mind language ready are linguistically special operations and principles then the fewer of these the better. In fact, in the best case there will be exactly a single relatively simple difference between the two, language ready minds just being non-language ready ones plus (at most) one linguistically special simple addition (the desideratum that it be simple motivated by the assumption that simple additions are more likely to become evolutionarily available than complex ones).[2]

So let’s assess: there are two closely related MP projects: unify the GB modules and unify them using largely non-linguistically proprietary operations and principles. How far has this project gotten? Well, IMO, quite far. Others are sure to disagree. But the projects though somewhat open textured have proven to be manageable and, the first in particular, has generated useful hypotheses (e.g. the Merge Hypothesis and extensions thereof, like the Movement Theory of Control and Construal), which even if wrong have the right flavor (Iknow, I know, this is self serving!). Indeed, IMO, trying to specify exactly where and how these theories go wrong (if they do, color me skeptical but I have dogs in these fights) and why they go wrong as they do, is a reasonable extension of the basic MP projects. It is a tribute to how little MP concerns drive contemporary syntax that such questions are, IMO, rarely broached. Let me rant a bit.

Darwin’s Problem (DP) currently enjoys as little interest among linguists today as Plato’s Problem (PP) does (and did, in earlier times). Indeed, from where I sit, even PP barely animates linguistic investigations. So, for example, people who study variation rarely ask how it might be fixed (though there are notable exceptions). Similarly, people who propose novel principles and operations rarely ask whether and how they might be integrated/unified with the rest of the features of FL. Indeed, most syntacticians take the basic apparatus as given and rarely critically examine it (e.g. how many people worry about the deep overlap between Agree and I-merge?). These are just not standard research concerns. IMO, sadly, most linguists could care less about the cognitive aspects of GG, let alone its possible bio-linguistic features. The object of study is language, not FL, and the technical apparatus is considered interesting to the degree that it provides a potentially powerful philological tool kit. 

Ok, so MP motivates two projects. There is one more, and it concerns variation. GB took variation to be bounded. It did this by conceiving UG as providing a finiteset of parameter values and conceived of language acquisition as fixing those parameters. So, even if the space of possible Gs is very large, for GB, it is finite. Now, given the linguistic specificityof the parameters, and given that GB treats them as internalto FL, the idea that variation is a matter of parameter setting proves to be a deep MP challenge. Indeed, I would go so far as to say, that ifMP is on the right track, thenFL does not contain a finite list of possible binary parameters and G acquisition cannot be a matter of parameter setting. It must be something else, something that is not specific to G acquisition. And this idea has caught on, big time. Let me explain.

I have many times mentioned the work by Berwick, Lidz and Yang on G acquisition. Each contains what is effectively a learning theory that constructs Gs from PLD using FL principles. It appears that this general idea is quite widely accepted now, with former parameter setting types (e.g. David Lightfoot) now arguing that “UG is open” and that there is “no evaluation of I-languages and no binary parameters” (1).[3]This view is much more congenial to MP as it removes the very specific parametric options fromFL and treats variation as entirely a “learning” problem. G learning is no different than other kinds, it is just aimed at Gs.[4]

Of course to make this work, will require specifying what kids come to the learning problem with, what kinds of data they exploit, and what the details of the G learning theory are. And this is hard. It requires more than pointing to differences in the PLD and attributing differences in Gs to these differences. However, this is a long way from an actual learning theory which specifies how PLD and properties of FL combine to give you a G. Not the least important fact is that there are many ways to generalize from PLD to Gs and kids only exploit some of these.[5]That said, if there is an MP “theory” of variation it will consist of adumbrating the innate assumptions the LAD uses to fix a particular G on the basis of PLD. To date, we have some interesting proposals (in particular from Lidz and Yang and their colleagues in syntax) but no overarching theory.

Interestingly, if this project can be made to fly, then it will also be the front end of an MP theory of variation. To date, the main focus of research has been on unifying and simplifying FL and trying to determine how much of FL is linguistically proprietary. However, there is no reason that the considerable current typological work on G variation shouldn’t feed into developing theories of learning aimed at explaining why we find the variation we do. It is just that thisproject is going to be very hard to execute well, as it will demand that linguists develop skills that are not currently part of standard PhD training, at least not in syntax (e.g. courses in stats, machine learning, and computation). But isn’t this as it should be? 

So, does taking MP seriously make a difference? Yes! It spawns three projects all animated by the MP problematic. These projects make sense in the context of trying to specify the internal structure of an FL that couldhave evolved from earlier minds. It suggests three concrete projects. So the programmatic aspects of MP are quite fecund, which is all that we can ask of a program.

And results? Well, here too I believe that we have made substantial progress as regards the first project, some as regards the third (though it is very very hard) and a little concerning the second.  IMO, this is not bad for 25 years and suggests that the DPish way of framing the MP issues has more than paid for itself.


[1]It’s worth adding that this sort of exercise is quite common in the real sciences. Ideal gases are not actual gases, planets are not point masses, and our universe may not be the only possible one but figuring out how they work has been very useful. 
[2]There is a lot of hand waving going on here. Thus, what evolves are genomes and what we are talking about here are phenotypic expressions thereof. We are assuming that simple genotypic difference reflect simple genetic differences. Who knows if this is right. However, it is the standard assumption for this kind of biological speculation so it would be a form of methodological dualism to treat it as suspect onlyin the linguistic case. See herefor discussion of this “phenotypic gambit” and its role in evolutionary thinking.
[3]See “Discovering New Variable Properties without Parameters,” in Massimo Piattelli-Palmarini and Simin Karimi, eds., “Parameters: What are they? Where are they?” Linguistic Analysis 41, special edition (2017).
            A very terse version of this view is advanced in Hornstein (2009) on entirely MP grounds. The main conceptual difference between approaches like Lightfoot’s and the one I advanced is that the former relies on the idea that “children DISCOVER variable properties of their language through parsing” (1), whereas I waved my hands and mumbled something about curve fitting given an enhanced representation provided by FL (see herefor slightly more elaboration).
[4]This folds together various important issues, the most important being that there is no overall evaluation metric for parameter setting. Chomsky argued that the shift from evaluation metrics to parameter setting modules increased the latters feasibility because applying global evaluation metrics to Gs is computationally intractable. I think Chomsky might have though that parameter setting is more localized than G evaluation and so will not require fancy learning theories. It turns, as Dresher and Kaye long ago noted, that parameter setting models have their own tractability issues unless the parameters can be set independently of one another. If they are not independent, problems quickly arise (e.g. it is hard to fix parameters once and for all). 
Furthermore, it is not clear to me that something like global measures of G fitness can be entirely avoided, though Lightfoot insists that they should be. The main reason for my skepticism is empirical and revolves around the question of whether the space of G options is scattered or not. At least in syntax, it seems that different Gs are kept relatively separate (e.g. bilinguals might code switch between French and English but they don’t syntactically blend them to get an “average” of the two in Frenglish. Why not?). This suggests that Gs enjoy an integrity and this is what keeps them cognitively apart. Bill Idsardi tells me that this might be less true on the sound side of things. But as regards the syntax, this looks more or less correct. If it is, then some global measure distinguishing different Gs might be required. 
I should add that more recently, if I recall correctly, Fodor and Sakas have argued that the evaluation metric cannot be completely dispensed with even on their “parsing” account.
[5]So, for example, invoking “parsing” as the driver behind acquisition does not do much unless one specifies howparsing works. Recall that standard parsers (e.g. the Marcus Parser) embody Gs that guide how it is that input data is analyzed. No G, no parsing. But if the aim is to explain how Gs are acquired then one cannot presuppose that the relevant G already exists as part of the parser. So what does a parse consist in in detail? This is a hard problem and it turns out that there are many factors that the child uses to analyze a string so as to recover a meaning. The MP project is to figure out what this is, not to name it.

Wednesday, August 26, 2015

It's not really possible that UG doesn't exist

There are four big facts that undergird the generative enterprise:

1.     Species specificity: Nothing talks like humans talk, not even sorta kinda.
2.     Linguistic creativity: “a mature native speaker can produce a new sentence on the appropriate occasion, and other speakers can understand it immediately, though it is equally new to them’ (Chomsky, Current Issues: 7). In other words, a native speaker of a given L has command over a discrete (and for all practical and theoretical purposes) infinity of differently interpreted linguistic expressions.
3.     Plato’s Problem: Any human child can acquire any language with native proficiency if placed in the appropriate speech community.
4.     Darwin’s Problem: Human linguistic capacity is a very recent biological innovation (roughly 50-100 kya).

These four facts have two salient properties. First, they are more or less obviously true. That’s why nobody will win a Nobel prize for “discovering” any of them. It is obvious that nothing does language remotely like humans do, and that any kid can learn any language, and that there is for all practical and theoretical purposes an infinity of sentences a native speaker can use, and that the kind of linguistic facility we find in humans is a recentish innovation in biological terms (ok, the last is slightly more tendentious, but still pretty obviously correct). Second, these facts can usefully serve as boundary conditions on any adequate theory of language. Let’s consider them in a bit more detail.

(1) implies that there is something special about humans that allows them to be linguistically proficient in the unique way that they are. We can name the source of that proficiency: humans (and most likely only humans) have a linguistically dedicated faculty of language (FL) and “designed” to meet the computational exigencies peculiar to language. 

(2) implies that native speakers acquire Gs (recursive /procedures or rules) able to generate an unbounded number of distinct linguistic objects that native speakers can use to express their thoughts and to understand the expressions that other native speakers utter. In other words, a key part of human linguistic proficiency consists in having an internalized grammar of a particular language able to generate an unbounded number of different linguistic expressions. Combined with (1), we get to the conclusion that part of what makes humans biologically unique is a capacity to acquire Gs of the kind that we do.

(3) implies that all human Gs have something in common; they are all acquirable by humans. This strongly suggests that there are some properties P that all humans have that allow them acquire Gs in the effortless reflexive way that they do.

Indeed, cursory inspection of the obvious facts allows us to say a bit more: (i) we know that the data available to the child vastly underdetermines the kinds of Gs that we know humans can acquire thus (ii) it must be the case that some of the limits on the acquirable Gs reflect “the general character his [the acquirers NH] learning capacity rather that the particular course of his experience” (Chomsky, Current Issues; 112). (1), (2) and (3) imply that FL consists in part of language specific capacities that enable humans to acquire some kinds of Gs more easily than others (and, perhaps, some not at all).

Here’s another way of saying this. Let’s call the linguo-centric aspect of FL, UG. More specifically, UG is those features of FL that are linguistically specific, in contrast to those features of FL that are part of human or biological cognition more generally. Note that this allows for FL to consist of features that are not parts of UG. All that it implies is that there are some features of FL that are linguistically proprietary. That some such features exist is a nearly apodictic conclusion given facts (1)-(3). Indeed, it is a pretty sure bet (one that I would be happy to give long odds on) that human cognition involves some biologically given computational capacities unique to humans that underlie our linguistic facility. In other words, the UG part of FL is not null.[1]

The fourth fact implies a bit more about the “size” of UG: (4) implies that the UG part of FL is rather restricted.[2] In other words, though there are some cognitively unique features of FL (i.e. UG is not empty), FL consists of many operations that FL shares with other cognitive components and that are likely shared across species. In other words, though UG has content, most of FL consists of operations and conditions not unique to FL.

Now, the argument outlined above is often taken to be very controversial and highly speculative. It isn’t. That humans have an FL with some unique UGish features is a trivial conclusion from very obvious facts. In short, the conclusion is a no-brainer, a virtual truism! What is controversial, and rightly so, is what UG consists in. This is quite definitely not obvious and this is what linguists (and others interested in language and its cognitive and biological underpinnings) are (or should be) trying to figure out. IMO, linguists have a pretty good working (i.e. effective) theory of FL/UG and have promising leads on its fundamental properties. But, and I really want to emphasize this, even if many/most of the details are wrong the basic conclusion that humans have an FL with UGish touches must be right. To repeat, that FL/UG is a human biological endowment is (or should be) uncontroversial, even if what FL/UG consists in isn’t.[3] 

Let me put this another way, with a small nod to 17th and 18th century discussions about skepticism. Thinkers of this era distinguished logical certainty from moral certainty. Something is logically certain iff its negation is logically false (i.e. only logical truths can be logically certain). Given this criteria, not surprisingly, virtually nothing is certain. Nonetheless, we can and do judge that many more or less certain that are neither tautologies nor contradictions. Those things that enjoy a high degree of certainty but are not logically certain are morally certain. In other words, it is worth a sizable bet with long odds given. My claim is the following: that UG exists is morally certain. That there is a species specific dedicated capacity based on some intrinsic linguistically specific computational capacities is as close to a sure thing as we can have. Of course, it might be wrong, but only in the way that our bet that birds are built to fly might be wrong, or fish are built to swim might be. Maybe there is nothing special about birds that allow them to fly (maybe as Chomsky once wrly suggested, eagles are just very good jumpers). Maybe fish swim like I do only more so. Maybe. And maybe you are interested in this beautiful bridge in NYC that I have to sell you. That FL/UG exists is a moral certainty. The interesting question is what’s in it, not if it’s there.

Why do I mention this? Because in my experience, discussions in and about linguistics often tend to run the whether/that and the what/how questions together. This is quite obvious in discussions of the Poverty of Stimulus (PoS). It is pretty easy to establish that/whether a given phenomenon is subject to PoS, i.e. that there is not enough data in the PLD to fix a given mature capacity. But this does not mean that any given solution for that problem is correct. Nonetheless, many regularly conclude that because a proposed solution is imperfect (or worse) that there is no PoS problem at all and that FL/UG is unnecessary. But this is a non-sequitur. Whether something has a PoS profile is independent of whether any of the extant proposed solutions are viable.

Similarly with evolutionary qualms regarding rich UGs: that something like island effects fall under the purview of FL/UG is, IMO, virtually uncontestable. What the relevant mechanisms are and how they got into FL/UG is a related but separable issue. Ok, I want to walk this back a bit: that some proposal runs afoul of Darwin’s Problem (or Plato’s) is a good reason for re-thinking it. But, this is a reason for rethinking the proposed specific mechanism, it is not a reason to reject the claim that FL has internal structure of a partially UGish nature. Confusing questions leads to baby/bathwater problems, so don’t do it!

So what’s the take home message: we can know that something is so without knowing how it is so. We know that FL has a UGish component by considering very simple evident facts. These simple evident facts do not suffice to reveal the fine structure of FL/UG but not knowing what the latter is does not undermine the former conclusion that it exists. Different questions, different data, different arguments. Keeping this in mind will help us avoid taking three steps backwards for every two steps forward.



[1] Note that this is a very weak claim. I return to this below.
[2] Actually, this needs to be very qualified, as done here and here.
[3] This is very like the mathematical distinction between an existence proof vs a constructive proof. We often have proofs that something is the case without knowing what that something is.

Tuesday, May 12, 2015

Darwin's problem; some reasonable

Karthik Durvasula has pointed me to a thoughtful blog post by the Confused Academic (CA) critical of using Darwin’s Problem (DP) kind of considerations as a criterion in the evaluation of linguistic proposals (see here).[1] The main thrust of the (to repeat, very reasonable) points made is that we know very little about how evolutionary considerations apply to cognitive phenomena in general and linguistic phenomena in particular and, as such, we should not expect too much from DP kinds of considerations. Indeed, as I noted here, there are several problems with giving a detailed account how FL evolved. Let me remind you of some of the more serious issues, as CA adverts to similar ones.

The most obvious is the remove between cognitive powers and genetic ones. In particular, for an account of the evolution of FL we need a story of how minds are incarnated in brains and how brains are coded in genes. Why? Because evolution rejiggers genomes, which in turn grow brains, which in turn secrete cognition.  Sadly, every link of this chain is weak, most particularly in the domain of language (though the rest of cognition is not in much better shape as Lewontin’s famous piece (noted by CA) emphasizes). We really don’t know much about the genetic bases of brain development, nor do we know much about how brains realize FL. So, though we do know a fair bit about the cognitive structure of FL, we don’t have any really good linking hypotheses taking us from this to the brain and the genome, which are the structures that evolution manipulates to work its magic. In other words, to really explain how FL evolved (at least in detail) we need to account for how the brains structures that embody FL rose via alterations in our ancestors’ genomes (broadly construed to include epigenetic factors), and right now, though we have decent cognitive descriptions of FL, we have no good way of linking these up to brains and genes.

Second, if Lewontin is right (and I for one found his discussion entirely persuasive) then the prospects of giving standard evo accounts of how FL evolved will be largely nugatory due to the virtual impossibility of finding the relevant evidence, e.g. there really exist no “fossil” records to exploit and there is really nothing like our linguistic facility evident in any of our primate “cousins.” This makes constructing a standard evolutionary account empirically very dicey. 

In sum, things do not look good, so there arises the very reasonable question of what good DP thinking is for the practicing linguist. That’s the question CA asks. Here’s what I think (acknowledging that the problems CA notes are serious): Despite these problems, I still think that DP thinking is useful. Let me say why.

As I’ve noted (I suspect too many times) before (e.g. here) there is a tension between Plato’s Problem (PP) and DP. The former encourages packing as much as possible into UG to make the acquisition problem easier while the latter eschews this strategy to make the evolvability problem more tractable. Put another way: the more linguistic knowledge is given rather than acquired[2], the easier it is to account for why acquisition is as easy and effortless as it seems to be. However, the more there is packed into FL/UG the more challenging it is to explain how this knowledge could have evolved. This is the tension, and here is why I like DP: this tension is a creative one. These two problems together generate a very interesting theoretical problem: how to have one’s DP cake and PP’s too (i.e. how to allow for a solution to both PP and DP). I have suggested several strategies about how this might be accomplished that leads, IMO, to an interesting research program (e.g. here). My claim is that if you find this program attractive, then you need to take DP semi-seriously. What do I mean by “semi” here? Well, nobody expects to explain how FL actually evolved given how little we know about the relevant bridging assumptions (see above), but by thinking of the problem in tandem with PP we know the kinds of things we need to do (e.g. effectively eliminate the G internal modularity of characteristic of GB style theories and show that all the apparently different linguistic dependencies found outlined in the separate GB modules are effectively one and the same). That’s the first necessary step needed to reconcile DP and PP.[3]

The second, also motivated by DP, is yet more interesting (and challenging): to try and factor out those G operations, principles and primitives that are not linguistically specific. Thus, given DP we want not only a simpler (more elegant, more beautiful yadda, yadda, yadda) theory, but a particular kind of simpler etc. theory. We want one with as little linguistic specificity as possible. Maybe an example would help here.

The cyclic nature of derivations has long been a staple of GG theory. Ok, how to explain this? Earlier GG theory simply stipulated it: rules apply cyclically. The Minimalist Program (MP) has tried to offer slightly deeper motivation. There are two prominent accounts in the literature: the cycle as expression of the Extension Condition (EC) and the cycle as the expression of feature checking (Featural Cyclicity (FC)). FC is the idea that “bad” (unvalued or un-interpretable) features must be discharged quickly (e.g. in the phase or phrase that introduces them). EC says that derivations are monotonic (i.e. constituents that are inputs to a G operation must be constituents in the output of the operation).

There are various empirical linguistic-theory internal reasons that have been offered for preferring one or the other of these ideas. Both apply over a pretty general domain of cases and, IMO, it is hard to argue that either is in any relevant sense “simpler” than the other. However, IMO, the world would be a better place minimalistically were the EC the right way to conceptually ground the cycle. Why? Because it has the right generic feel to it because it looks like a very generic property of computations. In other words, IMO, it would be natural to find that cognitive rules systems in general are monotonic (information preserving) so that to find this holding of Gs would not be a surprise. FC, on the other hand, strikes me as relying on quite linguistically special assumptions about the toxicity of certain linguistic features and how quickly they need to be neutralized (some examples of very linguistically special properties: only probes contain toxic features, only phase heads have toxic features, toxic features must be very quickly eliminated, Gs contain toxic features at all). Personally, I find it hard to see FC and its special conception of features generalizing to other cognitive domains of third factor considerations. Of course I could be wrong (indeed given my track record, I am likely wrong!). What’s nice about a DP perspective is that it encourages us to try and make this kind of evaluation (i.e. ask how generic/specific a proposed operation/principle is) and, if my reasoning is on the right track, it suggests that we should try to maintain EC as a core feature precisely because it is plausibly not linguistic specific (i.e. not tied to special properties of human Gs).[4]

You could rightly object that such considerations are hardly dispositive, and I would agree. But if the above form of reasoning is even moderately useful, it suggests that DP considerations can have some theoretical utility. So, DP points to certain kinds of accounts and encourages one to develop theories with a certain look. It encourages the unification of GB modules and the elimination of linguistically specific features of Gs. It does this by highlighting the tension between DP and PP. One might construe all of this in simplicity terms, but from where I sit, DP encourages a very specific kind of simple, elegant theory (i.e. it favors some simple theories over others) and for this alone it earns its theoretical keep.[5]

CA has one other objection to DP that I would like to briefly touch on: DP encourages reductionism and reductionism is not a great methodological stance. I have two comments.

First, I am not personally against reduction if you can get it. My problem with it is that it very hard to come by, and I do not expect it to occur any day soon in my neck of the scientific woods. However, I do like unification and think that we should be encouraged to seek it out. As my good and great friend Elan Dresher once sagely noted: there are really only two kinds of linguistic papers. The first shows that two things that appear completely different are roughly the same. The second shows that two things that are roughly the same are in fact identical. I don’t know about linguistic papers in general, but this is a pretty good description of what a good many theory papers look like. Unification is the name of the game. So if by reduction we mean unification, then I am for it. Nothing wrong with it and a great deal that is right In fact, there is nothing wrong with real reduction either, if you can pull it off. Sadly, it is very very hard to do.

Second, I think that for MP to succeed then we should expect lots of reduction/unification. We will need to unify the modules (as noted above) and unify certain cognitive and computational operations. If we cannot manage this, then I would conclude that the main ideas/intuitions behind MP are untenable/unworkable/sterile. So, maybe unlike CA I welcome the reductionist/unificationist challenge.  Not only is this good science, DP is betting that it is the next important step for GG.

This said, there is something bad about reduction, and maybe this is what CA was worried about. Some reductionism takes it as obvious that the reducing science is more epistemologically privileged than the reduced one. But I see no reason for thinking that the metaphysics of the case (e.g. if A reduces to B) implies that the reducing theory is more solidly grounded epistemologically than the reduced one. So the fact that A might reduce to B does not mean that B is the better theory and that A’s results must bow to B’s theoretical diktats. Like all science, reduction/unification is a completed affair. It is best attempted when there is a reasonable body of doctrine in the theories to be related. I suspect that CA (and I am pretty sure Karthik) thinks that this is actually the main problem with DP at this time. It is premature (and if Lewontin is right, it will remain premature for the foreseeable future) precisely because of the problems I noted above concerning the assumptions required to bridge genes and cognition. My only response to this is that I am (slightly) more optimistic. I think DP considerations, even if inchoate, have purchase, though I agree that we should proceed carefully given how little we know of the details. So, we should make haste very very slowly.

Let me end. I think that DP has raised important questions for theoretical linguistics. Like most questions, they are yet somewhat undefined and fuzzy. Our job is to try and make them clearer and find ways of making them empirically viable. I believe that MP has succeeded in this to a degree. This said, CA (and Karthik) are right to point out the problems. Needless to say (or as I have said), I remain convinced that DPish considerations can and should play a role in how we develop theory moving forward for if nothing else (and IMO this is enough) it helps imbue the all to vague notions of elegance and simplicity with some local linguistic content. In other words, DP clarifies what kinds of elegant and simple theories of FL and UG we should be aiming for.





[1] Importantly, CA is not anti minimalist and believes that general criteria like elegance and simplicity (suitably contextualized for linguistics) can play a useful role. It’s DP that bothers CA not theoretical desiderata on syntactic theories.
[2] Indeed the more linguistic specific this knowledge is, the easier it is to explain the facility of acquisition despite the limitations in the PLD for G construction.
[3] I should be careful here: it is conceptually possible that one small genetic change creates a brain that looks GBish. Recall, we really don’t know how a fold here and there in a brain unlocks cognition. But, if we assume that the small thing that happened genetically to change our brains resulted in a simple new cognitive operation being added to our previous inventory, we are home free. This seems like a reasonable assumption, though it might be wrong. Right now, it is the simplest least convoluted assumption. So it is a good place to start.
[4] To say what need not be said: none of this implies that EC is right and FC wrong. It means that EC has more MPish value than FC does if you share my views. So, if we want to pursue an MPish line of inquiry, then EC is a very good way to go. Or, you should demand lots of good empirical reasons for rejecting it. DP, like PP, when working well, conceptually orders hypotheses making some more desirable than others all things being equal.
[5] These considerations were prompted by a very fruitful e-mail exchange with Karthik. He pointed out that notions like simplicity etc., in order to be useful, need to be crafted to apply to a domain of inquiry. The above amounts to suggesting that DP helps in isolating the right kind of simplicity considerations.