Comments

Showing posts with label Darwin's Problem. Show all posts
Showing posts with label Darwin's Problem. Show all posts

Friday, May 25, 2018

Three quintessentially minimalist projects

I was in Barcelona last week giving some lectures on the triumphant march forward of the Minimalist Program (MP). As readers may know, I believe that MP has been a success in its own terms in that it has gone a fair way towards answering the questions it first posed for itself, viz. Why do we have the FL we actually have and not some other? Others are more skeptical, but I believe that this is mainly because critics demand that MP address questions not central to its research mission. Of course, answering the MP question might leave many others untouched, but that hardly seems like a reason to disparage MP so much as a reason for pursuing other programs simultaneously. At any rate, this was what the lectures were about and I thank the gracious audience at the Autonomous University of Barcelona for letting me outline these views for their delectation. 

In getting these lectures into shape I started thinking about a question prompted by recent comments from Peter Svenonius (thx, see here). Peter thinks, if I understand him correctly, that the MP obsession (ok, my obsession) with Darwin’s Problem (DP) really adds nothing to the MP enterprise. Things could proceed more or less as we see them without this biological segue. This got me thinking about the following question: Which MP projects are largely motivated by DP concerns? I can think of three. They may well be motivated on other grounds as well. But they seem to me a direct consequence of taking the DP perspective on the emergence of FL seriously. This post is a first stab at enumerating these and explaining why I think they are intimately tied to DP (in much the same way that the P&P project was intimately tied to Plato’s Problem (PP)). So what are the three lines of inquiry? A warning: The order of discussion does not imply anything about their relative salience or importance to MP. I should note, that many of the points I make below, I have made elsewhere and before. So do not expect to be enlightened. This is probably more for my benefit than for yours.

First, Unification of the Modules (UoM). MP is based on the success of the P&P program, in particular the perceived success of GBish conceptions of FL/UG. Another way of saying this is that if you don’t think that GB managed to limn a fairly decent picture of the fine structure (i.e. the universals) of FL/UG then the MP project will seem to you to be, at best, premature and, at worst, hubristic. 

I believe that a lot of the problems that linguists have with MP has less to do with its failure to make progress in answering the MP question above, than with the belief that the whole project presupposes accepting as roughly right a deeply flawed conception of FL/UG (viz. roughly the GB conception). So, for example, if you don’t like classical case theory (and many syntacticians today do not) then you won’t like a project that takes it to be more or less accurate and tries to derive its properties from deeper principles. If you don’t think that classical binding theory is more or less correct then you won’t like a project that tries to reduce it to something else. The problem for many is that what MP presupposes (namely that GB was roughly correct, but not fundamental) is precisely what they believe ought to be very much up for grabs. 

I personally have a lot of sympathy for this attitude. However, I also think that it misses a useful feature of the presupposition. MP motivated unification/reduction doesn’t require that the GB description be correct so much that it be plausible enough (viz. that it be of roughly the right order of complexity) as to make deriving its properties a useful exercise (i.e. an exercise that if successful will provide a useful modelfor future projects armed with a more accurate conception of FL/UG). To put this another way: the GB conception of FL/UG has identified design features of FL/UG (aka, universals) that are theoretically plausible and empirically justifiable and so it is worth asking why principles such as theseshould govern the workings of our linguistic capacities. In Poeppel’s immortal phrase, these GB principles are the right grain size for analysis, have non-trivial empirical backing and so the exercise of showing why they should be part of FL would be doing something useful even should they fail to reflect the exact structure of FL.[1]

So, a central project animated by DP is unification of the disparate GB modules. And this is a very non-trivial project. As many of you know, GB attributes a rather high degree of internalmodularity to FL. There are diverse principles regulating binding vs control vs movement vs selection/subcategorization vs theta role assignment vs case assignment vs phrase structure. From the perspective of Plato’s Problem, the diversity of the modules does not much matter as their operations and principles are presumed to be innate (and hence not learned). In fact, the main impetus behind P&P architectures was to isolate the plausibly invariant features of Gs, and explain them by attributing them to the internal workings of FL thereby constraining the Gs FL produces to invariably respect these features. Thus the reason that Gs are always structure dependent is that FL has the property of only being able to construct Gs that are structure dependent. The reason that movement and binding require c-command is that FL imposes c-command as a condition on these diverse modular operations. The aim of GB was to identify and factor out the invariant properties of specific Gs and treat them as fixed features of FL/UG so that they did not have to be acquired on the basis of PLD (and a good thing too as there is not sufficient data in the PLD to fix them (aka PoS considerations apply to these)). The problem of G acquisition could then focus on the variable parts (where Gs differ) using the invariant parts as Archimedean fixed points for leveraging the PLD into specific Gs. That was the picture. And for this P&P project to succeed, it did not much matter how “complex” FL/UG was so long as it was innate. 

All of this changes, and changes dramatically, once one asks how this system couldhave arisen. Then the internal complexity matters, and matters a lot. Indeed, once one asks this question there is a great premium on simple FL architectures, with fewer modules and fewer disparate principles for the simpler the structure of FL, the easier it is to imagine how it mighthave arisen from the cognitive architecture of predecessors that did not have one. 

If this is correct, then one central MP project is to show that the diversity of the GB modules is only apparent and that they are only different reflections of the same underlying operations and principles. In other words, the project of unifying the modules is central to MP and it is central becauseof DP. A solution to DP requiresthat what appearsto be a very complex FL system (i.e. what GB depicts) is actually quite simple and what appearto be very different modules with different operations and regulative principles are really all reflections of the same underlying generative procedures. Why? Because short of this it will be impossible to explain how the system that GB describes could have arisen from a mind without it. 

This is entirely analogous, in its logic, to Plato’s Problem. How can kids acquire the Gs they do with the properties they have despite a poverty of the linguistic stimulus? Because much of what they know they do not have to learn. How could humans have evolved an FL from non-FL cognitive minds? Because FL minds are only a very small very simple step away from the minds that they emerged from and this requires that the modular complexity GB attributes to FL is only apparent. It’s what you get when you add to the contents of non-linguistic minds the small simple addition MP hypothesizes bridged the ling/non-ling gap.

Are there other plausible motives for such a project, the project of unifying the modules? Well perhaps. One might argue that an FL with unified modules are in some methodological sense better than one with non-unified ones. Something like a principle that says fewer modules are better than more. Again, I think that this is probably correct, but let’s face it, this kind of methodological Ockamist accounting is very weak (or at least perceived to be so). When push comes to shove data coverage (almost?) always trumps such niceties (remember the ceteris paribus clausethat always accompanies such dicta). So it is worth having a big empiricalfact of interest driving the agenda as well. And there are few facts bigger and heftier than the fact that FL arose from non-FL capable minds and it is easier to explain how this could havehappened if FL capable minds are only mildly different from non-FL capable minds and this means that the complex modularity that GB attributes to FL capable minds is almost certainly incorrect. That’s the line of argument. It rests on DPish assumptions and, to my mind, provides a powerful empirical motivation for module unification, which is what makes unification a central MP project.

It suggests a second related project: not only must the modules be unified, but the unification should makes use of the fewest possible linguistically proprietary operations and principles. In other words, linguistically capable minds, ones withFLs should be as minimally linguisticallyspecial as possible. Why? Because evolution proceeds most smoothly when there is minimal qualitative difference between the evolved states. If the aim is to explain how language ready minds appeared from non language ready minds than the fewer the differences between the two, the easier it will to be to account for the emergence of the former form the latter. If one assumes that what makes an FL mind language ready are linguistically special operations and principles then the fewer of these the better. In fact, in the best case there will be exactly a single relatively simple difference between the two, language ready minds just being non-language ready ones plus (at most) one linguistically special simple addition (the desideratum that it be simple motivated by the assumption that simple additions are more likely to become evolutionarily available than complex ones).[2]

So let’s assess: there are two closely related MP projects: unify the GB modules and unify them using largely non-linguistically proprietary operations and principles. How far has this project gotten? Well, IMO, quite far. Others are sure to disagree. But the projects though somewhat open textured have proven to be manageable and, the first in particular, has generated useful hypotheses (e.g. the Merge Hypothesis and extensions thereof, like the Movement Theory of Control and Construal), which even if wrong have the right flavor (Iknow, I know, this is self serving!). Indeed, IMO, trying to specify exactly where and how these theories go wrong (if they do, color me skeptical but I have dogs in these fights) and why they go wrong as they do, is a reasonable extension of the basic MP projects. It is a tribute to how little MP concerns drive contemporary syntax that such questions are, IMO, rarely broached. Let me rant a bit.

Darwin’s Problem (DP) currently enjoys as little interest among linguists today as Plato’s Problem (PP) does (and did, in earlier times). Indeed, from where I sit, even PP barely animates linguistic investigations. So, for example, people who study variation rarely ask how it might be fixed (though there are notable exceptions). Similarly, people who propose novel principles and operations rarely ask whether and how they might be integrated/unified with the rest of the features of FL. Indeed, most syntacticians take the basic apparatus as given and rarely critically examine it (e.g. how many people worry about the deep overlap between Agree and I-merge?). These are just not standard research concerns. IMO, sadly, most linguists could care less about the cognitive aspects of GG, let alone its possible bio-linguistic features. The object of study is language, not FL, and the technical apparatus is considered interesting to the degree that it provides a potentially powerful philological tool kit. 

Ok, so MP motivates two projects. There is one more, and it concerns variation. GB took variation to be bounded. It did this by conceiving UG as providing a finiteset of parameter values and conceived of language acquisition as fixing those parameters. So, even if the space of possible Gs is very large, for GB, it is finite. Now, given the linguistic specificityof the parameters, and given that GB treats them as internalto FL, the idea that variation is a matter of parameter setting proves to be a deep MP challenge. Indeed, I would go so far as to say, that ifMP is on the right track, thenFL does not contain a finite list of possible binary parameters and G acquisition cannot be a matter of parameter setting. It must be something else, something that is not specific to G acquisition. And this idea has caught on, big time. Let me explain.

I have many times mentioned the work by Berwick, Lidz and Yang on G acquisition. Each contains what is effectively a learning theory that constructs Gs from PLD using FL principles. It appears that this general idea is quite widely accepted now, with former parameter setting types (e.g. David Lightfoot) now arguing that “UG is open” and that there is “no evaluation of I-languages and no binary parameters” (1).[3]This view is much more congenial to MP as it removes the very specific parametric options fromFL and treats variation as entirely a “learning” problem. G learning is no different than other kinds, it is just aimed at Gs.[4]

Of course to make this work, will require specifying what kids come to the learning problem with, what kinds of data they exploit, and what the details of the G learning theory are. And this is hard. It requires more than pointing to differences in the PLD and attributing differences in Gs to these differences. However, this is a long way from an actual learning theory which specifies how PLD and properties of FL combine to give you a G. Not the least important fact is that there are many ways to generalize from PLD to Gs and kids only exploit some of these.[5]That said, if there is an MP “theory” of variation it will consist of adumbrating the innate assumptions the LAD uses to fix a particular G on the basis of PLD. To date, we have some interesting proposals (in particular from Lidz and Yang and their colleagues in syntax) but no overarching theory.

Interestingly, if this project can be made to fly, then it will also be the front end of an MP theory of variation. To date, the main focus of research has been on unifying and simplifying FL and trying to determine how much of FL is linguistically proprietary. However, there is no reason that the considerable current typological work on G variation shouldn’t feed into developing theories of learning aimed at explaining why we find the variation we do. It is just that thisproject is going to be very hard to execute well, as it will demand that linguists develop skills that are not currently part of standard PhD training, at least not in syntax (e.g. courses in stats, machine learning, and computation). But isn’t this as it should be? 

So, does taking MP seriously make a difference? Yes! It spawns three projects all animated by the MP problematic. These projects make sense in the context of trying to specify the internal structure of an FL that couldhave evolved from earlier minds. It suggests three concrete projects. So the programmatic aspects of MP are quite fecund, which is all that we can ask of a program.

And results? Well, here too I believe that we have made substantial progress as regards the first project, some as regards the third (though it is very very hard) and a little concerning the second.  IMO, this is not bad for 25 years and suggests that the DPish way of framing the MP issues has more than paid for itself.


[1]It’s worth adding that this sort of exercise is quite common in the real sciences. Ideal gases are not actual gases, planets are not point masses, and our universe may not be the only possible one but figuring out how they work has been very useful. 
[2]There is a lot of hand waving going on here. Thus, what evolves are genomes and what we are talking about here are phenotypic expressions thereof. We are assuming that simple genotypic difference reflect simple genetic differences. Who knows if this is right. However, it is the standard assumption for this kind of biological speculation so it would be a form of methodological dualism to treat it as suspect onlyin the linguistic case. See herefor discussion of this “phenotypic gambit” and its role in evolutionary thinking.
[3]See “Discovering New Variable Properties without Parameters,” in Massimo Piattelli-Palmarini and Simin Karimi, eds., “Parameters: What are they? Where are they?” Linguistic Analysis 41, special edition (2017).
            A very terse version of this view is advanced in Hornstein (2009) on entirely MP grounds. The main conceptual difference between approaches like Lightfoot’s and the one I advanced is that the former relies on the idea that “children DISCOVER variable properties of their language through parsing” (1), whereas I waved my hands and mumbled something about curve fitting given an enhanced representation provided by FL (see herefor slightly more elaboration).
[4]This folds together various important issues, the most important being that there is no overall evaluation metric for parameter setting. Chomsky argued that the shift from evaluation metrics to parameter setting modules increased the latters feasibility because applying global evaluation metrics to Gs is computationally intractable. I think Chomsky might have though that parameter setting is more localized than G evaluation and so will not require fancy learning theories. It turns, as Dresher and Kaye long ago noted, that parameter setting models have their own tractability issues unless the parameters can be set independently of one another. If they are not independent, problems quickly arise (e.g. it is hard to fix parameters once and for all). 
Furthermore, it is not clear to me that something like global measures of G fitness can be entirely avoided, though Lightfoot insists that they should be. The main reason for my skepticism is empirical and revolves around the question of whether the space of G options is scattered or not. At least in syntax, it seems that different Gs are kept relatively separate (e.g. bilinguals might code switch between French and English but they don’t syntactically blend them to get an “average” of the two in Frenglish. Why not?). This suggests that Gs enjoy an integrity and this is what keeps them cognitively apart. Bill Idsardi tells me that this might be less true on the sound side of things. But as regards the syntax, this looks more or less correct. If it is, then some global measure distinguishing different Gs might be required. 
I should add that more recently, if I recall correctly, Fodor and Sakas have argued that the evaluation metric cannot be completely dispensed with even on their “parsing” account.
[5]So, for example, invoking “parsing” as the driver behind acquisition does not do much unless one specifies howparsing works. Recall that standard parsers (e.g. the Marcus Parser) embody Gs that guide how it is that input data is analyzed. No G, no parsing. But if the aim is to explain how Gs are acquired then one cannot presuppose that the relevant G already exists as part of the parser. So what does a parse consist in in detail? This is a hard problem and it turns out that there are many factors that the child uses to analyze a string so as to recover a meaning. The MP project is to figure out what this is, not to name it.

Friday, December 2, 2016

What's a minimalist analysis

The proliferation of handbooks on linguistics identifies a gap in the field. There are so many now that there is an obvious need for a handbook of handbooks consisting of papers that are summaries of the various handbook summaries. And once we take this first tiny recursive step, as you all know, sky’s the limit.

You may be wondering why this thought crossed my mind. Well, it’s because I’ve been reading some handbook papers recently and many of those that take a historical trajectory through the material often have a penultimate section (before a rousing summary conclusion) with the latest minimalist take on the relevant subject matter. So, we go through the Standard Theory version of X, the Extended Standard Theory version, the GB version and finally an early minimalist and late minimalist version of X. This has naturally led me to think about the following question: what makes an analysis minimalist? When is an analysis minimalist and when not? And why should one care?

Before starting let me immediately caveat this. Being true is the greatest virtue an analysis can have. And being minimalist does not imply that an analysis is true. So not being minimalist is not in itself necessarily a criticism of any given proposal. Or at least not a decisive one. However, it is, IMO, a legit question to ask of a given proposal whether and how it is minimalist. Why? Well because I believe that Darwin’s Problem (and the simplicity metrics it favors) is well-posed (albeit fuzzy in places) and therefore that proposals dressed in assumptions that successfully address it gain empirical credibility. So, being minimalist is a virtue and suggestive of truth, even if not its guarantor.[1]

Perhaps I should add that I don’t think that anything guarantees truth in the empirical sciences and that I also tend to think that truth is the kind of virtue that one only gains slantwise. What I mean by this is that it is the kind of goal one attains indirectly rather than head on. True accounts are ones that economically cover reasonable data in interesting ways, shed light on fundamental questions and open up new avenues for further research.[2] If a story does all of that pretty well then we conclude it is true (or well on its way to it). In this way truth is to theory what happiness is to life plans. If you aim for it directly, you are unlikely to get it. Sort of like trying to fall asleep. As insomniacs will tell you, that doesn’t work.

That out of the way, what are the signs of a minimalist analysis (MA)? We can identify various grades of minimalist commitment.

The shallowest is technological minimalism. On this conception an MA is minimalist because it expresses its findings in terms of ‘I-merge’ rather than ‘move,’ ‘phases’ rather than ‘bounding nodes’/‘barriers,’ or ‘Agree’ rather than ‘binding.’ There is nothing wrong with this. But depending on the details there need not be much that is distinctively minimalist here. So, for example, there are versions of phase theory (so far as I can tell, most versions) that are isomorphic to previous GB theories of subjacency, modulo the addition of v as a bounding node (though see Barriers). The second version of the PIC (i.e. where Spell Out is delayed to the next phase) is virtually identical to 1-subjacency and the number of available phase edges is identical to the specification of “escape hatches.”

Similarly for many Agree based theories of anaphora and/or control. In place of local coindexing we express the identical dependency in terms of Agree in probe/goal configurations (antecedents as probes, anaphors as goals)[3] subject to some conception of locality. There are differences, of course, but largely the analyses inter-translate and the novel nomenclature serves to mask the continuity with prior analyses of the proposed account. In other words, what makes such analyses minimalist is less a grounding in basic features of the minimalist program, then a technical isomorphism between current and earlier technology. Or, to put this another way, when successful, such stories tell us that our earlier GB accounts were no less minimalist than our contemporary ones. Or, to put this yet another way, our current understanding is no less adequate than our earlier understanding (i.e. we’ve lost nothing by going minimalist). This is nice to know, but given that we thought that GB left Darwin’s Problem (DP) relatively intact (this being the main original motivation for going Minimalist (i.e. beyond explanatory adequacy) then analyses that are effectively the same as earlier GB analyses likely leave DP in the same opaque state. Does this mean that translating earlier proposals into current idiom is useless? No. But such translations often make a modest contribution to the program as a whole given the suppleness of current technology.

There is a second more interesting kind of MA. It starts from one of the main research projects that minimalism motivates. Let’s call this “reductive” or “unificational minimalism” (UM). Here’s what I mean.

The minimalist program (MP) starts from the observation that FL is a fairly recent cognitive novelty and thus what is linguistically proprietary is likely to be quite meager. This suggests that most of FL is cognitively or computationally general, with only a small linguistically specific residue. This suggests a research program given a GB backdrop (see here for discussion). Take the GB theory of FL/UG to provide a decent effective theory (i.e. descriptively pretty good but not fundamental) and try to find a more fundamental one that has these GB principles as consequences.[4] This conception provides a two pronged research program: (i) eliminate the internal modularity of GB (i.e. show that the various GB modules are all instances of the same principles and operations (see here)) and (ii) show that of the operations and principles that are required to effect the unification in (i), all save one are cognitively and/or computationally generic. If we can successfully realize this research project then we have a potential answer to DP: FL arose with the adventitious addition of the linguistically proprietary operation/principle to the cognitive/computational apparatus the species antecedently had.

That’s the main contours of the research program. UF concentrates on (i) and aims to reduce the different principles and operations within FL to the absolute minimum. It does this by proposing to unify domains that appear disparate on the surface and by reducing G options to an absolute minimum.[5] A reasonable heuristic for this kind of MA is the idea that Gs never do things in more than one way (e.g. there are not two ways (viz. via matching or raising) to form relative clauses). This is not to deny different surface patterns obtain, only that they are not the products of distinctive operations.

Let me put this another way: UM takes the GB disavowal of constructions to the limit. GB eschewed constructions in that it eliminated rules like Relativization and Topicalization, seeing both as instances of movement. However, it did not fully eliminate constructions for it proposed very different basic operations for (apparently) different kinds of dependencies. Thus, GB distinguishes movement from construal and binding from control and case assignment from theta checking. In fact, each of the modules is defined in terms of proprietary primitives, operations and constraints. This is to treat the modules as constructions. One way of understanding UM is that it is radically anti-constructivist and recognizes that all G dependencies are effected in the same way. There is, grammatically speaking, only ever one road to Rome.

Some of the central results of MP are of this ilk. So, for example, Chomsky’s conception of Merge unifies phrase structure theory and movement theory. The theory of case assignment in the Black Book unifies case theory and movement theory (the latter being just a specific reflex of movement) in much the way that move alpha unifies question formation, relativization, topicalization etc. The movement theory of control and binding unifies both modules with movement. The overall picture then is one in which binding, structure building, case licensing, movement, and control “reduce” to a single computational basis. There aren’t movement rules versus phrase structure rules versus binding rules versus control rules versus case assignment rules. Rather these are all different reflexes of a single Merge effected dependency with different features being licensed via the same operation. It is the logic of On wh movement writ large.

There are other examples of the same “less is more” logic: The elimination of D-structure and S-structure in the Black Book, Sportiche’s recent proposals to unify promotion and matching analyses of relativization, unifying reconstruction and movement via the copy theory of movement (in turn based on a set theoretic conception of Merge), Nunes theory of parasitic gaps, and Sportiche’s proposed elimination of late merger to name five. All of these are MAs in the specific sense that they aim to show that rich empirical coverage is compatible with a reduced inventory of basic operations and principles and that the architecture of FL as envisioned in GB can be simplified and unified thereby advancing the idea that a (one!) small change to the cognitive economy of our ancestors could have led to the emergence of an FL like the one that we have good (GB) evidence to think is ours.  Thus, MAs of the UM variety clearly provide potential answers to the core minimalist DP question and hence deserve their ‘minimalist’ modifier.

The minimalist ambitions can be greater still. MAs have two related yet distinct goals. The first is to show that svelter Gs do no worse than the more complex ones that they replace (or at least don’t do much worse).[6] The second is to show that they do better. Chomsky contrasted these in chapter three of the Black Book and provided examples illustrating how doing less with more might be possible. I would like to mention a few by way of illustration, after a brief running start.

Chomsky made two methodological observations. First, if a svelter account does (nearly) empirically as well as a grosser one then it “wins” given MP desiderata. We noted why this was so above regarding DP, but really nobody considers Chomsky’s scoring controversial given that it is a lead footed application of Ockham. Fewer assumptions are always better than more for the simple reason that for a given empirical payoff K an explanation based on N assumptions leaves each assumption with greater empirical justification than one based on N+1 assumptions. Of course, things are hardly ever this clean, but often they are clean enough and the principle is not really contestable.[7]

However, Chomsky’s point extends this reasoning beyond simple assumption counting. For MP it’s not only the number of assumptions that matter but their pedigree. Here’s what I mean.  Let’s distinguish FL from UG. Let ‘FL’ designate whatever allows the LAD to acquire a particular GL based on PLDL. Let ‘UG’ designate those features of FL that are linguistically proprietary (i.e. not reflexes of more generic cognitive or computational operations). A MA aims to reduce the UG part of FL. In the best case, it contains a single linguistically specific novelty.[8] So, it is not just a matter of counting assumptions. Rather what matters is counting UG (i.e. linguistically proprietary) assumptions. We prefer those FLs with minimal UGs and minimal language specific assumptions.

An example of this is Chomsky’s arguments against D-structure and S-structure as internal levels. Chomsky does not deny that Gs interface with interpretive interfaces, rather he objects to treating these as having linguistically special properties.[9] Of course, Gs interface with sound and meaning. That’s obvious (i.e. “conceptually necessary”). But this assumption does not imply that there need be anything linguistically special about the G levels that do the interfacing beyond the fact that they must be readable by these interfaces. So, any assumption that goes beyond this (e.g. the theta criterion) needs defending because it requires encumbering FL with UG strictures that specify the extras required. 

All of this is old hat, and, IMO, perfectly straightforward and reasonable. But it points to another kind of MA: one that does not reduce the number of assumptions required for a particular analysis, but that reapportions the assumptions between UGish ones and generic cognitive-computational ones. Again, Chomsky’s discussions in chapter 3 of the Black Book provide nice examples of this kind of reasoning, as does the computational motivation for phases and Spell Out.

Let me add one more (and this will involve some self referentiality). One argument against PRO based conceptions of (obligatory) control is that they require a linguistically “special” account of the properties of PRO. After all, to get the trains to run on time PRO must be packed with features which force it to be subject to the G constraints it is subject to (PRO needs to be locally minimally bound, occurs largely in non-finite subject positions, and  has very distinctive interpretive properties). In other words, PRO is a G internal formative with special G sensitive features (often of the possibly unspecified phi-varierty) that force it into G relations. Thus, it is MP problematic.[10] Thus a proposal that eschews PRO is prima facie an MA story of control for it dispenses with the requirement that there exists a G internal formative with linguistically specific requirements.[11] I would like to add, precisely because I have had skin in this game, that this does not imply that PRO-less accounts of control are correct or even superior to PRO based conceptions.  No! But it does mean that eschewing PRO has minimalist advantages over accounts that adopt PRO as they minimize the UG aspects of FL when it comes to control.

Ok, enough self-promotion.  Back to the main point. The point is not merely to count assumptions but to minimize UGish ones. In this sense, MAs aim to satisfy Darwin more than Ockham. A good MA minimizes UG assumptions and does (about) as well empirically as more UG encumbered alternatives. A good sign that a paper is providing an MA of this sort, is manifest concern to minimize the UG nature of the principles assumed.

Let’s now turn to (and end with) the last most ambitious MA: it is one that not merely does (almost) as well as more UG encumbered accounts, but does better. How can one do better. Recall that we should expect MAs to be more empirically brittle than less minimalist alternatives given that MP assumptions generally restrict an account’s descriptive apparatus.[12]  So, how can a svelter account do better? It does so by having more explanatory oomph (see here). Here’s what I mean.

Again, the Black Book provides some examples.[13] Recall Chomsky’s discussion of examples like (1) with structures like (2):

(1)  John wonders how many pictures of himself Frank took
(2)  John wonders [[how many pictures of himself] Frank took [how many pictures of himself]]

The observation is that (1) has an idiomatic reading just in case Frank is the antecedent of the reflexive.[14] This can be explained if we assume that there is no D-structure level or S-structure level. Without these binding and idiom interpretation must be defined over that G level that is input to the CI interface. In other words, idiom interpretation and binding are computed over the same representation and we thus expect that the requirements of each will affect the possibilities of the other.

More concretely, to get the idiomatic reading of take pictures requires using the lower copy of the wh phrase. To get the John as potential antecedent of the reflexive requires using the higher copy. If we assume that only a single copy can be retained on the mapping to CI, this implies that if take pictures of himself is understood idiomatically, Frank is the only available local antecedent of the reflexive. The prediction relies on the assumption that idiom interpretation and binding exploit the same representation. Thus, by eliminating D-structure, the theory can no longer make D-structure the locus of idiom interpretation and by eliminating S-structure, the theory cannot make it the locus of binding. Thus by eliminating both levels the proposal predicts a correlation between idiomaticity and reflexive antecedence.

It is important to note that a GBish theory where idioms are licensed at D-structure and reflexives are licensed at S-structure (or later) is compatible with Chomsky’s reported data, but does not predict it. The relevant data can be tracked in a theory with the two internal levels. What is missing is the prediction that they must swing together. In other words, the MP story explains what the non-MP story must stipulate. Hence, the explanatory oomph. One gets more explanation with less G internal apparatus.

There are other examples of this kind of reasoning, but not that many.  One of the reasons I have always liked Nunes’ theory of parasitic gaps is that it explains why they are licensed only in overt syntax. One of the reasons that I like the Movement Theory of Control is that it explains why one finds (OC) PRO in the subject position of non-finite clauses. No stipulations necessary, no ad hoc assumptions concerning flavors of case, no simple (but honest) stipulations restricting PRO to such positions. These are minimalist in a strong sense.

Let’s end here. I have tried to identify three kinds of MAs. What makes proposals minimalist is that they either answer or serve as steps towards answering the big minimalist question: why do we have the FL we have? How did FL arise in the species?  That’s the question of interest. It’s not the only question of interest, but it is an important one. Precisely because the question is interesting it is worth identifying whether and in what respects a given proposal might be minimalist. Wouldn’t it be nice if papers in minimalist syntax regularly identified their minimalist assumptions so that we could not not only appreciate their empirical virtuosity, but could also evaluate their contributions to the programmatic goals.


[1] If pressed (even slightly) I might go further and admit that being minimalist is a necessary condition of being true. This follows if you agree that the minimalist characterization of DP in the domain of language is roughly accurate. If so, then true proposals will be minimalist for only such proposals will be compatible with the facts concerning the emergence of FL. That’s what I would argue, if pressed.
[2] And if this is so, then the way one arrives at truth in linguistics will plausibly go hand in hand with providing answers to fundamental problems like DP. This, proposals that are minimalist may thereby have a leg up on truth. But, again, I wouldn’t say this unless pressed.
[3] The agree dependency here established accompanied by a specific rule of interpretation whereby agreement signals co-valuation of some sort. This, btw, is not a trivial extra.
[4] This parallels the logic of On wh movement wrt islands and bounding theory. See here for discussion.
[5] Sportiche (here) describes this as eliminating extrinsic theoretical “enrichments” (i.e. theoretical additions motivated entirely by empirical demands).
[6] Note a priori one expects simpler proposals to be empirically less agile than more complex ones and to therefore cover less data. Thus, if a cut down account gets roughly the same coverage this is a big win for the more modest proposal.
[7] Indeed, it is often hard to individuate assumptions, especially given different theoretical starting points. However (IMO surprisingly), this is often doable in practice so I won’t dwell on it here.
[8] I personally don’t believe that it can contain less for it would make the fact that nothing does language like humans do a complete mystery. This fact strongly implies (IMO) that there is something UGishly special about FL. MP reasoning implies that this UG part is very small, though not null. I assume this here.
[9] That’s how I understand the proposal to eliminate G internal levels.
[10] It is worth noting that this is why PRO in earlier theories was not a lexical formative at all, but the residue of the operation of the grammar. This is discussed in the last chapter here if you are interested in the details.
[11] One more observation: this holds even if the proposed properties of PRO are universal, i.e. part of UG. The problem is not variability but linguistic specificity.
[12] Observe that empirical brittleness is the flip side of theoretical tightness. We want empirically brittle theories.
[13] The distinction between these two kinds of MAs is not original with me but clearly traces to the discussion in the Black Book.
[14] I report the argument. I confess that I do not personally get the judgments described. However, this does not matter for purposes of illustration of the logic.

Monday, February 29, 2016

Hauser reviews "Why only Us"

Though my interest in Darwin's Problem (DP) is deep, my "expertise," such as it is, is restricted to the logic of the argument. The logic is well known: (i) hierarchical recursion is a distinctive hallmark of human I-languages, (ii) there is no evidence that any other animal displays such recursive powers, (iii) what nonlinguistic evidence there is concerning such powers in humans is of rather recent vintage (roughly 100kya), (iv) logically speaking recursion is an all or nothing affair. The conclusion from (i)-(iv) is that something simple occurred roughly 100kya that in combination with the nonlinguistic cognitive and computational powers extant at the tome in our ancestors allowed for this human species specific capacity to emerge. That's the logic. 

It looks a lot like the logic of PoS arguments in that it starts from a specification of the capacity own interest and argues backwards to the causal mechanisms that could produce it. In other words, just as GGers investigate human linguistic cognition by first describing what it is that has been acquired and inferring from this what the system of acquisition must look like, so too we investigate evolutionary possibilities concerning language by first specifying what it is that has evolved. Sadly, this is not the general methods of investigation. Empiricists (both in psychology and evolutionary biology) seem to think that the direction of argument should be reversed: given that we know what learning and evolution is they conclude that a species specific FL or a species specific characteristics cannot grow in human minds nor have evolved there. The arguments they provide are awful, and not for sophisticated reasons. They are awful because they fail to address what we know about human linguistic capacities. They are based on the premiss, in other words, that the facts that GG has discovered over the last 60 years are bogus. As anyone but a flat earther knows this to be wrong,…[1]

Discussion at this level is where my expertise ends. However, DP is more than just logically interesting (it has potential empirical ramifications) and there is more to the logic than I outlined above (Merge may not be the only unique capacity our linguistic facility manifests (see the review below)). Berwick and Chomsky's Why Only Us (WOU) goes into these issues, and Marc Hauser has thought about them hard. So what better way to get into them more deeply than to ask Marc to do a blog post on WOU. He graciously agreed. Here it is. 

*****

Berwick & Chomsky’s Why only us (2016):
Challenges to the what, when, and why?

Marc D. Hauser

Why only us  [WOU] is a wonderful, slim, engaging, and clearly written book by Robert Berwick and Noam Chomsky.  From the authors’ perspective, it is a book about language and evolution. And of course it is.  However, I think it is actually about something much bigger.  It is an argument about the evolution of thought itself, with language being not only one form of thought, but a domain that can impact thought itself, in ways that are truly unique in the animal kingdom.  Seen in this light, WOU provides a framework for thinking about the evolution of thought and a challenge to Darwin’s claim that the human mind is only quantitatively different from other animals. Since this is an idea that I have championed (Hauser, 2009), I am of course a bit partial! Let me unpack all of this by working through Berwick and Chomsky’s arguments, especially those where we don’t quite agree. 

One caveat up front:  as I have written before, including with Berwick and Chomsky (Hauser et al., 2014), I am not convinced that the ideas put forward here or in WOU are testable: animal capacities are far too impoverished to shed any comparative light on the evolution of human language, and the hominid fossil record is either silent or too recent to be of interest. My goal here, therefore, is to focus on the fascinating ideas raised in WOU,  leaving to the side how or whether such ideas might be confronted by significant empirical tests. 

One of the essential moves in WOU is to argue that MERGE —the simplest recursive operation — is the bedrock of our capacity for infinite expression by finite means, one that generates hierarchical structure. Because no other animal has MERGE, and because MERGE  is simple and the essence of language, the evolutionary process may well have occurred rapidly, appearing suddenly in only one species: modern humans or Homo sapiens sapiens (Hss).  To accept this argument, you have to accept at least five premises:

            1- MERGE is the essence of language
            2- No other animal has MERGE
            3- No other hominid has MERGE
  4- Due to the simplicity of MERGE, it could evolve quickly, perhaps
 due to mutation
            5- Because you either have or don’t have MERGE (there is no
                  demi-MERGE), there is no option for proto-language.

I accept 2 because the comparative literature shows nothing remotely like MERGE.  Whether one looks at data from natural communication, artificial language learning experiments, or animal training studies with human language or language-like tokens, there is simply no evidence of anything remotely recursive.  As Berwick and Chomsky note, the closest one gets is the combinatoric gymnastics observed in birdsong, but these are neither recursive nor do they generate hierarchical structures that shape or generate the variety of meaningful expressions observed in all human languages. 

I also accept 3, though here we don’t really have the evidence to say one way or the other, and even if we did, and it turned out that say Neanderthals had MERGE, it wouldn’t really make much of a difference to the argument.  That is, the fossil record for Neanderthal, though richer than we once thought, says nothing about recursive operations, and nor for that matter does the fossil record for Hss. Both records show interesting signs of creative thought — a topic to which I return — but nothing that would indicate recursive thought or expression.  If evidence emerges that Neanderthals had MERGE, that would simply push back the date of origin for Berwick and Chomsky’s evolutionary account, without changing the core details.    

Let’s turn to 1, 4 and 5 then.  What is interesting about the core argument in WOU is that although Berwick and Chomsky place significant emphasis on MERGE, they fully acknowledge that the recursive machinery must interface with the Conceptual-Intensional system on the one hand, and with the Sensory-Motor system on the other.  However, once one acknowledges the non-trivial roles of CI, SM, and the interfaces, while also recognizing the unique properties of each of these systems, it is no longer possible to accept premise 4, and challenges arise for premise 5.  This analysis lays open the door to some fascinating possibilities, many of which might be explored empirically. I consider a few next.

Berwick and Chomsky devote some of the early material of WOU to review work on vocal imitation in songbirds, including comparative genetic and neurobiological data.  In some ways, the songbird system is a lovely example because the work is exquisitely detailed and shows some nice parallels with our own.  In particular, songbirds learn their song in some of the same ways as young children learn language, including evidence of an innate system that constrains both the timing and material acquired.  However, there are elements of the songbird system that are strikingly different from our own, not mentioned in WOU, but when acknowledged, tell an even more interesting tale about the evolution of Hss — one that is at the same time supportive of the uniqueness claims in WOU while also raising questions about the nature of the uniqueness claim.  Specifically, the songbird system is a striking example of extreme modularity.  The capacity of a songbird to imitate or learn its species-specific song is not a capacity that extends to other calls in its vocal repertoire, nor to any visual display. That is, a songbird can imitate the song material it hears, but nothing else.  Not so for our species, where the capacity to imitate is amodal, or at least bimodal, with sounds and actions copied readily, and from birth. This disconnect from sensory modality is a trademark of human thought, and of course, is a critical feature of our language faculty:  at virtually all levels of detail, including syntax, semantics, phonology, acquisition, and pragmatics, there are no differences between signed and spoken languages. No other animal is like this.  Whether we observe songbirds, dolphins, or non-human primates, an individual born deaf does not emerge with a comparably expressive visual system of communication.  The systems of communicative expression are intimately tied to the modality, such that if one modality is damaged, other modalities are incapable of picking up the tab.  The fact that our language, and even more broadly, our thoughts, are detached from modality, suggests a fundamental reorganization in our representations and computations.  This takes us to CI, SM, MERGE and the interfaces.

Given the modularity of the songbird system, and the lack of imitative capacities in non-human primates, we also need an account of how a motor system capable of imitating sounds and actions evolved.  This is an account of how SM evolved, but also, about how and when SM interfaced with CI and MERGE.  There is virtually no evidence on offer, and it is hard to imagine what kind of evidence could emerge. For example, the suggestion that Neanderthals had a hyoid bone like Hss is interesting, but doesn’t tell us what they were doing with it, whether it was capable of being deployed in vocal imitation, and thus, of building up the lexicon.  And of course, we don’t know whether or how it was connected to CI or MERGE.  But whatever we discover about this account, it showcases the importance of understanding the evolution of at least one unique property of SM.

When we turn to CI, and in particular, lexical or conceptual atoms, we know extremely little about them, even in fully linguistics human adults.  Needless to say, this makes comparative and developmental work difficult.  But one observation seems fairly uncontroversial: many of our concepts are completely detached from sensory experiences, and thus can’t be defined by them. If we take this as a starting point, we can ask: do animals have anything remotely like this?  On one reading of Randy Gallistel’s elegant work, the answer is “Yes.”  All of the empirical work on number, time and space in animals suggests that such concepts are either not linked to or defined by a particular modality, or minimally, can be expressed in multiple modalities.  Similarly, there is evidence that animals are capable of representing some sense of identity or sameness that is not tied to a modality.  If this is right, and even if these concepts are not as abstract as ours, they suggest a potential comparative approach that at this point, seems closed off for our recursive capacity.   Having a comparative evolutionary landscape of inquiry not only aids in our analyses, it also raises a challenge to premises 4 and 5, as well as to Richard Lewontin’s comment (supported by Berwick and Chomsky) that we can’t study or understand the evolution of cognition.  Let me take a small detour to describe a gorgeous series of studies on the evolution of cognition to show what can and has been done, and then return to premises 4 and 5.

In most monogamous species, the male and female share the same home range or territory.  In polygynous species, in contrast, there are several females associated with one male, and thus, the male’s home range area encompasses all of the smaller female home ranges.  Based on this observation, Steve Gaulin and his colleagues (Gaulin & Wartell, 1990; Jacobs, Gaulin, Sherry, & Hoffman, 1990; Puts, Gaulin, & Breedlove, 2007) predicted that the spatial abilities of a monogamous vole would show no sex differences, whereas males would show greater abilities than females in a closely related polygynous vole species. Using a maze running task to test for spatial capacity, results provided strong support for the prediction.  Further, the size of the hippocampus — an area of the brain known to play an important role in spatial navigation — was significantly larger in males of the polygynous species when contrasted with females, whereas no sex differences were found for the monogamous species. This, and several other examples, reveal how one can in fact study the evolution of cognition. Lewontin is, I believe, flatly wrong.

Back to premises 4 and 5. If nonhuman animals have abstract, amodal concepts — as some authors suggest —  then we have a significant line of empirical inquiry into the evolution of this system.  If our concepts are unique — as authors such as Berwick and Chomsky believe —  then there may not be that many empirical options. Perhaps Neanderthals have such concepts, perhaps not. Either way, the evolutionary timescale is short, and the evidence thus far, relatively thin.  On either account, however, there is the pressing need to understand the nature of such concepts as they bear on what I believe is the most interesting side effect of this discussion, and the issues raised in WOU.  In brief, if one concedes that what is unique about language, and thus, its evolutionary history, is MERGE, CI, SM and the interfaces, then a different issue emerges:  are these four ingredients unique to language or part of all aspects of human thought?  Said differently, perhaps WOU is really an account of how our uniquely human system of thought evolved, with language being only one domain in terms of its internal and external systems of expression. Berwick and Chomsky often refer to our Language of Thought, as the core of language, and what is our most dominant use of language: internal thought.  On this view, externalization of this system in expressed language is not at the core of the evolutionary account.  On the one hand, I agree. On the other hand, I think the use of the term of Language of Thought or LOT has confused the issue because of the multiple uses of the word “language.” If the essence of the argument in WOU is about the computations and representations of thought, with linguistic thought being one flavor, then I would suggest we call this system the Logic of Thought.  I suggest this substitution of L-words for two reasons.  Language of Thought implies that the system is explicitly linguistic, and I don’t believe it is.  Further, I think Logic of Thought better captures the abstract nature of the ingredients, including both the recursive operations, concepts, motor routines, and interfaces. 

The Logic of Thought, I would argue, is uniquely human, and underpins not only language, but many other domains as well.  It explains, I believe, why actions that appear similar in other animals are actually not similar at all.  It also provides the ultimate challenge to Darwin’s argument that there is continuity in mental thought between humans and other animals, with differences attributable to quantity as opposed to quality.  In contrast, if the ideas discussed here, and ultimately raised by Berwick and Chomsky are right, then it is the Logic of Thought that is unique to humans.  The Logic of Thought includes all four ingredients: MERGE, CI, SM, and the interfaces. How these components are articulated in different domains is fascinating in its own right, and raises several additional puzzles. For example, if MERGE is the simplest recursive operation, is it one neural mechanism that interfaces with different, domain-specific concepts and actions, or were merge like circuits effectively cloned repeatedly, each subserving a different domain?  The first possibility suggests that damage to this singular MERGE circuit would reveal deficits in multiple domains.  The second option suggests that damage to the MERGE circuit in one domain would only reveal deficits in this domain. To my knowledge, there is no evidence of neuropsychological deficits or imaging studies that point to the nature or distribution of such recursive circuitry. 

In sum, WOU is really a terrific book. It is thought provoking and clear.  What more could you want?  My central challenge is that it paints an evolutionary account that can only work if the essence of language is simple, restricted to MERGE.  But language is much more than this.  As such, there has to be more to the evolutionary process.  By raising these issues, I believe Berwick and Chomsky have challenged us to think about another option, one that preserves their title, but focuses on the logic of thought.  Why only us? Much to think about.


Gaulin, S. J., & Wartell, M. S. (1990). Effects of experience and motivation on symmetrical-maze performance in the prairie vole (Microtus ochrogaster). Journal of Comparative Psychology, 104(2), 183–189.
Hauser, M. D. (2009). The possibility of impossible cultures. Nature, 460, 190–196.
Hauser, M. D., Yang, C., Berwick, R. C., Tattersall, I., Ryan, M. J., Watumull, J., et al. (2014). The mystery of language evolution. Frontiers in Psychology, 5(401), 1–12.
Jacobs, L. F., Gaulin, S. J., Sherry, D. F., & Hoffman, G. E. (1990). Evolution of spatial cognition: sex-specific patterns of spatial behavior predict hippocampal size. Proceedings of the National Academy of Sciences, 87(16), 6349–6352.

Puts, D. A., Gaulin, S. J., & Breedlove, S. M. (2007). Sex differences in spatial ability: evolution, hormones and the brain. Evolutionary Cognitive Neuroscience. MIT Press, pp 329-379.





[1] Talking about flat-earthers and limitless ignorance, it is worth comparing Hauser’s review below to one by V. Evans (yes, that V. Evans) here. The best that I can say is that it is consistent with what I have come to expect from Evans’ work, (viz. it is not worse than his other recent output (tant pis)).