Comments

Showing posts with label GB. Show all posts
Showing posts with label GB. Show all posts

Monday, June 6, 2016

Theory, again

It’s the start of the summer so it’s time to return to some pet peeves. Here’s the fortune cookie version of the history of Generative Grammar (GG): we have moved from the study of Gs to the study of possible Gs to the study of possible FL/UGs. The (bulk of the) earliest work in GG (e.g. Syntactic Structures, LSLT, the Standard Theory) aimed to adumbrate the kinds of rules that Gs contain by studying the actual recursive mechanisms that specific Gs embody. The next stage aimed to adumbrate not only the rules that Gs actually contain but also the principles restricting the kinds of operations a G could contain (this is what UG in GB was all about). Minimalism builds on the results of all of this earlier research and aims to limn the contours of a possible human Faculty of Language (FL). It, in effect, addresses the question: why do we have the FL/UG we in fact have rather than some conceivable others?

As is obvious (but this won’t stop me form pressing the point) these research questions are closely inter-related with connections in two directions.  First, each later question starts from answers provided by the earlier one. It’s pointless to wonder about possible rules without some candidate actual ones and it is futile to investigate the limits of FL/UG without some candidate principles of FL/UG. Second, answers to later questions limit the range of answers to earlier ones. If a rule is not FL/UG possible then a particular G cannot contain such a rule and if a principle is not a possible principle of FL/UG then no FL/UG can contain that kind of principle.

So, two observations: first, the three kinds of questions above are importantly different even if closely related (as such, they must be kept logically and conceptually distinct). Second, the dialectic from answer to answer moves in both directions from “lower” level to “higher” and back again. “Lower” and “higher” are not intended as evaluative. They are just used to mark the conceptual flow noted above.

Here’s a third observation: despite their interconnections, the methods used to study each of these questions are partially autonomous from each other. People who study particular Gs can do useful work without resort to the accepted/proposed principles of FL/UG and those interested in the universal properties of Gs (i.e. the structure of FL/UG) can get a good way into this problem without bothering too much with minimalist concerns. The methods used to investigate all three questions partially overlap, but the criteria for success are not the same and even some of the detailed kinds of arguments advanced can have somewhat different flavors. So not only are the questions different, but progress on addressing them is somewhat independent of progress in addressing the others. Just as there is no discovery procedure for Gs (no reduction of later levels to earlier ones), there is none for theories of GGG (no requirement that later questions uncritically respect the answers provided to earlier ones). The questions are related to one another in roughly the way that levels in a G are: they take in one another’s washing in complicated ways.

Why do I mention this? Because I believe that some of the unease in current syntax stems from misunderstanding what question is being addressed by a particular proposal and thus what counts as evidence for or against it. Or to put this another way: if the above is a roughly correct characterization of the conceptual GG landscape, then it is important to understand that many proposals, especially “higher” ones, are hidden conditionals. For example, minimalist proposals are of the form: Given that such and such is a plausible (better still, actual) principle of FL/UG then so and so is why this kind of principle obtains rather than others.

If this is so, then there are two ways to reject a specific proposal: (i) argue against the conditional as a whole or (ii) argue only against the antecedent. The former denies that the deductive link between premise and conclusion holds. The latter denies the relevance of the deductive link even if it does hold. As I see it, most critiques of minimalist proposals are of the second kind. They deny that what is taken as given should be so taken because the premise is empirically suspect. In other words, many objections are actually objections to the underlying “GB” principle being “explained” (and hence assumed) in minimalist terms rather than the explanation itself.[1]  These critiques deny the utility of the explanation rather than question its deductive validity. Thus they conclude that showing how to deduce the principle from more general considerations is valueless because the premise is false. IMO, this conclusion is unfortunate and it reflects a general disdain for theory characteristic of much work in contemporary “theoretical” syntax. Let me vent a bit (again).

In the real sciences, a lot of time is spent trying to find ways of tying together seemingly disparate principles. It really isn’t easy to show that two principles that look different are nonetheless fundamentally the same. And the problem is in large part conceptual. And one way that conceptual problems are investigated is by (often radically) simplifying them. Of course, the hope is that the simplification will preserve many of the core features of interest and so the simplification can “scale up” as we make the premises more realistic. Such simplifications often rest on “stylized” facts that are acknowledged to be (ahem) “incomplete” (aka: false). However, investigating such empirically inadequate simple problems based on stylized facts is often a vital step in advancing understanding even though the premises might be false (as simplifications almost always are). The same should hold true in syntax.

Btw, this sort of investigation (largely pencil and paper kind of stuff) is what is commonly called ‘theoretical.’ Theoretical work consists in investigating how simple concepts can be related to produce theories with rich deductive structure. Theory places a premium on (i) the reasonableness (rather than the truth) of the basic simplification (i.e. the rough accuracy of the stylized facts), (ii) the naturalness of the assumed basic concepts and (iii) the depth of the deductive structure that results.

A good example of this in GG is Chomsky’s recent proposals concerning Merge. It runs roughly as follows: if you assume that Merge is a very simple binary operation that takes two syntactic objects (SOs) and combines them into a set of those SOs (i.e. If A is an So and B is an SO then {A,B} is an SO) then you can generate objects with unbounded hierarchical structure with the following “nice” properties: Merge must be structure dependent (linear order irrelevant to syntax) and cyclic (e.g. no lowering rules), phrase structure building and movement are two faces of the self same basic Merge operation (E and I-merge), movement (aka I-Merge) must target c-commanding positions (due to Extension), and the products of I-Merge necessarily produce copies (due to Inclusiveness and hence producing structures supporting operator-variable structures and allowing for reconstruction effects). So, from a simple idea concerning the recursive mechanism, Chomsky derives a bunch of plausible properties of Gs and UG that GGers have proposed over the last 50 years of research.

However, the generalizations deduced (cyclicity, c-command, copies etc.) are not perfect (e.g. tucking-in is not strictly speaking cyclic in the standard usage, there are many cases in which reconstruction is impossible, movement is not the only operation for which c-command is relevant). Does that mean that the Chomsky’s unification of these properties in terms of Merge is a bad one? Not necessarily. Conceptually it is an achievement for it shows how to link certain salient (stylized) features of Gs together. Empirically, it is a step forward for it links properties that have non-negligible empirical backing and that are plausibly descriptive of our FL. Is it “true”? Well, that depends on how we eventually handle the (apparent) problems for the (lower level) principles that it has unified. Should these prove to be false, then this unification will not be what we ultimately want. However, and this is important, Chomsky’s unification provides a strong (explanatory) incentive for going back and reanalyzing the (empirical) “problems” for the lower level principles, and it provides a nice example of the kind of theory we want. We really do want to have our cake and eat it too and this is what the dialectic between empirical “coverage” and theoretical “explanation” aims to provide. The problem is that for this dialectic to gain a foothold we need to appreciate both sides of the going-and-froing. We need to concretely understand the tension between explanatory force and empirical coverage and understand that the right theory needs both. Right now, IMO, our attitudes over-prize (apparent) empirical coverage. We very seldom count (or even address) the cost of lost explanation when we evaluate our proposals.

This is not a new complaint, at least from me. I make it again because in my experience GGers have a low tolerance for theoretical ambition. I suspect that this is so for several reasons. First, we tend to confuse formal work with theoretical work and this muddies our sensitivity to the explanatory oomph of different approaches. Second, linguistics is a data rich field and so supporting theory means tolerating some empirical slack at least for a while. But, last, I think that we don’t actually spend enough time teaching and touting the explanatory virtues of our best accounts. We seldom go back and ask what we have lost or try to theoretically motivate the new principles we adopt to “capture” the data. Indeed, the whole idea that data is something that needs capturing (rather than explaining) is, to my mind, quite odd.

Does this mean that theory does not need empirical support? Nope. Theories need to be justified by facts. But, facts also need to be justified by theories. One of the original hopes of the minimalist program was that it would sensitize us to what a good explanation was. It would make us aware that our “explanations” (and these are scare quotes) are often as complex as the data they address. And this is not good. IMO, this appreciation is less vivid today than it was in the earliest days of the minimalist program. And part of the problem is un-interest in theory and a misplaced belief that lots of data signifies empirical progress. In this regard, GG work has been disimproving.



[1] “GB” is in quotes because I do not mean to invidiously distinguish between GB proper and its many theoretical twins (many of them identical IMO for most of the questions I am interested in). These include LFG, RG, GPSG, HPSG a.o. From where I sit, most of these theories are intertranslatable and make effectively the same distinctions in the same theoretical places. They are more notationally than notionally distinct.

Tuesday, February 11, 2014

Plato, Darwin, P&P and variation

Alex C (in the comment section here (Feb. 1)) makes a point that I’ve encountered before that I would like to comment on. He notes that Chomsky has stopped worrying about Plato’s Problem (PP) (as has much of “theoretical” linguistics as I noted in the previous post) and suggests (maybe this is too much to attribute to him, if so, sorry Alex) that this is due to Darwin’s Problems (DP) occupying center stage at present. I don’t want to argue with this factual claim, for I believe that there’s lots of truth to it (though IMO, as readers of the last several posts have no doubt gathered, theory of any kind is largely absent from current research). What I want to observe is that (1) there is a tension between PP and DP and (2) that resolving it opens an important place for theoretical speculation. IMO, one of the more interesting facets of current theoretical work is that it proposes a way of resolving this tension in an empirically interesting way. This is what I want to talk about.

First the tension: PP is the observation that the PLD the child uses in developing its G is impoverished in various ways when one compares it to the properties of Gs that children attain. PP, then, is another name for the Poverty of Stimulus Problem (POS).  Generative Grammarians have proposed to “solve” this problem by packing FL with principles of UG, many of which are very language specific (LS), at least if GB is taken as a guide to the content of FL.  By LS, I mean that the principles advert to very linguisticky objects (e.g. Subjects, tensed clauses, governors, case assigners, barriers, islands, c-command, etc) and very linguisticky operations (agreement, movement, binding, case assignment, etc.).  The idea has been that making UG rich enough and endowing it with LS innate structure will allow our theories of FL to attain explanatory adequacy, i.e. to explain how, say, Gs obey islands despite the absence of good and bad data relevant to fixing them present in the PLD. 

By now, all of this is pretty standard stuff (which is not to say that everyone buys into the scheme (Alex?)), and, for the most part, I am a big fan of POS arguments of this kind and their attendant conclusions. However, even given this, the theoretical problem that PP poses has hardly been solved. What we do have (again assuming that the POS arguments are well founded (which I do believe)) is a list of (plausibly) invariant(ish) properties of Gs and an explanation for why these can emerge in Gs in the absence of the relevant data in the PLD required to fix them. Thus, why do movement rules in a given G resist extraction from islands? Because something like the Subjacency/Barriers theory is part of every Language Acquisition Device’s (LAD) FL, that’s why.

However, even given this, what we still don’t have is an adequate account of how the variant properties of Gs emerge when planted in a particular PLD environment. Why is there V to T in French but not in English? Why do we have inverse control in Tsez but not Polish? Why wh-in-situ in Chinese but multiple wh to C in Bulgarian. The answer GB provided (and so far as I can tell, the answer still) is that FL contains parameters that can be set in different ways on the basis of PLD and the various Gs we have are the result of differential parameter setting. This is the story, but we have known for quite a while that this is less a solution to the question of how Gs emerge in all their variety than it is an explanation schema for a solution. P&P models, in other words, are not so much well worked out theories than they are part of a general recipe for a theory that were we able to cook it, would produce just the kind of FL that could provide a satisfying answer to the question of how Gs can vary so much. Moreover, as many have observed (Dresher and Janet Fodor are two notable examples, see below) there are serious problems with successfully fleshing out a P&P model.

Here are two: (i) the hope that many variant properties of Gs would hinge on fixing a small number of parameters seems increasingly empirically uncertain. Cederic Boeckx and Fritz Newmeyer have been arguing this for a while, and while their claims are debated (and by very intelligent people so, at least for a non-expert like me, the dust is still too unsettled to reach firm conclusions), it seems pretty clear that the empirical merits of earlier proposed parameterizations are less obvious than we took them to be. Indeed, there appears to some skepticism about whether there are any macro-parameters (in Baker’s sense[1]) and many of the micro-parametric proposals seem to end up restating what we observe in the data: that languages can differ. What made early macro-parameter theories interesting is the idea that differences among Gs come in largish clumps. The relation between a given parameter setting and the attested surface differences was understood as one to many. If, however, it turns out that every parameter correlates with just a single difference then the value of a parametric approach becomes quite unclear, at least so far as acquisition considerations are concerned. Why? Because it implies that surface differences are just due to differing PLD, not to the different options inherent in the structure of FL. In other words, if we end up with one parameter per surface difference then variation among Gs will not be as much of a window into the structure of FL as we thought it could be.

Here’s another problem: (ii) the likely parameters are not independent. Dresher (and friends) has demonstrated this for stress systems and Fodor (and friends) has provided analogous results for syntax.  The problem with a theory where parameters are not independent is that they make it very hard to see how acquisition could be incremental. If it turns out that the value of any parameter is conditional on the value of every other parameter (or very many others) then it would seem that we are stuck with a model in which all parameters must be set at once (i.e. instantaneous learning). This is not good! To evade this problem, we need some way of imposing independence on the parameters so that they can be set piecemeal without fear of having to re-set them later on. Both Dresher and Fodor have proposed ways of solving this independence problem (both elaborate a richer learning theory for parameter values to accommodate this problem). But, I think that it is fair to say that we are still a long way from a working solution. Moreover, the solutions provided all involve greatly enriching FL in a very LS way. This is where PP runs into DP. So let’s return to the aforementioned tension between PP and DP.

One way to solve PP is to enrich FL. The problem is that the richer and more linguistically parochial FL is, the harder it becomes to understand how it might have evolved. In other words, our standard GB tack in solving PP (LS enrichment of FL) appears to make answering DP harder. Note I say ‘appears.’ There are really two problems, and they are not equally acute. Let me explain.

As noted above, we have two things that a rich FL has been used to explain; (a) invariances characteristic of all Gs and (b) the attested variation among Gs. In a P&P model, the first ‘P’ handles (a) and the second (b). I believe that we have seen glimmers of how to resolve the tension between PP’s demands on FL versus DP’s as regards the principles part of P&P. Where things have become far more obscure (and even this might be too kind) involves the second parametric P. Here’s what I mean.

As I’ve argued in the past, one important minimalist project has been to do for the principles of GB what Chomsky did for islands and movement via the theory of subjacency in On Wh Movement (OWM). What Chomsky did in this paper is theoretically unify the disparate island effects by unifying all non-local (A’) dependency constructions by proposing that they have a common movement core (viz. move WH) subject to locality restrictions characterized by Bounding Theory (BT). This was terrifically inventive theory and aside from rationalizing/unifying Ross’s very disparate Island Effects, the combination of Move WH + BT predicted that all long movement would have to be successive cyclic (and even predicted a few more islands, e.g. subject islands and Wh-islands).[2]

But to get back to PP and DP, one way of regarding MP work over the last 20 years is as an attempt to do for GB modules what Chomsky did for Ross’s Islands. I’ve suggested this many times before but what I want to emphasize here is that this MP project is perfectly in harmony with the PP observation that we want to explain many of the invariances witnessed across Gs in terms of an innately structured FL. Here there is no real tension if this kind of unification can be realized. Why not? Because if successful we retain the GB generalizations. Just as Move WH + BT retain Ross’s generalizations, a successful unification within MP will retain GB’s (more or less) and so we can continue to tell the very same story about why Gs display the invariances attested as we did before. Thus, wrt this POS problem, there is a way to harmonize DP concerns with PP concerns. Of course, this does not mean that we will successfully manage to unify the GB modules in a Move WH + BT way, but we understand what a successful solution would look like and, IMO, we have every reason to be hopeful, though this is not the place to defend this view.

So, the principles part of P&P is, we might say, DP compatible (little joke here for the cognoscenti). The problem lies with the second P. FL on GB was understood to provide not only the principles of invariance but also to specify all the possible ways that Gs could differ. The parameters in GB were part of FL! And it is hard to see how to square this with DP given the terrific linguistic specificity of these parameters. The MP conceit has been to try and understand what Gs do in terms of one (perhaps)[3] linguistically specific operation (Merge) interacting with many general cognitive/computational operations/principles.  In other words, the aim has been to reduce the parochialism of the GB version of FL. The problem with the GB conception of parameters is that it is hard to see how to recast them in similarly general terms. All the parameters exploit notions that seem very very linguo-centric. This is especially true of micro parameters, but it is even true of macro ones. So, theoretically, parameters present a real problem for DP, and this is why the problems alluded to earlier have been taken by some (e.g. me) to suggest that maybe FL has little to say about G-variation. Moreover, it might explain why it is that, with DP becoming prominent, some of the interest in PP has seemed to wane. It is due to a dawning realization that maybe the structure of FL (our theory of UG) has little to say directly about grammatical variation and typology. Taken together PP and DP can usefully constrain our theories of FL, but mainly in licensing certain inferences about what kinds of invariances we will likely discover (indeed have discovered). However, when it comes to understanding variation, if parameters cannot be bleached of their LSity (and right now, this looks to me like a very rough road), it looks to me like they will never be made to fit with the leading ideas of MP, which are in turn driven by DP. 

So, Alex C was onto something important IMO. Linguists tend to believe that understanding variation is key to understanding FL. This is taken as virtually an article of faith. However, I am no longer so sure that this is a well founded presumption. DP provides us with some reasons to doubt that the range of variation reflects intrinsic properties of FL. If that is correct, then variation per se may me of little interest for those interested in liming the basic architecture of FL. Studying various Gs will, of course, remain a useful tool for in getting the details of the invariant principles and operations right. But, unlike earlier GB P&P models, there is at least an argument to be made (and one that I personally find compelling) that the range of G-variation has nothing whatsoever to do with the structure of FL and so will shed no light on two of the fundamental questions in Generative Grammar: what’s the structure of FL and why?[4]





[1] Though Baker, a really smart guy, thinks that there are so please don’t take me as endorsing the view that there aren’t any. I just don’t know. This is just my impression from linguist in the street interviews.
[2] The confirmation of this prediction was one of the great successes of generative grammar and the papers by, e.g. Kayne and Pollock, McCloskey, Chung, Torrego, and many others are still worth reading and re-reading. It is worth noting that the Move WH + BT story was largely driven by theoretical considerations, as Chomsky makes clear in OWM. The gratifying part is that the theory proved to be so empirically fecund.
[3] Note the ‘perhaps.’ If even merge is in the current parlance “third factor” then there is nothing taken to be linguistically special about FL.
[4] Note that this quite a bit of room for “learning” theory. For if the range of variation is not built into FL then why we see the variation we do must be due to how we acquire Gs given FL/UG.  The latter will still be important (indeed critical) in that any larning theory will have to incorporate the isolated invariances. However, a large part of the range of variation will fall outside the purview of FL. I discuss this somewhat in the last chapter of A theory if syntax for any of you with a prurient interest in such matters. See, in particular, the suggestion that we drop the switch analogy in favor of a more geometrical one.

Sunday, January 27, 2013

Joining the Fun; A Ramble on Parameters


There a very interesting pair of posts and a long thread of insightful comments relating to parameters, both the empirical support for such as well as their suitability given current theoretical commitments.  Cederic, commenting on Neil’s initial post and then adding a longer elaboration, makes the point that nobody seems committed to parameters in the classical sense anymore. Avery and Alex C comment that whatever the empirical shortcomings of parameteric accounts, something is always better than nothing so they reasonably ask what do we replace it with. Alex D rightly points out that the success of parametric accounts is logically independent of the POS and claims about linguistic nativism. In this post, I want to reconstruct the history of how parameter theory arose so as to consider where we ought to go from here. The thoughts ramble on a bit, because I have been trying to figure this out for myself.  Apologies ahead of time.

In the beginning there was the evaluation metric (EM), and Chomsky looked on his work and saw that it was deficient.  The idea in Aspects was that there is a measure of grammatical complexity built into FL and that children in acquiring their I-language (an anachronism here) chose the simplest one compatible with the PLD (viz. the linguistic data available to and used by the child). EM effectively ordered grammars according to their complexity. The idea riffed on ideas of minimal description length around at the time but with the important addition that the aim of a grammatical theory with aspirations of explanatory adequacy was to find the correct UG for specifying the meta-language relevant to determining the correct notion of “description” and “length” in minimal description length. The problem was finding the right things to count when evaluating grammars. At any rate, on this conception, the theory of acquisition involved finding the best overall G compatible with PLD as specified by EM.  Chomsky concluded that this conception, though logically coherent, was not feasible as a learning theory, largely because it looked to be computationally intractable.  Nobody had (nor I believe, has) a good tractable idea of how to compare grammars overall so as to have a complete ordering. Chomsky in LSLT developed some pair-wise metrics for the local comparison of alternative rules, but this is a long way from having a total ordering of the alternative Gs that is required to make EM accounts feasible.  Chomsky’s remedy for this problem: divorce language acquisition from the evaluation of overall grammar formats.

The developments of the Extended Standard Theory, which culminated in GB theories, allows for an alternative conception of acquisition, one that divorces it from measuring overall grammar complexity.  How so? Well, first we eliminated the idea that Gs were compendia of constructions specific rules. And second, we proposed that UG consists of biologically provided schemata (hence part of UG and hence not in need of acquisition) that specify the overall shape of a particular G. On this view, acquisition consists in filling in values for the schematic variables.  Filling in values of UG specified variables is a different task from figuring out the overall shape of the grammar and, on the surface at least, a far more tractable task. The number of parameters being finite already distinguished this from earlier conceptions. In the earlier EM view of things there was no reason to think that the space of grammatical possibilities was finite. Now, as Chomsky emphasized, within a parameter setting model, the space of alternatives, though perhaps very large, was still finite and hence the computational problem was different in kind from the one lightly limned in Aspects. So, divorcing the question of grammatical formats (via the elimination of rules or their reduction to a bare minimum form like ‘move alpha’) from the question of acquisition allowed for what looked like a feasible solution to Plato’s Problem. In place of Gs being sets of constructions specific rules with EMs measuring their overall collective fitness, we had the idea that Gs were vectors of UG specified variables with two possible values (and hence “at most” 2n possible grammars, a finite number of options). Finding the values was divorced from evaluating sets of rules and this looked feasible.

Note that this is largely a conceptual argument. There is a reasonable hunch but no “proof.” I mention this because other conceptual considerations (we will get to them) can serve to challenge the conclusion and make parameter theories less appealing.

In addition to these conceptual considerations, the comparative grammar research in the 70s, 80s, and 90s provided wow-inducing empirical confirmation of parameter based conceptions. It is hard for current (youngish) practitioners of the grammatical dark arts to appreciate how exciting early work on parameter setting models was. There were effectively three lines of empirical support.

1.     The comparative synchronic grammar research. For example:
a.     The S versus S’ parameter distinguishing Italian from English islands (Rizzi, Sportiche, Torrego)).
b.     The pro drop parameter (correlating, null subjects, inversion and long movement apparently violating the fixed subject/that-t condition (Rizzi, Brandi and Cordin)).
c.     The parametric discussions of anaphoric classes (local and long distance anaphors (Wexler, Borer), to name just three, all uncovered a huge amount of new linguistic data and argued for the fecundity of parametric thinking.
2.     Crain’s “continuity thesis,” which provided evidence that kids “mistakes” in acquiring their particular Gs all actually conform to actual adult Gs. This provided evidence that the space of G options was pretty circumscribed, as a parameter theory implies it is.
3.     The work on diachronic change by Kroch, Lightfoot, Roberts (and more formal work by Berwick and Niyogi) a,o., which indicated that large shifts in grammatical structure over time (e.g. SOV to SVO) could be analyzed as changes in a small number of simple parameter changes.

So, there was a good conceptual reason for moving to parameter models of UG and the move proved to be empirically very fecund. Why the current skepticism?  What’s changed?

To my mind, three changes occurred. As usual, I will start with the conceptual challenges and then proceed to the empirical ones.

The first one can be traced to work first by Dresher and Kaye, and then taken up and further developed with great gusto by Fodor (viz. Janet) and Sakas. This work shows that finite parameter setting can present tractability problems almost as difficult as the ones that Chomsky identified in his rejection of EM models.  What this work demonstrates is that given current envisioned parameters, parameter setting cannot be incremental. Why not? Because parameter values are not independent.  In other words, the value of one parameter in a particular G may depend crucially on that of another. Indeed, the value of any may depend on the value of each and this makes for an explosive combinatory problem. It also makes incremental acquisition mysterious; how do the parameter values get set if any bit of later PLD can completely overturn values previously set?

There have been ingenious solutions to this problem, my favorite being cue-based conceptions (developed by Dresher, Fodor, Lightfoot a.o.). These rely on the notion that there is some data in the PLD that unambiguously determines the value of a parameter. Once set on the basis of this data, the value need never change.  Triggers effectively impose independence on the parameter space. If this is correct, then it renders UG yet more linguistically specific; not only are the parameters very linguistically specific, but the apparatus required to fix these is very linguistically specific as well. Those that don’t like linguistically parochial UGs should really hate both parameter theories and this fix to them. Which brings us to the second conceptual shift: Minimalism.

The minimalist conceit is to eliminate the parochialism of FL and show that the linguistically specific structure of UG can be accounted for in more general cognitive/computational terms. This is motivated both on general methodological grounds (factoring out what is cognitively general from what is linguistically specific is good science) and as a first step to answering Darwin’s Problem, as we’ve discussed at length in other posts. FL internal parameters are a very big challenge to this project. Why? Because UG specified parameters encumber FL with very linguistically specific information (e.g. it’s hard to see how the pro drop parameter (if correct) could possibly be stated in non linguistically specific terms!).

This is what I meant earlier when I noted that conceptual reasons could challenge Chomsky’s earlier conceptual arguments.  Even if parameters made addressing Plato’s Problem more tractable, they may not be a very good solution to the feasibility problem if they severely compromise any approach to Darwin’s. This is what motivates Cederic’s concerns (and others, e.g. Terje Lohndal) I believe, and rightly so.  So, the conceptual landscape has changed and it is not surprising that parameter theories have become less appealing and so open to challenge.

Moreover, as Cederic also stresses, the theoretical landscape has changed as well. A legacy of the GB era that has survived into Minimalism is the agreement that Gs do not consist of construction based rules. Rather, there are very general operations (Merge) with very general constraints (e.g. Extension, Minimality) that allow for a small set of dependencies universally.  Much of this (not all, but much) can be reanalyzed in non linguistically specific terms (or so I believe). With this factored out, there are featural idiosyncracies located in demands made by specific lexical items, but this kind of idiosyncracy may be tolerable as it is segregated to the lexicon, a well known repository of eccentrics.[1]  At any rate, it is easy to see what would motivate a reconsideration of UG internal parameters.

The tractability problems related to parameter setting noted by Dresher-Fodor and company simply add to these motivations. 

That leaves us with the empirical arguments. These alone are what make parameter accounts worth endorsing, if they are well founded, and this is what is currently up for grabs and way beyond my pay grade. Cederic and Fritz Newmeyer (among others) have challenged the empirical validity of the key results. The most important discoveries amounted to the clumping of surface effects with the settings of single values, e.g. pro drop+subject inversion+no that-t effects together as a unit. Find one, you find them all.  However, this is what has been challenged. Is it really true that the groupings of phenomena under single parameter settings is correct?  Do these patterns coagulate as proposed? If not, and this I believe is Newmeyer’s point and strongly empahasized by Cedric, then it is not clear what parameters buy us.  Yes I-languages are different. So? Why think that this difference is due to different parameter settings? So, there is an empirical argument: are there data groupings of the kind earlier proposals advocated? Is the continuity thesis accurate and if so how does one explain this without parameters? These are the two big empirical questions and it is likely to be where the battle over parameters has been joined and, one hopes, will ultimately get resolved.

I’d like to epmpahsize that this is an empirical question.  If the data falls on the classical side then this is a problem for minimalists and exacerbates our task of addressing Darwin’s problem. So be it. Minimalism as I understand it has an empirical core and if it turns out that there is richer structure to UG than I would like, well tough cookies on me (and you if your sympathies tend in the same direction)!

Last point and I will end the rambling here. One nice feature of parameter models is the pretty metaphor it afforded for language acquisition as parameter setting. The switch box model is intuitive and easy to grasp. There is no equivalent for EM models and this is partly why nobody knew what to do with the damn thing.  EM never really got used to generate actual empirical research the way parameter setting models did, at least not in syntax. So can we envision a metaphor for non parameter setting models. I think we can. I offered one in A theory of syntax that I’d like to try and push it again here (I know that this is self aggrandizing, but tooting one’s own horn can be so much fun).  Here’s what I said there (chapter 7):

Assume for a moment that the idea of specified parameters is abandoned. What then?  One attractive property of the GB story was the picture that it came with.  The LAD was analogized to a machine with open switches.  Learning amounts to flipping the switches ‘on’ or ‘off’.  A specific grammar is then just a vector of these switches in one of the two positions.  Given this view there are at most 2P grammars (P=number of parameters).  There is, in short, a finite amount of possible variation among grammars.
            We can replace this picture of acquisition with another one.  Say that FL provides the basic operations and conditions on their application (e.g. like minimality).  The acquisition process can now be seen as a curve fitting exercise using these given operations.  There is no upper bound on the ways that languages might differ though there are still some things that grammars cannot do.  A possible analogy for this conception of grammar is the variety of geometrical figures that can be drawn using a straight edge and compass.  There is no upper bound on the number of possible different figures.  However, there are many figures that cannot be drawn (e.g. there will be no triangles with 20 degree angles).  Similarly, languages may contain arbitrarily many different kinds of rules depending on the PLD they are trying to fit.

So think of the basic operations and conditions as the analogues of the straight edge and compass and think of language acquisition as fitting the data using these tools. Add to this a few general rules for figure fitting: add a functional category if required, pronounce a bottom copy of a chain rather than a top copy, add an escape hatch to a phase head. These are general procedures that can allow the LAD to escape the strictures of the limited operations a minimalistically stripped down FL makes available.  The analogy is not perfect. But the picture might be helpful in challenging the intuitive availability of the switch box metaphor.

That’s it. This post has also been way too long. Kudos to Neil and Cedric and the various very articulate commenters for making this such a fruitful topic for thought, at least for me. 


[1] Though I won’t discuss this now, it seems to me that the Cartographic Project and its discovery of what amounts to a universal base for all Gs is not so easily dismissed. The best hope is to see these substantive universals explicated in semantic terms, not something I am currently optimistic will soon appear.