Comments

Showing posts with label P&P. Show all posts
Showing posts with label P&P. Show all posts

Thursday, October 2, 2014

The Aspects acquisition model

Some of this post is thinking out loud.  I am not as sure as I would like to be about certain things (e.g. how to understand feasibility for example, see below). This said, I thought I’d throw it up and see if comments etc. allow me to clarify my own thoughts.

In Aspects (chapter 1;30ff), Chomsky outlines an abstract version of an “acquisition model.”[1] I want to review some of its features here. I do this for two reasons. First, this model was later replaced with a principles and parameters (P&P) account and in order to see why this happened it’s useful to review the theory that P&P displaced. Second, it seems to me that the Aspects model is making a comeback, often in a Bayesian wrapper, and so reviewing the features of the Aspects model will help clarify what Bayesians are bringing to the acquisition party beyond what we already had in the Aspects model. BTW, in case I haven’t mentioned this before, chapter 1 of Aspects is a masterpiece. Everyone should read it (maybe once a year, around Passover where we commemorate our liberation from the tyranny of Empiricism). If you haven’t done so yet, stop reading this post and do it! It’s much better than anything you are going to read below.

Chomsky presents an idealized acquisition model in §6 (31).  The model has 5 parts:

i.               an enumeration of the class s1, s2… of possible sentences
ii.              an enumeration of the class SD1, SD2… of possible structural descriptions
iii.            an enumeration of the class G1, G2,… of possible generative grammars
iv.            specifiction of the f such that SDf(i,j) is the structural description to sentence si by grammar Gj for arbitrary i, j
v.              specification of a function m such that m(i) is an integer associated with the grammar Gi as its value[2]

 (i-v) describe the kinds of capacities a language acquisition device (LAD) must have to use primary linguistic data (PLD) to acquire a G. It must (i) have a way of representing the input signals, (ii) a way of assigning structures to these signals, (iii) a way of restricting the class of possible structures available to languages, (iv) a way of figuring out what each hypothetical G implies for each sentence (i.e. a input/structure pair) and (v) a method for selecting one of the very very many (maybe infinitely many) hypothesis allowed by (iii) that are compatible with the given PLD.  So, we need a way of representing the input, matching that input to a Gish representation of that input and a way of choosing the “right” match (the correct G) from the many logically possible G-matches (i.e. a way of “evaluating alternative proposed grammars”).

How would (i-v) account for language acquisition? A LAD with structure (i-v) could use PLD to search the space of Gs to find the one that generates that PLD.  The PLD, given (i, ii) is a pairing of inputs with (partial) SDs. (iii, iv) allows these SDs to be related to particular Gs. As Chomsky says it (32):

The device must search through the set of possible hypotheses G1, G2,…, which are available to it by virtue of condition (iii), and must select grammars that are compatible with the primary linguistic data, represented in terms if (i) and (ii). It is possible to test compatibility by virtue of the fact that the device meets condition (iv).

The last step is to select one of these “potential grammars” using the evaluation measure provided by (v). Thus, if a LAD has these five components, the LAD has the capacity to build a “theory of the language of which the primary linguistic data are a sample” (32).

As Chomsky notes, (i-v) packs a lot of innate structure into the LAD. And, interestingly, what he proposes in Aspects matches pretty closely how our thoroughly modern Bayesians would describe the language acquisition problem: A space of possible Gs, a way of matching empirical input to structures that the Gs in the space generate, and a way of choosing the right G among the available Gs given the analyzed input and the structure of the space of Gs.[3] The only thing “missing” from Chomsky’s proposal is Bayes rule, but I have no doubt that were it useful to add, Chomsky would have had no problem adding it. Bayes rule would be part of (v), the rule specifying how to choose among the possible Gs given PLD. It would say: “Choose the G with the highest posterior probability.”[4] The relevant question is how much this adds? I will return to this question anon.

Chomsky describes theories that meet conditions (i-iv) as descriptive and those adding (v) as well, as being explanatory. Chomsky further notes that gaining explanatory power is very hard, the reason being that there are potentially way too many Gs compatible with given PLD.  If so, then choosing the right G (the needle) given the PLD (a very large haystack) is not a trivial task.  In fact, in Chomsky’s view (35):

… the real problem is almost always to restrict the range of possible hypotheses [i.e. candidate Gs, NH] by adding additional structure to the notion “generative grammar.” For the construction of a reasonable acquisition model, it is necessary to reduce the class of attainable grammars compatible with given primary linguistic data to the point where selection among them can be made by a formal evaluation measure. This requires a precise and narrow delimitation of the notion “generative grammar”- a restrictive and rich hypothesis concerning the universal properties that determine the form of language…[T]he major endeavor of the linguist must be to enrich the theory of linguistic form by formulating more specific constraints and conditions on the notion “generative grammar.”

So, the main explanatory problem, as Chomsky sees it, is to so circumscribe (and articulate) the grammatical hypothesis space such that for any given PLD, only a very few candidate Gs are possible acquisition targets.[5] In other words, the name of the explanatory game is to structure the hypothesis space (either by delimiting the options or biasing the search (e.g. via strong priors)) so that, for any given PLD, very few candidates are being simultaneously evaluated.  If this is correct, the focus of theoretical investigation is the structure of this space, which as Chomsky further argues, amounts to finding the universals that suitably structure the space of Gs. Indeed, Chomsky effectively identifies the task of achieving explanatory adequacy with the “attempt to discover linguistic universals” (36), principles of G that will deliver a space of possible Gs that for any given PLD a very small number of candidate Gs need be considered.

I have noted that the Aspects model shares many of the features that a contemporary Bayesian model of acquisition would also assume. Like the Aspects model, a Bayesian one would specify a structured hypothesis space that ordered the available alternatives in some way (e.g. via some kind of simplicity measure?). It would also add a rule (viz. Bayes rule) for navigating this space (i.e. by updating values of Gs) given input data and a decision rule that roughly enjoins that the one choose (at some appropriate time) the highest valued alternative.  Here’s my question: what does Bayes add to Aspects?

In one respect, I believe that it reinforces Chomsky’s conclusion: that we really really need a hypothesis space that focuses LAD’s attention on a very small number of candidates. Why?

The answer, in two words, is computational tractability. Doing Bayes proud is computationally expensive. A key feature of Bayesian models is that with each new input of data the whole space of alternatives (i.e. all potential Gs) is updated. Thus, if the there are, say, 100 possible grammars, then for each datum D all 100 are evaluated with respect to D (i.e. Bayes computes a posterior for each G given D). And this is known to be computationally so expensive as to not be feasible if the space of alternatives is moderately large.[6] Here, for example, is what O’Reilly, Jbabdi and Behrens (OJB) say (see note 5):

…it is well known that “adding parameters to a model (more dimensions to the model) increases the size of the state space, and the computing power required to represent and update it, exponentially (1171).

As OJB further notes, the computational problems arise even when there are only “a handful of dimensions of state spaces” (1175, my emphasis, NH).

This would not be particularly problematic if only a small number of relevant alternatives were the focus of Bayesian attention, as would be the case given Chomsky’s conception of the problem, and that’s why I say that the Aspects formulation of what’s needed seems to fit well with Bayesian concerns. Or, to put this another way: if you want to be Bayesian then you’d better hope that something like Chomsky’s position is correct and that we can find a way of using universals to develop an evaluation measure that serves to severely restrict the relevant Gs under consideration for any given PLD.

There is one way, however, in which Chomsky’s guess in the Aspects model and contemporary Bayesians seem to part ways, or at least seem to emphasize different parts of the research problem (I say ‘seem’ because what I note does not follow from Bayes like assumptions. Rather it is characteristic of what I have read (and, recall, I have not read tons in this area, just some of the “hot” papers by Tenenbaum and company). Chomsky says the following (36-7):

It is logically possible that the data might be sufficiently rich and the class of potential grammars sufficiently limited so that no more than a single permitted grammar will be compatible with the available data at the moment of successful language acquisition…In this case, no evaluation procedure will be necessary as part of linguistic theory – that is, as an innate property of an organism or a device capable of language acquisition. It is rather difficult to imagine how in detail this logical possibility might be realized, and all concrete attempts to formulate an empirically adequate linguistic theory certainly leave ample room for mutually inconsistent grammars, all compatible with the primary linguistic data of any conceivable sort. All such theories therefore require supplementation by an evaluation measure if language acquisition is to be accounted for and the selection of specific grammars is to be justified: and I shall continue to assume tentatively…that this is an empirical fact about the innate human faculté de langage and consequently about general linguistic theory as well.

In other words, the HARD acquisition problem in Chomsky’s view resides in figuring out the detailed properties of the evaluation metric.  Once we have this, the other details will fall into place. So, the emphasis in Aspects strongly suggests that serious work on the acquisition problem will focus on elaborating the properties of this innate metric. And this means working on developing “a restrictive and rich hypothesis concerning the universal properties that determine the form of language.”

Discussions of this sort are largely missing from Bayesian proposals. It’s not that they are incompatible with these and it’s not even that nods in this direction are not frequently made (see here). Rather most of the effort seems placed on Bayes Rule, which, from the outside (where I sit) looks a lot like bookkeeping. The rule is fine, but its efficacy rests on a presupposed solution to the hard problem. And this looks as if Bayesians worry more about how to navigate the space (on the updating procedure) given its structure rather than on what the space looks like (it’s algebraic structure and the priors on it).[7] So, though Bayes and Chomsky in Aspects look completely compatible, what they see as the central problems to be solved look (or seem to look) entirely different.[8]

What happened to the Aspects model? In the late 70s and early 80s, Chomsky came to replace this “acquisition model” with a P&P model. Why did he do this and how are they different? Let’s consider these questions in turn.

Chomsky came to believe that the Aspects approach was not feasible. In other words, he despaired of finding a formal simplicity metric that would so order the space of grammars as required.[9]  Not that he didn’t try. Chomsky discusses various attempts in §7, including ordering grammars in accord with the number of symbols they use to express their rules (42).[10] However, it proved to be very hard (indeed impossible) to come up with a general formal way of so ordering Gs.

So, in place of general formal evaluation metrics, Chomsky proposed P&P systems where the class of available grammars are finitely specified by substantive 2 valued parameters. P&P parameters are not formally interesting. In fact, there have been no general theories (not even failed ones) of what a possible parameter is (a point made by Gert Webelhuth in this thesis, and subsequently).[11] In this sense, P&P approaches to acquisition are less theoretically ambitious than earlier theories based on evaluation measures. In effect, Chomsky gave up on the Aspects model because it proved to be hard to give a general definition of “generative grammar” that served to order the infinite variety of Gs according to some general (symbol counting) metric. So, in place of this, he proposed that all Gs have the same general formal structure and only differ in a finite number of empirically pre-specified ways. On this revised picture, Gs as a whole are no longer formally simpler than one another. They are just parametrically different. Thus, in place of an overall simplicity measure, P&P theories concentrate on the markedness values of the specific parameter values; some values being more highly valued than others.

Let me elaborate a little. The P&P “vision” as illustrated by GB (as one example) is that Gs all come with a pre-specified set of rules (binding, case marking, control, movement etc.). Languages differ, but not in the complexity of these rules. Take movement rules as an example. They are all of the form ‘move alpha’ with the value of alpha varying across languages. This very simple rule has no structural description and no structural change, unlike earlier rules like passive, or raising or relativization. In fact, the GB conceit was that rules like Passive did not really exist! Constructions gave way to congeries of simple rules with no interesting formal structure.  As such there was little for an evaluation measure to do.[12] There remains a big role for markedness theory (which parameter values are preferred over others (i.e. priors in Bayes speak), but these do not seem to have much interesting formal structure.

Let me put this one more way: the role of the evaluation metric in Aspects was to formally order the relevant Gs by making some rule formats more unnatural than others. As rules become more and more simple, the utility of counting symbols to differentiate them becomes less and less useful. A P&P theory need not order Gs as it specifies the formal structure of every G: it’s a vector with certain values for the open parameters.  The values may be ranked, but formally, the rules in different Gs look pretty much the same.  The problem of acquisition moves from ordering the possible Gs by considering the formally distinct rule types, to finding the right values for pre-specified parameters.

As it turns out, even though P&P models are feasible in the required formal sense, they still have problems. In particular, setting parameters incrementally has proven to be a non-trivial task (as people like Dresher and Fodor & Sakas have shown) largely because the parameters proposed are not independent of one another. However, this is not the place to rehearse this point. What is of interest here is why evaluation metrics gave way to P&P models, namely that it proved to be impossible to find general evaluation measures to order the set of possible Gs and hence impossible to specify (v) above and thus attain explanatory adequacy.[13]

Let me end here for now (I really want to return to these issues later on).  The Aspects model outlines a theory of acquisition in which a formal ordering of Gs is a central feature. With such a theory, the space of possible Gs can be infinite, acquisition amounting to going up the simplicity ladder looking for the simplest G compatible with the PLD.  The P&P model largely abandoned this vision, construing acquisition instead as filling a finite number of fixed parameters (some values being “better” than others (i.e. unmarked)). The ordering of all possible Gs gave way to a pre-specification of the formal structure of all Gs.  Both stories are compatible with Bayesian approaches. The problem is not their compatibility, but what going Bayes adds. It’s my impression that Bayesians as a practical matter slight the concerns that both Aspects style models and P&P models concentrate on. This is not a matter of principle, for any Bayesian story needs what Chomsky has emphasized is both central and required.  What is less clear, at least to me, is what we really learn from models that concentrate more on Bayes rule than the structures that the rule is updating. Enlighten me.



[1] It’s wroth emphasizing that what is offered here is not an actual learning theory, but an idealized one. See his note 19 and 22 for further discussion.
[2] Chomsky suggests the convention that lower valued Gs are associated with higher numbers.
[3] One thing I’ve noticed though is that many Bayesians seem reluctant to conclude that this information about the hypothesis space and the decision rule are innately specified. I have never understood this (so maybe if someone out there thinks this they might drop a comment explaining it).  It always seemed to me that were they not part of the LAD then we had not acquisition explanation. At any rate, Chomsky did take (i-v) to specify innate features of the LAD that were necessary for acquisition.
[4] With the choice being final after some period of time t.
[5] See especially note 22 where Chomsky says:
What is required of a significant linguistic theory…is that given primary linguistic data D, the class of grammars compatible with D be sufficiently scattered, in terms of value, so that the intersection of the class of grammars compatible with D and the class of grammars which are highly valued be reasonably small. Only then can language learning actually take place.
[6] See OJB discussed here. It is worth noting that many Bayesians take Bayesian updating over the full parameter space to be a central characteristic of the Bayes perspective.  Here again is OJB:
It is a central characteristic of fully Bayesian models that they represent the full state space (i.e. the full joint probability distribution across all parameters) (1171).
It is worth noting, perhaps, that a good part of what makes Bayesian modeling “rational” (aka: “optimal,” and it’s key purported virtue) is that it considers all the consequences of all the evidence. One can truncate this so that only some of the consequences of only some of the evidence is relevant, but then it is less clear what makes the evaluations “rational/optimal.”  Not that there aren’t attempts to truncate the computation citing resource constraints and evaluating optimality wrt these constraints. However, this has the tendency of being a mug’s game as it is always possible to add just enough to get the result that you want, whatever these happen to be.  See Glymour here. However, this is not the place to go into these concerns. Hoepfully, I can return to them sometime later.
[7] Indeed, many of the papers I’ve seen try to abstract from the contributions of the priors. Why? Because sufficient data washes out any priors (so long as they are not set to 1, a standard assumption in the modeling literature precisely to allow the contribution of priors to be effectively ignored).  So, the papers I’ve seen say little of a general nature about the hypothesis space and little about the possible priors (viz. what I have been calling the hard problem).
[8] See for example the Perfors et. al. paper (here). The space of options is pretty trivial (5 possible grammars (three regular and two PCFGs) and it is hand coded in. It is not hard to imagine a more realistic problem: say including all possible PCFGs. Then the choice of the right one becomes a lot more challenging. In other words, seen as an acquisition model, this one is very much a toy system.
[9] Chomsky emphasizes that the notion of simplicity is proprietary to Gs, it is not some general notion of parsimony (see p. 38).  It would be interesting to consider how this fits with current Minimalist invocations of simplicity, but I won’t do so, or at least not now, not here.
[10] This model was more fully developed in Sound Patterns and earlier in The morphophonemics of modern Hebrew. It also plays a role in Syntactic Structures and the arguments for a transformational approach to the auxiliary system. Lasnik (here) has some discussion of this. I hope to write something up on this in the near future (I hope!). 
[11] Which precisely what makes parameters internal to FL minimalistically challenging.
[12] This is not quite right: a G that has alpha = any category is more highly valued than one that limits alpha’s reach to, e.g. just DPs.  Similary for things like the head parameter. Simplicity is then a matter of specifying more or less generally the domain of the rule. Context specifications, which played a large part in the earlier theory, however, are no longer relevant given such a slimed down rule structure. So the move to simple rules does not entirely eliminate considerations of formal simplicity, but it truncates it quite a bit. 
[13] The Dresher-Fodor/Sakas problem is a problem that arises on the (very reasonable and almost certainly correct) assumption that Gs are acquired incrementally. The problem is that unless the parameter values are independent, no parameter setting is fixed unless all the data is in. P&P models abstract away from real time learning. So too with Aspects style models. They were not intended as models of real time learning. Halle and Chomsky note this p. 331 of Sound Patterns where they describe acquisition as an “instantaneous process.” When Chomsky concludes that evaluation measures are not feasible, he abstracts away from the incrementality issues that Dresher-Fodor/Sakas zero in on.

Tuesday, February 11, 2014

Plato, Darwin, P&P and variation

Alex C (in the comment section here (Feb. 1)) makes a point that I’ve encountered before that I would like to comment on. He notes that Chomsky has stopped worrying about Plato’s Problem (PP) (as has much of “theoretical” linguistics as I noted in the previous post) and suggests (maybe this is too much to attribute to him, if so, sorry Alex) that this is due to Darwin’s Problems (DP) occupying center stage at present. I don’t want to argue with this factual claim, for I believe that there’s lots of truth to it (though IMO, as readers of the last several posts have no doubt gathered, theory of any kind is largely absent from current research). What I want to observe is that (1) there is a tension between PP and DP and (2) that resolving it opens an important place for theoretical speculation. IMO, one of the more interesting facets of current theoretical work is that it proposes a way of resolving this tension in an empirically interesting way. This is what I want to talk about.

First the tension: PP is the observation that the PLD the child uses in developing its G is impoverished in various ways when one compares it to the properties of Gs that children attain. PP, then, is another name for the Poverty of Stimulus Problem (POS).  Generative Grammarians have proposed to “solve” this problem by packing FL with principles of UG, many of which are very language specific (LS), at least if GB is taken as a guide to the content of FL.  By LS, I mean that the principles advert to very linguisticky objects (e.g. Subjects, tensed clauses, governors, case assigners, barriers, islands, c-command, etc) and very linguisticky operations (agreement, movement, binding, case assignment, etc.).  The idea has been that making UG rich enough and endowing it with LS innate structure will allow our theories of FL to attain explanatory adequacy, i.e. to explain how, say, Gs obey islands despite the absence of good and bad data relevant to fixing them present in the PLD. 

By now, all of this is pretty standard stuff (which is not to say that everyone buys into the scheme (Alex?)), and, for the most part, I am a big fan of POS arguments of this kind and their attendant conclusions. However, even given this, the theoretical problem that PP poses has hardly been solved. What we do have (again assuming that the POS arguments are well founded (which I do believe)) is a list of (plausibly) invariant(ish) properties of Gs and an explanation for why these can emerge in Gs in the absence of the relevant data in the PLD required to fix them. Thus, why do movement rules in a given G resist extraction from islands? Because something like the Subjacency/Barriers theory is part of every Language Acquisition Device’s (LAD) FL, that’s why.

However, even given this, what we still don’t have is an adequate account of how the variant properties of Gs emerge when planted in a particular PLD environment. Why is there V to T in French but not in English? Why do we have inverse control in Tsez but not Polish? Why wh-in-situ in Chinese but multiple wh to C in Bulgarian. The answer GB provided (and so far as I can tell, the answer still) is that FL contains parameters that can be set in different ways on the basis of PLD and the various Gs we have are the result of differential parameter setting. This is the story, but we have known for quite a while that this is less a solution to the question of how Gs emerge in all their variety than it is an explanation schema for a solution. P&P models, in other words, are not so much well worked out theories than they are part of a general recipe for a theory that were we able to cook it, would produce just the kind of FL that could provide a satisfying answer to the question of how Gs can vary so much. Moreover, as many have observed (Dresher and Janet Fodor are two notable examples, see below) there are serious problems with successfully fleshing out a P&P model.

Here are two: (i) the hope that many variant properties of Gs would hinge on fixing a small number of parameters seems increasingly empirically uncertain. Cederic Boeckx and Fritz Newmeyer have been arguing this for a while, and while their claims are debated (and by very intelligent people so, at least for a non-expert like me, the dust is still too unsettled to reach firm conclusions), it seems pretty clear that the empirical merits of earlier proposed parameterizations are less obvious than we took them to be. Indeed, there appears to some skepticism about whether there are any macro-parameters (in Baker’s sense[1]) and many of the micro-parametric proposals seem to end up restating what we observe in the data: that languages can differ. What made early macro-parameter theories interesting is the idea that differences among Gs come in largish clumps. The relation between a given parameter setting and the attested surface differences was understood as one to many. If, however, it turns out that every parameter correlates with just a single difference then the value of a parametric approach becomes quite unclear, at least so far as acquisition considerations are concerned. Why? Because it implies that surface differences are just due to differing PLD, not to the different options inherent in the structure of FL. In other words, if we end up with one parameter per surface difference then variation among Gs will not be as much of a window into the structure of FL as we thought it could be.

Here’s another problem: (ii) the likely parameters are not independent. Dresher (and friends) has demonstrated this for stress systems and Fodor (and friends) has provided analogous results for syntax.  The problem with a theory where parameters are not independent is that they make it very hard to see how acquisition could be incremental. If it turns out that the value of any parameter is conditional on the value of every other parameter (or very many others) then it would seem that we are stuck with a model in which all parameters must be set at once (i.e. instantaneous learning). This is not good! To evade this problem, we need some way of imposing independence on the parameters so that they can be set piecemeal without fear of having to re-set them later on. Both Dresher and Fodor have proposed ways of solving this independence problem (both elaborate a richer learning theory for parameter values to accommodate this problem). But, I think that it is fair to say that we are still a long way from a working solution. Moreover, the solutions provided all involve greatly enriching FL in a very LS way. This is where PP runs into DP. So let’s return to the aforementioned tension between PP and DP.

One way to solve PP is to enrich FL. The problem is that the richer and more linguistically parochial FL is, the harder it becomes to understand how it might have evolved. In other words, our standard GB tack in solving PP (LS enrichment of FL) appears to make answering DP harder. Note I say ‘appears.’ There are really two problems, and they are not equally acute. Let me explain.

As noted above, we have two things that a rich FL has been used to explain; (a) invariances characteristic of all Gs and (b) the attested variation among Gs. In a P&P model, the first ‘P’ handles (a) and the second (b). I believe that we have seen glimmers of how to resolve the tension between PP’s demands on FL versus DP’s as regards the principles part of P&P. Where things have become far more obscure (and even this might be too kind) involves the second parametric P. Here’s what I mean.

As I’ve argued in the past, one important minimalist project has been to do for the principles of GB what Chomsky did for islands and movement via the theory of subjacency in On Wh Movement (OWM). What Chomsky did in this paper is theoretically unify the disparate island effects by unifying all non-local (A’) dependency constructions by proposing that they have a common movement core (viz. move WH) subject to locality restrictions characterized by Bounding Theory (BT). This was terrifically inventive theory and aside from rationalizing/unifying Ross’s very disparate Island Effects, the combination of Move WH + BT predicted that all long movement would have to be successive cyclic (and even predicted a few more islands, e.g. subject islands and Wh-islands).[2]

But to get back to PP and DP, one way of regarding MP work over the last 20 years is as an attempt to do for GB modules what Chomsky did for Ross’s Islands. I’ve suggested this many times before but what I want to emphasize here is that this MP project is perfectly in harmony with the PP observation that we want to explain many of the invariances witnessed across Gs in terms of an innately structured FL. Here there is no real tension if this kind of unification can be realized. Why not? Because if successful we retain the GB generalizations. Just as Move WH + BT retain Ross’s generalizations, a successful unification within MP will retain GB’s (more or less) and so we can continue to tell the very same story about why Gs display the invariances attested as we did before. Thus, wrt this POS problem, there is a way to harmonize DP concerns with PP concerns. Of course, this does not mean that we will successfully manage to unify the GB modules in a Move WH + BT way, but we understand what a successful solution would look like and, IMO, we have every reason to be hopeful, though this is not the place to defend this view.

So, the principles part of P&P is, we might say, DP compatible (little joke here for the cognoscenti). The problem lies with the second P. FL on GB was understood to provide not only the principles of invariance but also to specify all the possible ways that Gs could differ. The parameters in GB were part of FL! And it is hard to see how to square this with DP given the terrific linguistic specificity of these parameters. The MP conceit has been to try and understand what Gs do in terms of one (perhaps)[3] linguistically specific operation (Merge) interacting with many general cognitive/computational operations/principles.  In other words, the aim has been to reduce the parochialism of the GB version of FL. The problem with the GB conception of parameters is that it is hard to see how to recast them in similarly general terms. All the parameters exploit notions that seem very very linguo-centric. This is especially true of micro parameters, but it is even true of macro ones. So, theoretically, parameters present a real problem for DP, and this is why the problems alluded to earlier have been taken by some (e.g. me) to suggest that maybe FL has little to say about G-variation. Moreover, it might explain why it is that, with DP becoming prominent, some of the interest in PP has seemed to wane. It is due to a dawning realization that maybe the structure of FL (our theory of UG) has little to say directly about grammatical variation and typology. Taken together PP and DP can usefully constrain our theories of FL, but mainly in licensing certain inferences about what kinds of invariances we will likely discover (indeed have discovered). However, when it comes to understanding variation, if parameters cannot be bleached of their LSity (and right now, this looks to me like a very rough road), it looks to me like they will never be made to fit with the leading ideas of MP, which are in turn driven by DP. 

So, Alex C was onto something important IMO. Linguists tend to believe that understanding variation is key to understanding FL. This is taken as virtually an article of faith. However, I am no longer so sure that this is a well founded presumption. DP provides us with some reasons to doubt that the range of variation reflects intrinsic properties of FL. If that is correct, then variation per se may me of little interest for those interested in liming the basic architecture of FL. Studying various Gs will, of course, remain a useful tool for in getting the details of the invariant principles and operations right. But, unlike earlier GB P&P models, there is at least an argument to be made (and one that I personally find compelling) that the range of G-variation has nothing whatsoever to do with the structure of FL and so will shed no light on two of the fundamental questions in Generative Grammar: what’s the structure of FL and why?[4]





[1] Though Baker, a really smart guy, thinks that there are so please don’t take me as endorsing the view that there aren’t any. I just don’t know. This is just my impression from linguist in the street interviews.
[2] The confirmation of this prediction was one of the great successes of generative grammar and the papers by, e.g. Kayne and Pollock, McCloskey, Chung, Torrego, and many others are still worth reading and re-reading. It is worth noting that the Move WH + BT story was largely driven by theoretical considerations, as Chomsky makes clear in OWM. The gratifying part is that the theory proved to be so empirically fecund.
[3] Note the ‘perhaps.’ If even merge is in the current parlance “third factor” then there is nothing taken to be linguistically special about FL.
[4] Note that this quite a bit of room for “learning” theory. For if the range of variation is not built into FL then why we see the variation we do must be due to how we acquire Gs given FL/UG.  The latter will still be important (indeed critical) in that any larning theory will have to incorporate the isolated invariances. However, a large part of the range of variation will fall outside the purview of FL. I discuss this somewhat in the last chapter of A theory if syntax for any of you with a prurient interest in such matters. See, in particular, the suggestion that we drop the switch analogy in favor of a more geometrical one.