Comments

Showing posts sorted by relevance for query what would GB say. Sort by date Show all posts
Showing posts sorted by relevance for query what would GB say. Sort by date Show all posts

Wednesday, November 21, 2012

How I became a minimalist and why or What would GB say?


It was apparently Max Planck who discovered the unit time of scientific change to be the funeral (the new displacing the old one funeral at a time).  In the early 1990s, I discovered a second driving force, boredom.  As some of you may know, since about the mid-1990s I have been a minimalist enthusiast. For the record, I became one despite my initial inclinations. On first reading A minimalist program for linguistic theory (a Korean bootlegged version purportedly whisked of Noam’s desk and quickly disseminated), I was absolutely convinced that it had to be on the wrong track, if not the aspirations, then the tentative conclusions. I was absolutely certain that one of the biggest discoveries of generative grammar had been the centrality of government as a core relation and S-structure as the indispensible level (I can still see myself making just these points in graduate intro syntax). Thus the idea that we dispense with government as a fundamental relation (it’s called Government-Binding theory after all!), or that we eliminate S-structure as a fundamental level (D-structure, I confess, I was willing to throw under the bus) struck me as nuts, just another manoeuver by Chomsky to annoy former graduate students.

Three things worked together to open (more accurately, pry open) my mind.

First, my default strategy is to agree with Chomsky, even if I have no idea what he’s talking about. In fact, I often try to figure out where he’s heading so that I can pre-agree with him. Sadly, he tends not to run in a straight line so I can often be seen going left when he zags right or right when he zigs left. This has proven to be both healthful (I am very fit!) and fruitful. More often than not, Chomsky identifies fecund research directions, or at least ones that in retrospect I have found interesting.  No doubt this is just dumb luck on Chomsky’s part, but if someone is lucky often enough, it is worth paying very careful attention (as my mother says: “better lucky than smart”).  So, though I have often found my work at a slant (even perpendicular) to his detailed proposals (e.g. just look at how delighted Noam is with Movement Theory of Control, a theory near and dear to my heart), I have always found it worthwhile to try to figure out what he is proposing and why. 

Second, fear: when the first minimalist paper began to circulate in the early 1990s I was invited to teach a graduate syntax seminar at Nijmegen (populated by eager, smart, hungry (and so ill-tempered) grad students from Holland and the rest of Europe) and I needed something new to talk about. If you just get up and repeat what you’ve already done, they could be ready for you. Better to move in some erratic direction and keep them guessing. Chomsky’s recent minimalist musings seemed like perfect cover.

Third, and truth be told I believe that this is the main reason, the GB stuff I/we had been exploring had become really boring. Why? For the best of possible reasons: viz. we really understood what made GB style theories tick and we/I needed something new to play with, something that would allow me/us to approach old questions in a different way (or at least not put us/me to sleep). That new thing was the Minimalist Program. I mention this, because at the time there was a lot of toing and froing about why so many had lemming-like (this is apparently a rural legend; they don’t fling themselves off cliffs) jumped off of the GB bandstand and onto the minimalist bandwagon. As I faintly recall, there was an issue of the Linguistic Review dedicated to this timely question with many authoritative voices giving very reasonable explanations for why they were taking the minimalist turn.  And most of these reasons were in fact good ones. However, if my conversion was not completely atypical, the main thrust came from simple thasaphobia and the discovery of the well-established fact that intensive study of the Barriers framework could be deleterious to one’s health (good reason to avoid going there again all you phase-lovers out there!).

These three motivations joined to prompt me, as an exercise, to stow the skepticism, at least for the duration of the Dutch lectures, assume that this minimalist stuff was on the right track and see how far I could get with it.  Much to my surprise, it did not fall apart on immediate inspection (a surprisingly good reason to persist in my experience), it was really fun to play with, and, if you got with the program, there was a lot to do given that few GB details survived minimalism’s dumping of government as a core grammatical relation (not so surprising given that it is government-binding theory).  So I was hooked, and busy. (p.s. I also enjoyed the fact that, at the time, playing minimalist partisan could get one into a lot of arguments and nothing is more fun than heated polemics).

These were the basic causes for my theoretical conversion. Were there any good reasons? Yes, one.  Minimalism was the next natural scientific step to take given the success of the GB enterprise.

This actually became more apparent to me several years later, than it was on my road to Damascus Nijmegen.  The GB era produced a rich description of the structure of UG; internally modular with distinctive conditions, primitives and operations characterizing each sub-part.  In effect, GB delivered a dozen or so “laws” of grammar (e.g. subjacency, ECP, principles A-C of binding theory, X’-theory etc.), of pretty good (no, not perfect, but pretty good) empirical standing (lots of cross linguistic support). This put generative grammar in a position to address a new kind of question: why these laws and not others? Note: you can’t ask this question if there are no “laws.” Attacking it requires that we rethink the structure of UG in a new way; not only to ask “what’s in UG ?” but also “what that is in UG is distinctively linguistic and what traceable to more general powers, cognitive, computational, or physical?”. This put a version of what we might call Darwin’s Problem (the logical problem of language evolution) on the agenda along side Plato’s Problem (the logical problem of language acquisition).  The latter has not been solved, not by a long shot, but fortunately adding a question to the research agenda does not require that previous problems have been put to bed and snuggly tucked in. So though in one sense, minimalism was nothing new, just the next reasonable scientific step to take, it was also entirely new in that it raised to prominence a question whose time, we hoped, had come. [1]

Chomsky has repeatedly emphasized the programmatic aspects of minimalism.  And, as he has correctly noted, programs are not true or false but fecund or barren. However, after 20 years, it’s perhaps (oh what a weasel word!) time to sit back and ask how fertile the minimalist turn has been? In my view, very, precisely because it has spawned minimalist theories that advance the programmatic agenda, theories that can be judged not merely in terms of their fertility but also in terms of their verisimilitude. I have my own views about where the successes lie, and I suspect that they may not coincide with either Noam’s or yours.  However, I believe that it is time that we identified what we take to be our successes and ask ourselves how (or whether?) they reflect the principle ambitions and intuitions of the minimalist program.

Let me put this another way: in one sense minimalism and GB are not competitors for the aims of the former presuppose the success of the latter.  However, minimalist theories and GB theories often are (or can be) in direct competition and it is worth evaluating them against each other.  So for example, to take an example at random (haha!), GB has a theory of control and current minimalism has several. We can ask, for example: In what ways do the GB and minimalist accounts differ? How do they stack up empirically? What minimalist precepts do the minimalist theories reflect?  What GB principles are the minimalist accounts (in)compatible with? What larger minimalist goals do the minimalist theories advance?  What does the minimalist story tells us that the earlier GB story didn’t? And vice versa? Etc. etc. etc.

IMHO, these are not questions that we have asked often enough. I believe that we have failed to effectively use GB as the foil (and measuring rod) it can be. Why? I’m not sure. Perhaps because we have concluded that because the minimalist program is worth pursuing that specific minimalist theories that brandish distinctive minimalist technology (feature checking, merge, Agree, probe-goal architecture, phases etc.) are “better” or “truer” than those exploiting the quaint out of date GB apparatus.  If so, we were wrong.  We always need to measure our advances. One good way to do this is to compare your spanking new minimalist proposal with the model T GB version. I hereby propose that going forward we adopt the mantra “What would GB say?” (WWGBS; might even make for a good license plate) and compare our novel proposals with this standard to make clear to ourselves and others where and how we’ve progressed.

I will likely blog more on this topic soon and identify what I take to be some of the more interesting lines of investigation to date.  However, I am very interested in what others take the main minimalist successes to be.  What are the parade case achievements? Let me know. After 20 years, it seems reasonable to try to make a rough estimate of how far we’ve come.



[1] Here Sean Carroll goes minimalist in a different setting:
The actual laws of nature are interesting, but it’s also interesting that there are laws at all…We want to know what those laws are. More ambitiously, we’d like to know if those laws could possibly have been different…We may or may not be able to answer such a grandiose question, but it’s the kind of thing that lights the imagination of the working scientist (p.23)
This is what I mean by the next obvious scientific step to take.  First find laws, then ask why these laws and not others. That’s the way the game is played, at least by the real pros.

Tuesday, February 11, 2014

Plato, Darwin, P&P and variation

Alex C (in the comment section here (Feb. 1)) makes a point that I’ve encountered before that I would like to comment on. He notes that Chomsky has stopped worrying about Plato’s Problem (PP) (as has much of “theoretical” linguistics as I noted in the previous post) and suggests (maybe this is too much to attribute to him, if so, sorry Alex) that this is due to Darwin’s Problems (DP) occupying center stage at present. I don’t want to argue with this factual claim, for I believe that there’s lots of truth to it (though IMO, as readers of the last several posts have no doubt gathered, theory of any kind is largely absent from current research). What I want to observe is that (1) there is a tension between PP and DP and (2) that resolving it opens an important place for theoretical speculation. IMO, one of the more interesting facets of current theoretical work is that it proposes a way of resolving this tension in an empirically interesting way. This is what I want to talk about.

First the tension: PP is the observation that the PLD the child uses in developing its G is impoverished in various ways when one compares it to the properties of Gs that children attain. PP, then, is another name for the Poverty of Stimulus Problem (POS).  Generative Grammarians have proposed to “solve” this problem by packing FL with principles of UG, many of which are very language specific (LS), at least if GB is taken as a guide to the content of FL.  By LS, I mean that the principles advert to very linguisticky objects (e.g. Subjects, tensed clauses, governors, case assigners, barriers, islands, c-command, etc) and very linguisticky operations (agreement, movement, binding, case assignment, etc.).  The idea has been that making UG rich enough and endowing it with LS innate structure will allow our theories of FL to attain explanatory adequacy, i.e. to explain how, say, Gs obey islands despite the absence of good and bad data relevant to fixing them present in the PLD. 

By now, all of this is pretty standard stuff (which is not to say that everyone buys into the scheme (Alex?)), and, for the most part, I am a big fan of POS arguments of this kind and their attendant conclusions. However, even given this, the theoretical problem that PP poses has hardly been solved. What we do have (again assuming that the POS arguments are well founded (which I do believe)) is a list of (plausibly) invariant(ish) properties of Gs and an explanation for why these can emerge in Gs in the absence of the relevant data in the PLD required to fix them. Thus, why do movement rules in a given G resist extraction from islands? Because something like the Subjacency/Barriers theory is part of every Language Acquisition Device’s (LAD) FL, that’s why.

However, even given this, what we still don’t have is an adequate account of how the variant properties of Gs emerge when planted in a particular PLD environment. Why is there V to T in French but not in English? Why do we have inverse control in Tsez but not Polish? Why wh-in-situ in Chinese but multiple wh to C in Bulgarian. The answer GB provided (and so far as I can tell, the answer still) is that FL contains parameters that can be set in different ways on the basis of PLD and the various Gs we have are the result of differential parameter setting. This is the story, but we have known for quite a while that this is less a solution to the question of how Gs emerge in all their variety than it is an explanation schema for a solution. P&P models, in other words, are not so much well worked out theories than they are part of a general recipe for a theory that were we able to cook it, would produce just the kind of FL that could provide a satisfying answer to the question of how Gs can vary so much. Moreover, as many have observed (Dresher and Janet Fodor are two notable examples, see below) there are serious problems with successfully fleshing out a P&P model.

Here are two: (i) the hope that many variant properties of Gs would hinge on fixing a small number of parameters seems increasingly empirically uncertain. Cederic Boeckx and Fritz Newmeyer have been arguing this for a while, and while their claims are debated (and by very intelligent people so, at least for a non-expert like me, the dust is still too unsettled to reach firm conclusions), it seems pretty clear that the empirical merits of earlier proposed parameterizations are less obvious than we took them to be. Indeed, there appears to some skepticism about whether there are any macro-parameters (in Baker’s sense[1]) and many of the micro-parametric proposals seem to end up restating what we observe in the data: that languages can differ. What made early macro-parameter theories interesting is the idea that differences among Gs come in largish clumps. The relation between a given parameter setting and the attested surface differences was understood as one to many. If, however, it turns out that every parameter correlates with just a single difference then the value of a parametric approach becomes quite unclear, at least so far as acquisition considerations are concerned. Why? Because it implies that surface differences are just due to differing PLD, not to the different options inherent in the structure of FL. In other words, if we end up with one parameter per surface difference then variation among Gs will not be as much of a window into the structure of FL as we thought it could be.

Here’s another problem: (ii) the likely parameters are not independent. Dresher (and friends) has demonstrated this for stress systems and Fodor (and friends) has provided analogous results for syntax.  The problem with a theory where parameters are not independent is that they make it very hard to see how acquisition could be incremental. If it turns out that the value of any parameter is conditional on the value of every other parameter (or very many others) then it would seem that we are stuck with a model in which all parameters must be set at once (i.e. instantaneous learning). This is not good! To evade this problem, we need some way of imposing independence on the parameters so that they can be set piecemeal without fear of having to re-set them later on. Both Dresher and Fodor have proposed ways of solving this independence problem (both elaborate a richer learning theory for parameter values to accommodate this problem). But, I think that it is fair to say that we are still a long way from a working solution. Moreover, the solutions provided all involve greatly enriching FL in a very LS way. This is where PP runs into DP. So let’s return to the aforementioned tension between PP and DP.

One way to solve PP is to enrich FL. The problem is that the richer and more linguistically parochial FL is, the harder it becomes to understand how it might have evolved. In other words, our standard GB tack in solving PP (LS enrichment of FL) appears to make answering DP harder. Note I say ‘appears.’ There are really two problems, and they are not equally acute. Let me explain.

As noted above, we have two things that a rich FL has been used to explain; (a) invariances characteristic of all Gs and (b) the attested variation among Gs. In a P&P model, the first ‘P’ handles (a) and the second (b). I believe that we have seen glimmers of how to resolve the tension between PP’s demands on FL versus DP’s as regards the principles part of P&P. Where things have become far more obscure (and even this might be too kind) involves the second parametric P. Here’s what I mean.

As I’ve argued in the past, one important minimalist project has been to do for the principles of GB what Chomsky did for islands and movement via the theory of subjacency in On Wh Movement (OWM). What Chomsky did in this paper is theoretically unify the disparate island effects by unifying all non-local (A’) dependency constructions by proposing that they have a common movement core (viz. move WH) subject to locality restrictions characterized by Bounding Theory (BT). This was terrifically inventive theory and aside from rationalizing/unifying Ross’s very disparate Island Effects, the combination of Move WH + BT predicted that all long movement would have to be successive cyclic (and even predicted a few more islands, e.g. subject islands and Wh-islands).[2]

But to get back to PP and DP, one way of regarding MP work over the last 20 years is as an attempt to do for GB modules what Chomsky did for Ross’s Islands. I’ve suggested this many times before but what I want to emphasize here is that this MP project is perfectly in harmony with the PP observation that we want to explain many of the invariances witnessed across Gs in terms of an innately structured FL. Here there is no real tension if this kind of unification can be realized. Why not? Because if successful we retain the GB generalizations. Just as Move WH + BT retain Ross’s generalizations, a successful unification within MP will retain GB’s (more or less) and so we can continue to tell the very same story about why Gs display the invariances attested as we did before. Thus, wrt this POS problem, there is a way to harmonize DP concerns with PP concerns. Of course, this does not mean that we will successfully manage to unify the GB modules in a Move WH + BT way, but we understand what a successful solution would look like and, IMO, we have every reason to be hopeful, though this is not the place to defend this view.

So, the principles part of P&P is, we might say, DP compatible (little joke here for the cognoscenti). The problem lies with the second P. FL on GB was understood to provide not only the principles of invariance but also to specify all the possible ways that Gs could differ. The parameters in GB were part of FL! And it is hard to see how to square this with DP given the terrific linguistic specificity of these parameters. The MP conceit has been to try and understand what Gs do in terms of one (perhaps)[3] linguistically specific operation (Merge) interacting with many general cognitive/computational operations/principles.  In other words, the aim has been to reduce the parochialism of the GB version of FL. The problem with the GB conception of parameters is that it is hard to see how to recast them in similarly general terms. All the parameters exploit notions that seem very very linguo-centric. This is especially true of micro parameters, but it is even true of macro ones. So, theoretically, parameters present a real problem for DP, and this is why the problems alluded to earlier have been taken by some (e.g. me) to suggest that maybe FL has little to say about G-variation. Moreover, it might explain why it is that, with DP becoming prominent, some of the interest in PP has seemed to wane. It is due to a dawning realization that maybe the structure of FL (our theory of UG) has little to say directly about grammatical variation and typology. Taken together PP and DP can usefully constrain our theories of FL, but mainly in licensing certain inferences about what kinds of invariances we will likely discover (indeed have discovered). However, when it comes to understanding variation, if parameters cannot be bleached of their LSity (and right now, this looks to me like a very rough road), it looks to me like they will never be made to fit with the leading ideas of MP, which are in turn driven by DP. 

So, Alex C was onto something important IMO. Linguists tend to believe that understanding variation is key to understanding FL. This is taken as virtually an article of faith. However, I am no longer so sure that this is a well founded presumption. DP provides us with some reasons to doubt that the range of variation reflects intrinsic properties of FL. If that is correct, then variation per se may me of little interest for those interested in liming the basic architecture of FL. Studying various Gs will, of course, remain a useful tool for in getting the details of the invariant principles and operations right. But, unlike earlier GB P&P models, there is at least an argument to be made (and one that I personally find compelling) that the range of G-variation has nothing whatsoever to do with the structure of FL and so will shed no light on two of the fundamental questions in Generative Grammar: what’s the structure of FL and why?[4]





[1] Though Baker, a really smart guy, thinks that there are so please don’t take me as endorsing the view that there aren’t any. I just don’t know. This is just my impression from linguist in the street interviews.
[2] The confirmation of this prediction was one of the great successes of generative grammar and the papers by, e.g. Kayne and Pollock, McCloskey, Chung, Torrego, and many others are still worth reading and re-reading. It is worth noting that the Move WH + BT story was largely driven by theoretical considerations, as Chomsky makes clear in OWM. The gratifying part is that the theory proved to be so empirically fecund.
[3] Note the ‘perhaps.’ If even merge is in the current parlance “third factor” then there is nothing taken to be linguistically special about FL.
[4] Note that this quite a bit of room for “learning” theory. For if the range of variation is not built into FL then why we see the variation we do must be due to how we acquire Gs given FL/UG.  The latter will still be important (indeed critical) in that any larning theory will have to incorporate the isolated invariances. However, a large part of the range of variation will fall outside the purview of FL. I discuss this somewhat in the last chapter of A theory if syntax for any of you with a prurient interest in such matters. See, in particular, the suggestion that we drop the switch analogy in favor of a more geometrical one.

Thursday, December 12, 2013

Simultaneous rule application: Help!!!

Lately I have been thinking of something and have gotten stuck. Very stuck. This post is a request for help.  Here’s the problem. It relates to some current minimalist technology and how it relates to the bigger framework assumptions of the enterprise.  Here’s what I don’t quite get: what’s it means to say that rules apply “all at once” at Spell Out.  Let me elaborate.

A recent minimalist innovation is the proposal that congeries of operations apply simultaneously at Spell Out (SO).  The idea of operations applying all at once is not in and of itself problematic, for it is easy to imagine that many rules can apply “in parallel.”  However, when rules so apply, they are informationally encapsulated in the sense that the output of rule A does not condition the application of rule B. ‘Condition’ here means neither feeds nor bleeds its application. When rules do feed and bleed one another, then the idea that they all apply “all at once” is hard (at least for me) to understand, for if the application of B logically requires information about the output of A then how could they apply “in parallel.” But if they are not applying “in parallel” what exactly does it mean to say that the rules apply “all at once”?

One answer to this question is that I am getting entangled in a preconception, namely my confusion is the consequence of a “derivational” mindset (DM). The DM picture treats derivations like proofs, each line licensed by some rule applying to the preceding lines.[1] The “all at once” idea is rejecting this picture and is suggesting in its place a more model theoretic idiom in which sets of constraints together vet a given object for well-formedness. An object is well formed not if derivable from rules sequentially applied, but no matter how constructed it meets all the relevant constraints.  This should be familiar to those with GBish or OTish educations, for GB and OT are very much “freely generate and filter” kinds of models, the filters being the relevant constraints.[2] If this is correct, then the current suggestion about simultaneous rule application at SO is more accurately understood as a proposal to dump the derivational conception of grammar characteristic of earlier minimalism in favor of a constraint based approach of the GB variety.

Note, that to get something like this to work, we would need some way of constructing the objects that the constraints inspect.  In GB this was the province of Phrase Structure Rules and ‘Move alpha.’ These two kinds of rules applied freely and generated the structures and dependencies that filters like Principle A/B/C, ECP, Subjacency, etc. vetted. In an MP setting, it is harder to see how this gets done, at least to me.  Recall that in an MP setting, there is only one “rule,” (i.e. Merge). So, I assume that it would generate the relevant structures and these would be subsequently vetted. In effect the operations of the computational system (i.e. Merge, Agree and anything else, e.g. Feature Transfer, Probing, ???) would apply freely and then the result would be vetted for adequacy. What would this consist in? Well, I assume checking the resultant structures for Minimality, Extension, Inclusiveness, etc.  The problem, then, would be to translate these principles, which are easy enough to picture when thought of derivationally, into constraints on freely generated structures. I confess that I am not sure how to do this.  Consider the Extension condition. How is one to state this as a well-formedness condition on derived structures rather than on the operations that determine how the structures are derived? Ditto on steroids for Derivational Economy (aka: Merge over Move) or the idea that shorter derivations trump longer ones, or determining what constitutes a chain (which are the copies that form a chain?).  Are there straightforward ways of coding these as output conditions in freely generated objects?  If so, what are they?

There is another subsidiary more conceptual concern. In early Minimalism output conditions (aka filters) were understood as Bare Output Conditions (BOCs). BOCs were legibility conditions that interfaces, particularly CI, imposed on linguistic products. Now, BOCs were not intended to be linguistic, though they imposed conditions on linguistic objects. This means that whatever filter one proposes needs to have a BOC kind of interpretation. This was always my problem with, for example, the Minimal Link Condition (MLC). Do we really think that chains are CI objects and that “thoughts” impose locality conditions on their interacting parts? Maybe, but, I’m dubious. I can see minimality arising naturally as a computational fact about how derivations proceed. I find it harder to see it as a reflection of how thoughts are constructed.  However, whatever one thinks of the MLC, understanding Economy or Extension or Phase Impenetrability or Inclusiveness as BOCs seems, at least to me, more challenging still.

Things get even hairier, I think, when one considers that range of operations supposed to happen “all at once.” So, for example, If the features of T are inherited from C (as currently assumed) and I-merge is conditioned by Agree, then this suggests that DPs move to Spec T conditional on C having merged with T. But any such movement must violate Extension. The idea seems to be that this is not a problem if all the indicated operations apply simultaneously. But how is this accomplished?  How can I-merge be conditioned (fed) by features that are only available under operations that require that a certain structure exists (i.e. C and “TP” have E-merged) but whose existence would preclude Merging the DP (doing so would violate Extension).  One answer: screw Extension. Is this what is being suggested?  If not, what?

So, I throw myself on the mercy of those who have a better grasp of the current technology. What is involved in doing operations “all at once”? Are we dumping derivations and returning to a generate-and-filter model? What do we do with apparent bleeding and feeding relations and the dependencies that exploit these notions. Which principles are we to retain, and which dispense with? Extension? Economy? Minimality? How to the rules/operations work? Sample examples of “derivations” would be nice to see.  If anyone knows the answer to all or any of these questions, please let me know.



[1] A strong version of this is that it is only the immediately preceding line can influence what happens “next.”
[2] OT’s filters are ranked, whereas GB filters were not.  However, I don’t believe that this difference makes a difference for my problem.

Thursday, September 28, 2017

Physics envy and the dream of an interpretable theory

I have long believed that physics envy is an excellent foundation for linguistic inquiry (see here). Why? Because physics is the paradigmatic science. Hence, if it is ok to do something there it’s ok to do it anywhere else in the sciences (e.g. including in the cog-neuro (CN) sciences, including linguistics) and if a suggested methodological precept fails for physics, then others (including CNers) have every right to treat it with disdain. Here’s a useful prophylactic against methodological sadists: Try your methodological dicta out on physics before you encumber the rest of us with them. Down with methodological dualism!

However, my envy goes further: I have often looked to (popular) discussions about hot topics in physical theory to fuel my own speculations. And recently I ran across a stimulating suggestive piece about how some are trying to rebuild quantum theory from the ground up using simple physical principles (QTFSPP) (here). The discussion is interesting for me in that it leads to a plausible suggestion for how to enrich minimalist practice. Let me elaborate.

The consensus opinion among physicists is that nobody really understands quantum mechanics (QM). Feynman is alleged to have said that anyone who claims to understand it, doesn’t. And though he appears not to have said exactly this (see here section 9), it's a widely shared sentiment. Nonetheless, QM (or the Standard Theory) is, apparently, the most empirically successful theory ever devised. So, we have a theory that works yet we have no real clarity as to why it works. Some (IMO, rightly) find this a challenge. In response they have decided to reconstitute QM on new foundations. Interestingly, what is described are efforts to recapture the main effects of QM within theories with more natural starting points/axioms. The aim, in other words, is reminiscent of the Minimalist Program (MP): construct theories that have the characteristic signature properties of QM but are grounded in more interpretable axioms. What’s this mean? First let’s take a peak at a couple of examples from the article and then return to MP.

A prominent contrast within physics is between QM and Relativity. The latter (the piece mentions special relativity) is based on two fundamental principles that are easy to understand and from which all the weird and wonderful effects of relativity follow. The two principles are: (1) the speed of light is constant and (2) the laws of physics are the same for two observers moving at constant speed relative to one another (or, no frame of reference is privileged when it comes to doing physics). Grant these two principles and the rest follows. As QTFSPP outs it: “Not only are the axioms simple, but we can see at once what they mean in physical terms” (my emphasis, NH) (5).

Standard theories of QM fail to be physically perspicuous and the aim of reconstructionists is to remedy this by finding principles to ground QM as natural and physically transparent as those that Einstein found for special relativity.  The proposals are fascinating. Here are a couple:

One theorist, Lucien Hardy, proposed focusing on “the probabilities that relate the possible states of a system with the chance of observing each state in a measurement” (6). The proposal consists of a set of probabilistic rules about “how systems can carry information and how they can be combined and interconverted” (7). The claim was that “the simplest possible theory to describe such systems is quantum mechanics, with all its characteristic phenomena such as wavelike interference and entanglement…” (8). Can any MPer fail to reverberate to the phrase “the simplest possible theory”? At any rate, on this approach, QMs is fundamentally probabilistic and how probabilities mediate the conversion between states of the system are taken as the basic of the theory.  I cannot say that I understand what this entails, but I think I get the general idea and how if this were to work it would serve to explain why QM has some of the odd properties it does.

Another reconstruction takes three basic principles to generate a theory of QM. Here’s QTFSPP quoting a physicist named Jacques Pienaar: “Loosely speaking, their principles state that information should be localized in space and time, that systems should be able to encode information about each other, and that every process should be in principle reversible, so that information is conserved.” Apparently, given these assumptions, suitably formalized, leads to theories with “all the familiar quantum behaviors, such as superposition and entanglement.” Pienaar identifies what makes these axioms reasonable/interpretable: “They all pertain directly to the elements of human experience, namely what real experimenters ought to be able to do with systems in their laboratories…” So, specifying conditions on what experimenters can do in their labs leads to systems of data that look QMish. Again, the principles, if correct, rationalize the standard QM effects that we see. Good.

QTFSPP goes over other attempts to ground QM in interpretable axioms. Frankly, I can only follow this, if at all, impressionistically as the details are all quite above my capacities. However, I like the idea. I like the idea of looking for basic axioms that are interpretable (i.e. whose (physical) meaning we can immediately grasp) not merely compact. I want my starting points to make sense too. I want axioms that make sense computationally, whose meaning I can immediately grasp in computational terms. Why? Because, I think that our best theories have what Steven Weinberg described as a kind of inevitability and they have this in virtue of having interpretable foundations. Here’s a quote (see here and links provided there):

…there are explanations and explanations.  We should not be satisfied with a theory that explains the Standard Model in terms of something complicated an arbitrary…To qualify as an explanation, a fundamental theory has to be simple- not necessarily a few short equations, but equations that are based on a simple physical principle…And the theory has to be compelling- it has to give us the feeling that it could scarcely be different from what it is. 

Sensible interpretable axioms are the source of this compulsion. We want first principles that meet the Wheeler T-shirt criteria (after John Wheeler): they make sense and are simple enough to be stated “in one simple sentences that the non sophisticate could understand,” (or, more likely, a few simple sentences). So, with this in mind, what about fundamental starting points for MP accounts. What might these look like?

Well, first, they will not look like the principles of GB. IMO, these principles (more or less) “work,” but they are just too complicated and complex to be fundamental. That’s why GB lacks Weinberg’s inevitability. In fact, it takes little imagination to imagine how GB could “be different.” The central problem with GB principles is that they are ad hoc and have the shape they do precisely because the data happens to have the shape it does. Put differently, were the facts different we could rejigger the principles so that they would come to mirror those facts and not be in any other way the worse off for that. In this regard, GB shares the problem QTFSPP identifies with current QM: “It’s a complex framework, but it’s also an ad hoc patchwork, lacking any obvious physical interpretation or justification” (5).

So, GB can’t be fundamental because it is too much of a hodgepodge. But, as I noted, it works pretty well (IMO, very well actually, though no doubt others would disagree). This is precisely what makes the MP project to develop a simple natural theory with a specified kind of output (viz. a theory with the properties that GB describes) worthwhile.

Ok, given this kind of GB reconstruction project, what kinds of starting points would fit?  I am about to go out on a limb here (fortunately, the fall, when it happens, will not be from a great height!) and suggest a few that I find congenial.

First, the fundamental principle of grammar (FPG)[1]: There is no grammatical action at a distance. What this means is that for two expressions A and B to grammatically interact, they must form a unit. You can see where this is going, I bet: for A and B to G interact, they must Merge.[2]

Second, Merge is the simplest possible operation that unitizes expressions. One way of thinking of this is that all Merge does is make A and B, which are heretofore separate, into a unit. Negatively, this implies that it in no way changes A and B in making them a unit, and does nothing more than make them a unit (e.g. negatively, it imposes no order on A and B as this would be doing more than unitizing them). One can represent this formally as saying that Merge takes A,B and forms the set {A,B}, but this is not because Merge is a set forming operation, but because sets are the kinds of objects that do nothing more than unitize the objects that form the set. They don’t order the elements or change them in any way. Treating Merge (A,B) as creating leaves of a Calder Mobile would have the same effect and so we can say that Merge forms C-mobiles just as well as we can say that it forms sets. At any rate, it is plausible that Merge so conceived is indeed as simple a unitizing operation as can be imagined.

Third, Merge is closed in the domain of its application (i.e. its domain and range are the same). Note that this implies that the outputs of Merge must be analogous to lexical atoms in some sense given the ineluctable assumption that all Merges begin with lexical atoms. The problem is that unitized lexical atoms (the “set”-likeoutputs of Merge) are not themselves lexical atoms and so unless we say something more, Merge is not closed. So, how to close it? By mapping the Merged unit back to one of the elements Merged in composing it. So if we map {A,B} back to A or to B we will have closed the operation in the domain of the primitive atoms. Note that by doing this, we will, in effect, have formed an equivalence class of expressions with the modulus being the lexical atoms. Note, that this, in effect, gives us labels (oh nooooo!), or labeled units (aka, constituents) and endorses an endocentric view of labels. Indeed, closing Merge via labeling in effect creates equivalence classes of expressions centered on the lexical atoms (and more abstract classes if the atoms themselves form higher order classes). Interestingly (at least to me) so closing Merge allows for labeled objects of unbounded hierarchical complexity.[3]

These three principles seem computationally natural. The first imposes a kind of strict locality condition on G interactions. E and I merge adhere to it (and do so strictly given labels). Merge is a simple, very simple, combination operation and closure is a nice natural property for formal systems of (arbitrarily complex) “equations” to have. That they combine to yield unbounded hierarchically structured objects of the right kind (I’ve discussed this before, see here and here) is good as this is what we have been aiming for. Are the principles natural and simple? I think so (at least form a kind of natural computation point of view), but I would wouldn’t I?  At any rate, here’s a stab at what interpretable axioms might look like. I doubt that they are unique, but I don’t really care if they aren’t. The goal is to add interpretatbility to the demands we make on theory, not to insist that there is only one way to understand things.

Nor do we have to stop here. Other simple computational principles include things like the following: (i) shorter dependencies are preferred to longer dependencies (minimality?), (ii) bounded computation is preferred to unbounded computation (phases?), (iii) All features are created equal (the way you discharge/check one is the way you discharge/check all). The idea is then to see how much you get starting from these simple and transparent and computationally natural first principles. If one could derive GBish FLs from this then it would, IMO, go some way towards providing a sense that the way FL is constructed and its myriad apparent complexities are not complexities at all but the unfolding of a simple system adhering to natural computational strictures (snowflakes anyone?). That, at least, is the dream.

I will end here. I am still in the middle of pleasant reverie, having mesmerized myself by this picture. I doubt that others will be as enthralled, but that is not the real point. I think that looking for general interpretable principles on which to found grammatical theory makes sense and that it should be part of any theoretical project. I think that trying to derive the “laws” of GB is the right kind of empirical target. Physics envy prompts this kind of search. Another good reason, IMO, to cultivate it.



[1] I could have said, the central dogma of syntax, but refrained. I have used FPG in talks to great (and hilarious) effect.
[2] Note, that this has the pleasant effect of making AGREE (and probe-goal architectures in general) illicit G operations. Good!
[3] This is not the place to go into this, but the analogy to clock arithmetic is useful. Here too via the notion of equivalence classes it is possible to extend operations defined for some finite base of expressions (1-12) to any number. I would love to be able to say that this is the only feasible way of closing a finite domain, but I doubt that this is so. The other suspects however are clearly linguistically untenable (e.g. mapping any unit to a constant, mapping any unit randomly to some other atom). Maybe there is a nice principle (statable on one simple sentence) that would rule these out.