Comments

Showing posts with label Chomsky vs Greenberg universals. Show all posts
Showing posts with label Chomsky vs Greenberg universals. Show all posts

Monday, April 9, 2018

Methodological sadism

Methodological sadism (MS) is quite fashionable nowadays and nothing gets practitioners more excited than the possibility that someone somewhere is proposing something interesting (i.e. something that reaches beyond the sensory surface of things and that might possibly reveal some of the underlying mechanics of reality). You’ve all met people like this,[1] and one of their distinctive character traits is a certain (smug?) assurance that when it comes to the philosophy of science, they are on the side of the angels. They love to methodologically demarcate the boundaries of legitimate inquiry so as to protect the weak minded from fake science.

Of course, the standard demeanor of MSers is severe. Yes they are tough. But standards must be maintained lest we slide joyfully to our scientific perdition. Like I said, you’ve all met MSers. Nowadays, at least in my little domain of inquiry, they are the media stars and have done a pretty good job convincing the outside world (and some on the inside) that GG is dead and that there is really nothing special about the cognitive powers required for language. I think that this is deeply wrong, and will write another brief arguing as much in the next post. But for now, I want to, once again, offer some prophylaxis against the most rabid form of MS, falsification.  Here is a useful short antidote, a paper (which I will refer to as ‘AB’ (Adam Becker is author)) that touches all the right themes. Its main claim is that trying to demarcate science from non-science is a mugs game that relies on ignoring how real successful domains of inquiry have grown.

So what are the main themes?

First AB points out that SMers (my term not AB’s) adhere to a basic erroneous principle: “that a new theory shouldn’t invoke the undetectable” (2).[2] Why? Because this makes it “unfalsifiable.”  So, observability underlies falsifiability and both are used by “self-appointed guardian[s], who relish dismissing some of the more fanciful notions in physics, cosmology and quantum mechanics [and linguistics! NH] as just so many castles in the sky” in order to protect science from “from all manner of manifestly unscientific nonsense” (2).

There are ways of understanding falsifiability that seem unobjectionable, namely that theories that never can have observable consequences are thereby undesirable. Well, yeah. The problem is that this is a very low bar, and any stronger version that “turn[s] ingenuity into fact” must be “much more nuanced” (2). Why? Because falsifiability is hardly ever possible and observability is undefinable. Let’s consider both of these facts seriatim.

First falsifiability. This is impossible for scientific theories for the simple reason that any falsification can be patched up with the right ad hoc statement, leaving the rest of the theory the same. As AB puts it (correctly) (2):

Falsifiability doesn’t work as a blanket restriction in science for the simple reason that there are no genuinely falsifiable scientific theories. I can come up with a theory that makes a prediction that looks falsifiable, but when the data tell me it’s wrong, I can conjure some fresh ideas to plug the hole and save the theory.

Any linguist knows how true this is. Moreover, it is easier the less brittle a theory is, and our current theories tend to be very labile. You can bend them in many directions without ever hearing a creak let alone inducing a crack or a break. You loose nothing when adding a bespoke principle to explain recalcitrant data because the only thing there is to loose in doing this is explanatory power and there was not much of this to begin with in flexible theories. However, even with good theories that have some oomph, it is generally possible (I would say “always possible” but I am being mealy mouthed here) to plug the hole and carry on.  One of the virtues of AB is that it provides some nice historical examples of this happening. As AB notes, the history of science is full of them.

AB recounts the famous one where Uranus’ odd (apparently non Newtonian) orbit begats Neptune, which in turn begats Vulcan to explain Mercury’s perihelion which finally fails when General Relativity replaces Newton. AB notes that each historical move makes sense and that looking for Neptune (victory!) and looking for Vulcan (failure!) were both rational despite the different outcomes.  Of course, with hindsight, Neptune is a bold prediction that strengthens the theory and Vulcan turns out to have been just an unfortunate wrong turn. But Vulcan did not lead to people jumping the Newtonian ship. Rather “astronomers of the time collectively shrugged and moved on” (4). And rightly so. Do you really want to give up Newton just because Vulcan was impossible to spot? What then do you do with all the other stuff it does explain?

Note that this means that there is an asymmetry between potentially falsifying experiments that succeed and those that don’t. The former are declared as triumphs of the scientific will, while the latter are quietly shelved and de-emphasized (or, more accurately, added to the ledger of anomalies that a discipline collects as targets for yet unknown superior explanations yet to come). No reason to show these off in public and take the shine from the powerful rational methods of scientific inquiry.

Of course, one might eventually hit the jackpot and find the anomalies resolved with the right new theory. Mercury was a feather in General Relativity’s cap. And well it should have been, for it allowed us to dump Vulcan and replace it with a story that allowed us to also keep all the good parts of Newton. So yes exceptions prove the rule in the sense that Newtonian exceptions prove (justify) Relativity’s rules.

AB provides other examples, one of the best being Pauli’s proposal to save the Law of the Conservation of Energy via the neutrino. It took over 25 years to prove him right (it’s good to have descendants like Fermi who get interested in your ideas). But what is interesting is not merely the long wait time, but why it took so long: the neutrino had no properties at the time that Pauli proposed it that could have allowed it to be detected. So when Pauli proposed it, the theory was unfalsifiable because its basics were unobservable. But, as history shows, wait 25 years and who knows what unobservables might become detectable. As AB again, rightly, puts it (5):

It’s certainly true that observation plays a crucial role in science. But this doesn’t mean that scientific theories have to deal exclusively in observable things. For one, the line between the observable and unobservable is blurry – what was once ‘unobservable’ can become ‘observable’, as the neutrino shows. Sometimes, a theory that postulates the imperceptible has proven to be the right theory, and is accepted as correct long before anyone devises a way to see those things.

Not only is ‘observable’ irreparably vague, but MSers also often make a second more extreme demand concerning its applicability. They require theory to only postulate constructs with directly observable magnitudes. Laws of nature then simply relate “directly observable quantities” and make no reference “to anything unobservable at all” (6). This was essentially Mach’s views and, in hindsight, they served to hamstring scientific insight. For example, as AB notes, Mach opposed atomic theory as unscientific. Why? Well you cannot “see” atoms. Of course, as many pointed out, there is lots to be gained by postulating them (e.g. you can derive the principles of thermodynamics and explain Brownian motion). But this was not good enough for Mach, nor for current day MSers. So how did Mach’s views hold up? Well, it cost Walter Kaufman a Nobel Prize apparently and might have led to Boltzmann’s suicide, but aside from that it did not do any real damage because the scientific community largely ignored Mach’s injunctions.

But isn’t postulating unseen elements unscientific? Well no. Note: even if we all agree that theory must have discernible (aka observable) consequences so that it can be tested/verified, this does not imply that every part of the theory must invoke elements whose properties are directly observable. Of course, it is nice if one can do this. It is always nice to be able to measure. But the idea that the only thing worth doing is relating (perhaps statistically) the magnitudes of observable quantities is something that would sink our best sciences if implemented. And this is a good reason not to do it![3]

One can go further, and AB does. The notion of an observable is itself fundamentally obscure. It cannot mean, observable given “current” technology, for that is too strong. As the history of the neutrino “discovery” indicates, taking this position would have led away from the truth, not towards it. But it cannot mean “observable in principle” for this is irremediably vague. If we mean that it is not “logically possible” to observe what but contradictions will fail. If we mean given current technology or current theory it is too strong. So what then? AB, quotes Grover Maxwell as observing” “There is no a priori or philosophical criteria for separating the observable form the unobservable.” The best we can say is that theories with observable consequences are better situated than those without ceteris paribus. But what goes into the ceteris paribus determination is forever up for grabs, and subject to the inconclusive (yet critically important) vagaries of judgment. No method, just mucking around, always.

AB makes a last observation I’d like to highlight: unlike many, AB emphasizes that part of the scientific enterprise is building “sky-castles.” This is not a scientific aberration, nor an example of science misfiring, but part of the central enterprise. For the scientist (including the linguist see here) “[s]pinning new ideas about how the world could be – or in some cases, how the world definitely isn’t – is central to their work” (2). Explanation leans heavily on the modal ‘could’ in the quote. Not just what you see or mild extensions thereof, but what could be and couldn’t. That’s the stuff of understanding and as AB notes, again rightly, “[the] goal of scientific theory is to understand [my emphasis, NH] the nature of the world with increasing accuracy over time.”

Methodological sadists, if given power, would sink scientific inquiry. They would make it nearly impossible to uncover unobservable mechanisms for they are an inherent part of all decent explanation in the sciences. MSers undervalue explanation and hence distrust the speculation required to get any. As AB notes, their dicta are at odds with the history of science. They are also deeply obscure. So historically misguided and irredeemably obscure? Yes, but also sadistically useful. MSers sound tough minded (just the facts kinda people) but really they are hopeless romantics, stuck with a view of method and inquiry that successful inquiry has largely ignored, as linguists should as well, at least if they want to get anywhere.



[1] Pullum, Haspelmath, Tomasello and Everett are prominent examples of such in my own little world.
[2] The strong form would say that no theory should, not only new ones. However, MSers generally aspire to be gatekeepers and, in practice, this means keeping out the new. Facing out, rather than in, also has one important advantage. It is pretty hard to argue that accepted results are suspect without making one’s methodological injunctions sound dumb (recall, that every modus ponens comes with an equally powerful modus tolens). Consequently, fire is reserved for the novel, which is always deemed to differ from the accepted in being methodologically deficient. Note that being methodologically deficient has its virtues in argument. It relieves the critic of actually having to go into details, of having to do the hard work of arguing against actual results. MSers generally paint with a broad methodological brush, and I would argue that this is the reason why.
[3] Linguists here should be thinking of those that take Greenberg universals to be the only kinds that are legit (e.g. the crowd in note 1). Why do this? Well because as MSers they demand that science eschew the unobservable. Greenberg universals just are estimates of co-occurrence (either categorical or probabilistic) among surface visible language properties. Chomsky universals are not, and this is why for MSers Chomsky Universals are verboten.

Sunday, September 11, 2016

Universals: a consideration of Everett's full argument

I have consistently criticized Everett’s Piraha based argument against Chomsky’s conception about Universal Grammar (UG) by noting that the conclusions only follow if one understands ‘universal’ in Greenberg rather than Chomsky terms (e.g. see here).  I have recently discovered that this is correct as far as it goes, but it does not go far enough. I have just read this Everett post, which indicates that my diagnosis was too hasty. There is a second part to the argument and the form is actually one of a dilemma: either you understand Chomsky’s claims about recursion as a design feature of UG in Greenbergian terms OR Chomsky’s position is effectively unfalsifiable (aka: vacuous). That’s the full argument.  I (very) critically discuss it in what follows. The conclusion is that not only does it fail to understand the logic of a CU conception of universal, it also presupposes a rather shallow Empiricist conception of science, one in which theoretical postulates are only legitimate if directly reflected in surface diagnostics. Thus, Everett’s argument gains traction only if one mistakes Chomsky Universals (CU) for Greenberg Universals (GUs), misunderstands what kind of evidence is relevant for testing CUs and/or tacitly assumes that only GUs are theoretically legit conceptions in the context of linguistic research. In short, the argument still fails, even more completely than I thought.

Let’s give ourselves a little running room by reviewing some basic material. GUs are very different from CUs. How so?

GUs concern the surface distributional properties of the linguistic objects (LO) that are the outputs of Gs. They largely focus on the string properties of these LOs.  Thus, one looks for GUs by, for example, looking at surface distributions cross linguistically. For example, one looks to see if languages are consistent in their directionality parameters (If ‘OP’ then ‘OV’). Or if there are patterns in the order in nominals of demonstratives, modifiers and numerals wrt the heads they modify (e.g. see here for some discussion).

CUs specify properties of the Faculty of Language (FL). FL is the name given to the mental machinery (whatever its fine structure) that outputs a grammar for L (GL) given Primary Linguistic Data from language L (PLDL). FL has two kinds of design features. The linguistically proprietary ones (which we now call UG principles) versus the domain general ones, which are part of FL but not specific to it. GGers  investigate the properties of FL by, first, investigating the properties of language particular Gs and second, via the Poverty of the Stimulus argument (POS). POS aims to fix the properties of FL by seeing what is needed to fill the gap between information provided about the structure of Gs in the PLD and the actual properties that Gs have. FL has whatever structure is required to get the Language Acquisition Device (LAD) from PLDL to GL for any L. Why any L? Because any kid can acquire any G when confronted with the appropriate PLD. 

Now on the face of it, GUs and CUs are very different kinds of things. GUs refer to the surface properties of G outputs. CUs refer to the properties of FL, which outputs Gs. CUs are ontologically more basic than GUs[1] but GUs are less abstract than CUs and hence epistemologically more available.

Despite the difference between GUs and CUs, GGers sometimes use string properties of the outputs of Gs to infer properties of the Gs that generate these LOs. So, for example, in Syntactic Structures Chomsky argues that human Gs are not restricted to simple finite state grammars because of the existence of sentences that allow non-local dependencies of the sort seen in sentences of the form ‘If S1 then S2’. Examples like this are diagnostic of the fact that the recursive Gs native speakers can acquire must be more powerful than simple FSGs and therefore that FL cannot be limited to Gs with just FSG rules.[2]  Nonetheless, though GUs might be useful in telling you something about CUs, the two universals are conceptually very different, and only confusion arises when they are run together.

All of this is old hat, and I am sorry for boring you. However, it is worth being clear about this when considering the hot topic of the week, recursion, and what it means in the context of GUs and CUs. Let’s recall Everett’s dilemma. Here is part 1:

1.     Chomsky claims that Merge is “a component of the faculty of language,” (i.e. that it is a Universal).[3]
2.     But if it is a universal then it should be part of the G of every language.
3.     Piraha does not contain Merge.
4.     Therefore Chomsky is wrong that Merge is a Universal.

This argument has quite a few weak spots. Let’s review them.

First, as regards the premises (1) and (2), the argument requires assuming that if something is part of FL, a CU, then it appears in every G that is a product of FL. For unless we assume this, it cannot be that the absence of Merge in Piraha is inconsistent with the conclusion that it is part of FL. But, the assumption that Merge is a CU does not imply that it is a GU. It simply implies that FL can construct Gs with that embody (recursive) Merge.[4] Recall, CUs describe the capacities of the LAD not its Gish outputs. FL can have the capacity to construct Merge containing Gs even if it can also construct Gs that aren’t Merge containing Gs. Having the capacity to do something does not entail that the capacity is always (or even ever) used. This is why a claim like Everett’s that argues from (3), the absence of Merge in the G of Piraha, does not argue against Merge as part of FL.

Second, what is the evidence that Merge is not part of Piraha’s G? Everett points to the absence of “recursive structures” in Piraha LOs (2). What are recursive structures? I am not sure, but I can hazard a guess. But before I do so, let me note that recursion is not properly a predicate of structures but of rules. It refers to rules that can take their outputs as inputs. The recursive nature of Merge can be seen from the inductive definition in (5):

5.   a. If a is a lexical item then a is a Syntactic Object (SO)
b. If a is an SO and b is an SO then Merge(a,b) is an SO

With (5) we can build bigger and bigger SOs, the recursive “trick” residing in the inductive step (5b). So rules can be recursive. Structures, however, not so much. Why? Well they don’t get bigger and bigger. They are what they are. However, GGers standardly illustrate the fact of recursion in an L by pointing to certain kinds of structures and these kinds have come to be fairly faithful diagnostics of a recursive operation underlying the illustrative structures. Here are two examples both of which have a phrase of type A embedded in another one of type A.

6.     S within an S: e.g. John thinks that Bill left in which the sentence Bill left is contained within the larger sentence John thinks that Bill left.
7.     A nominal within a nominal: e.g.  John saw a picture of a picture where the nominal a picture is contained within the larger nominal a picture of a picture.

Everett follows convention and assumes that structures of this sort are diagnostic of recursive rules. We might call them reliable witnesses (RW) for recursive rules. Thus (6)/(7) are RWs for the claim that the rule for S/nominal “expansion” can apply repeatedly (without bound) to their outputs.

Let’s say this is right. It does not imply that the absence of RWs implies the absence of recursive rules. As it is often said: the absence of evidence is not evidence of absence. Merge may be applying even though we can find no RWs diagnostic of this fact in Piraha.

Moreover, Everett’s post notes this. As it says: “…the superficial appearance of lacking recursion does not mean that the language culd not be derived from a recursive process like Merge. And this is correct” (2-3). Yes it is. Merge is sufficient to generate the structures of Piraha. So, given this, how can we know that Piraha does not employ a Merge like operation?

So far as I can tell, the argument that it doesn’t is based on the assumption that unless one has RWs for some property X one cannot assume that X is a characteristic of G. So absent visible “recursive structures” we cannot assume that Merge obtains in the G that generates these structures. Why?  Because Merge is capable of generating unboundedly big (long and deep) structures, and we have no RWs indicating that the rule is being recursively applied. But, and this I really don’t get; the fact that Merge could be used to generate “recursive structures” does not imply that in any given G it must so apply. So how exactly does the absence of RWs for recursive rule application in Piraha (note, I am here tentatively conceding that Everett’s factual claims might be right (which is likely incorrect, for they are likely wrong)) show that Merge is not part of a Piraha G? Maybe Piraha Gs can generate unboundedly large phrases but then applies some filters to the outputs to limit what surfaces overtly (this in fact appears to be Everett’s analysis).[5] In this sort of scenario, the Merge rule is recursive and can generate unboundedly large SOs but the interfaces (to use minimalist jargon) filters these out preventing the generated structures from converging. On this scenario, Piraha Gs are like English or Braizilain Portuguese or … Gs, but for the filters.

Now, I am not saying that this is correct. I really don’t know and I leave the relevant discussions to those that do.[6] But, it seems reasonable and if this is indeed what the right G analysis for Piraha is, then it too contains a recursive rule (aka, Merge) though because of the filters it does not generate RWs (i.e. the whole Piraha G does not output “recursive structures”).

Everett’s post rejects this kind of retort. Why? Because such “universals cannot be seen, except by the appropriate theoretician” (4). In other words, they are not surface visible, in contrast to GUs, which are. So, the claim is that unless you have a RW for a rule/operation/process you cannot postulate that rule/operation/process exists within that G. So, absence of positive evidence for (recursive) Merge within Piraha is evidence against (recursive) Merge being part of Piraha G. That’s the argument. The question is why anyone should accept this principle?

More exactly, why in the case of studying Gs should we assume that absence of evidence is evidence of absence, rather than, for example, evidence that more than Merge is involved is yielding the surface patterns attested. This is what we do in any other domain of inquiry. The fact that planes fly does not mean that we throw out gravity. The fact that balls stop rolling does not mean that we dumb inertia, the fact that there are complex living systems does not mean that entropy doesn’t exist. So why should the absence of RWs for recursive Merge in Piraha imply that Piraha does not contain Merge as an operation?

In fact, there is a good argument that it does. It is that many many other Gs have  RWs for recursive Merge (a point that Everett accepts). So, why not assume that Piraha Gs do too? This is surely the simplest conclusion (viz. that Piraha Gs are just like other Gs fundamentally) if it is possible to make this assumption and still “capture” the data that Everett notes.  The argument must be that this kind of reply is somehow illicit. What could license the conclusion that it is?

I can only think of one: that universals just are summaries of surface patterns. If so, then  without  surface patterns that are RWs for recursion in a given G means that there are no recursive rules in that G for there is nothing to “summarize.” All G generalizations must be surface “true.” The assumption is that it is scientifically illicit to postulate some operation/principle/process whose surface effects are “hidden” by other processes. The problem then with considering a theory according to which Merge applies in Piraha but its effects are blocked so that there are no RW-like “recursive structures” to reliably diagnose that it is there is that such an assumption is unscientific!

Note that this methodological principle applied in the real sciences would be considered laughable. Most of 19th century astronomy was dedicated to showing that gravitation regulates planetary motion despite the fact that planets do not appear to move in accord with the inverse square law. The assumption was made that some other mass was present and it was responsible for the deviant appearances. That’s how we discovered Neptune (see here for a great discussion). So unless one is a methodological dualist, there is little reason to accept Everett’s presupposed methodological principle.

It is worth noting that an Empiricist is likely to endorse this kind of methodological principle and be attracted to a Greenberg conception of universals. If universals are just summaries of surface patterns then absent the pattern there is no universal at play.

Importantly, adopting this principle runs against all of modern GG, not just the most recent minimalist bit. Scratch any linguist and s/he will note that you can learn a lot about language A by studying language B. In particular, modern comparative linguistics within GG assumes this as a basic operating principle. It is based on the idea that Gs are largely similar and that what is hard to see within the G of language A might be pretty easy to observe in that of language B. This, for example, is why we often conclude that Irish complementizer agreement tells us something about how WH movement operates in English, despite there being very little (no?) overt evidence for C to C movement in English. Everett’s arguments presuppose that all of this reasoning is fallacious. His is not merely an argument against Merge, but a broadside on virtually all of the cross-linguistic work within GG for the last 30+ years.

Thankfully, the argument is really bad. It either rests on a confusion between GUs and CUs or rests on bogus (dualist) methodological principles. Before ending, however, one more point.

Everett likes to say that a CU conception of universals is unfalsifiable. In particular, that a CU view of universals robs these universals of any “predictive power” (5). But this too is false.  Let’s go back to Piraha.

Say you take recursion to be a property of FL then what would you conclude if you ran into speakers that spoke a language without RWs for that universal? You would conclude that they could learn a language where that universal has clear overt RWs. So, assume (again only for the sake of discussion) that you find a Piraha speaker sans recursive G. Assuming that recursion is part of FL you predict that such speakers could acquire Gs that are clearly recursive. In other words, you would predict that Piraha kids would acquire, to take an example at random, Brazilian Portuguese just like non-Piraha kids do. And, as we know, they do! So, taking recursion to be a property of FL makes a prediction about the kinds of Gs LADs can/do acquire. And these predictions seem to be correct. So, postulating CUs does have empirical consequences and it does make predictions, it’s just that it does not make predictions about whether CUs will be surface visible in every L (i.e. provide RWs in every L) and there is no good reason that they should.

Everett complains in this post that people reject his arguments because they confuse GUs and CUs and that this is incorrect (i.e. they don’t make this confusion). However, it is clear that there is lots of other confusion and lots of methodological dualism and lots of failure to recognize the kinds of “predictions” a CU based understanding of universals does make. Both prongs of the dilemma that the argument against CUs rests on collapse on pretty cursory inspection. There is no there there.

Let me end with one more observation, one I have made before. That recursion is part of FL is not reasonably debatable. What kind of recursion there is, and how it operates is very debatable. The fact is not debatable because it is easy to see its effects all around you. It’s what Chomsky called linguistic productivity (LP). LP, as Chomsky has noted repeatedly, requires that linguistic competence involve knowledge of a G with recursive rules. Moreover, that any child can acquire any language implies that every child comes equipped to the language acquisition task with the capacity to acquire a recursive G. This means that the capacity to acquire a recursive G (i.e. to have an operation like Merge) must be part of every human FL.  This is a near truism and, as Chomsky (and many others, including moi) have endlessly repeated, it is not really contestable. But there is a lot that is contestable. What kind of rules/operations do Gs contain (e.g. FSGs, PSGs, TGs, MPs?)? Are these rules/operations linguistically proprietary (i.e. part of UG or not?)? How do Gs interact with other cognitive systems, etc.? These are all very hard and interesting empirical questions which are and should be vigorously debated (and believe me, they are). The real problem with Everett’s criticism is that it has wasted a lot of time by confusing the trivial issues with the substantive ones. That’s the real problem with the Piraha “debate.” It’s been a complete waste of time.



[1] By this I mean that whereas you are contingently a speaker of English and English is contingently SVO it is biologically necessary that you are equipped with an FL. So, Norbert is only accidentally a speaker of English (and so has a GEnglish) and it is only contingently the case that English is SVO (it could have been SOV as it once was). But it is biologically necessary that I have an FL. In this sense it is more basic.
[2] Actually, they are diagnostic on the assumption that they depict one instance of an unbounded number of sentences of the same type. If one only allows finite substitutions in the S positions then a more modest FSG can do the required work.
[3] Quote from 202 Science paper with Fitch and Hauser.
[4] I henceforth drop the bracketed modifier.
[5] Thanks to Alec Marantz for bringing this to my attention.
[6] This point has already been made By Nevins, Pesestsky and Rodrigues in their excellent paper.

Monday, October 19, 2015

What's in UG (part 3)

Here is the third and final post on the CHY paper (see here and here).

The second CHY argument goes as follows: (i) the clear categorical complementary distribution of BT-anaphors and pronominals that one finds in languages like English is merely a preference in other languages and (ii) ungrammaticality implies categorical unacceptability. In other words, mere preference (i.e. graded acceptability) is a sure indicator that the acceptability difference cannot reflect G structure.[1] This argument form is one that we’ve encountered before (see here) and it is no more compelling here than it was there, or so I will again argue. Let’s go to the videotape for details.

What is the CHY case of interest. It describes two dialects of Malay. In one the categorical judgments found in English are replicated (call this M-1). In the other, the same kinds of sentences evoke preference judgments rather than categorical judgments (call this M-2). The argument is that because Gs only license categorical judgments, M-2’s preferences cannot be explained Gishly. But as M-1 and M-2 are so similar, then whatever account offered for one must extend to the other. Thus, because the account of M-2 cannot be a Gish one, the account of M-1 can’t be either. That’s the argument. Not very good unless one accepts that categorical (un)acceptability is a necessary property of (un)grammaticality. Reject this and the argument goes nowhere. And reject this we should. Here’s why.

The basic judgment data that linguists use involve relative acceptability (usually under an interpretation). Sometimes, the relevant comparison class is obvious and the data is so clean that we can treat the data as categorical (as I argued here, I think that this is not at all uncommon). However, little goes awry if judgment data is glossed in terms of relative acceptability and virtually all the data can be so construed. Now, in these terms, the perception that some judgments (acceptability under an interpretation in this case) are preferable to others is perfectly serviceable. And it may (need not, but may) reflect underlying Gish properties. It will depend on the case at hand.

I mention this because, as noted, CHY describes M-1 and M-2 as making the same distinctions but with M-1 judgments being categorical and M-2 being preferences. CHY concludes that these should be treated in the same way. I agree. However, I do not see how this implies that the distinction is non-grammatical, unless one assumes that preferences cannot reflect underlying grammatical form.  CHY provides no argument for this. It takes it as obvious. I do not.

Is there anything to recommend CHY’s assumption? There is one line of reasoning that I can think of. How is one to explain gradient acceptability (aka preference) if one takes grammaticality to be categorical?  This is the question extensively discussed here (see the discussion thread in particular) and here. The problem in the domain of island phenomena is that even when we find the diagnostic marks of islandhood (super additivity effects) and we conclude that there is “subliminal” ungrammaticality, we are left asking why for speakers of one G the effects of ungrammaticality manifest themselves in stronger unacceptability judgments than for speakers of another G. In other words, why if island violations are ungrammatical do some find them relatively acceptable? The same question arises in the case that binding data that CHY discusses. And in both instances the question raised is a good one. What kind of answer should/might we expect?

Here’s a proposal: we should expect ungrammaticality ceteris paribus to get reflected in categorical judgments of unacceptability. However, ceteri are seldom paribused.  We know that lots goes into an acceptability judgment and it is hard to keep all things equal.  So for example, it is not inconceivable that sentences with many probable parses are more demanding performance wise than those without. More concretely, imagine a sentence where the visible functional surface vocabulary (FSV) fails to make clear what the underlying structure is. I am assuming, as is standard, that functional vocabulary can be a guide to underlying form (hence the relative “acceptability” of Jabberwocky). Say that in some languages the underlying surface morphology is more closely correlated to the underlying syntactic categories than in others. And say that this creates problems mapping from the utterance to the underlying G form. And say that this manifests itself in muddier (un)acceptability judgments. To say this another way; the less ambiguous the mapping from surface forms to underlying forms the more categorical the judgment will be. If something like this is right, then if we find a language where the morphology does not disambiguate BT-anaphors from exempt anaphors then we might expect acceptability to be less than categorical. Think of it as the acceptability judgment averaging over the two G possibilities (maybe a weighted average). On this scenario, then the absence of a “dedicated reflexive” form (see CHY p. 9) in M-2 will make it harder to apply BT than in a language where there is a dedicated form, as in M-1. Note, this is consistent with the assumption that in both languages the G distinguishes well-formed forms from ill-formed forms. However, it is harder to “see” this in M-2 given the obscurity of the surface FSVs than it is in M-1 where the distinction has been “grammaticalized.”[2]

I mention this option for it is consistent with everything that CHY discusses and, as I hope is evident, it leaves the question of the UG status of BT untouched. In short, in this particular case it is easy enough to cook up an explanation for why binding judgments in M-2 are murkier than those in M-1 without assuming that both reflect the operations of a common UG.[3] Thus, the CHY conclusion is not only based on a debatable premise, but in this particular case there is a pretty obvious way of explaining why the two dialects might provide different acceptability judgments. I should also add that this little story I’ve provided is more than CHY does. Here’s what I mean.

Curiously CHY does not explain how M-1 and M-2 are related except to say that M-1 has grammaticalized a distinction that M-2 has not. Which? M-1 has grammaticalized the notion notion of anaphor. What kind of process is “gramamticalizing”?  CHY does not say. It does not provide an account of what grammaticalization actually is, it only points to some of its effects in Malay and suggests that this goes on in creolization. Nor does CHY explain how undergoing grammataicalization renders preferences in the pre-grammaticalization period categorical in the post grammaticalization period? In CHY “grammaticalization” is Voltarian.[4] Let me offer a proposal of what gramamticalization is (actually this is implicit in CHY’s discussion).

Here’s one proposal: grammaticalization involves sharpening the FSV so that it more directly reflects the underlying G structure. In other words, grammaticalization is a process that aligns surface functional vocabulary with underlying grammatical forms. It might even be the case that language change is driven to sharpen this alignment (though I doubt that the force is very strong (personal opinion) Why? Because FSVs muddies overtime as well as sharpens and lots of FSV is very misleading). But if this is what grammaticalization is, then it can hardly challenge the UG nature of BT as it presupposes it. Grammaticalization is the process whereby the underlying categories of FL/UG act as attractors for overt functional morphology (i.e. LADs try to treat visible functional as reflecting underlying G categories and so over time the surface functional vocabulary will come to (more) perfectly delineate UG cleavages). In fact on this view, CHY, inadvertently, argues FOR the UG nature of BT for it assumes that grammaticalization is the operative process linking M-1 and M-2.

As noted, CHY does not explain what grammaticalization is (nor, to my knowledge has anybody else), though it does note what drives it. It is the usual suspect in such cases; the facilitation of processing (p. 17).  Unfortunately, even were this so (and I am skeptical that this actually means anything), it leaves unexplained how languages like M-2 could exist. After all, if processing ease is a good thing, then why should only M-1 partake?  The answer must be that something stops it from enjoying the fruits of parsing efficiency. What might this be? Well, how about the fact that the PLD only murkily maps the binding relevant FSVs (i.e. the surface forms of the anaphoric morphemes) onto the relevant underlying grammatical categories. But, as noted, if this is what grammaticalization is and what it does, then it is not merely compatible with the view that BT is part of UG and that it is innate, but virtually presupposes that something like this must be the case. Attractors cannot attract without existing.

Let me end here, with a diagnosis of what I take the fundamental error that drives the CHY discussion to be. It is not a new mistake, but one that is, sadly, endemic. It rests on the confusion between Greenberg and Chomsky Universals. CHY assumes that BT aims to catalogue surface distribution of overt morphemes. On this construal, BT is indeed not universal (as even a rabid nativist like me would concede). It is clearly not the case that languages all distinguish overt morphological categories subject to different BT principles. Some languages don’t clearly have a demarcated distinction between overt anaphors or pronominals among their FSVs, some don’t even have dedicated overt functional forms for reflexivization or pronominalization. If one understands UG as committing hostages to surface functional morphology, then CHY is right that BT is not universal. However, this is not how GGers ever understood (or, more accurately, ever should have understood) UG and universals. Chomsky universals are not Greenberg universals. They are more abstract and can be hard to discern from the surface (btw, this is what makes them interesting, IMO). Thus criticizing BT because it is wrong when understood in Greenbergian terms is not much of a fault given that it was not supposed to be so understood (i.e. another Dan Everett moment (see here)).  What is surprising is that the distinction between the two kinds of universals seems so difficult for linguists to grasp. Why is this?

Here’s an unfair (though I believe close to accurate) speculation: it results from the confluence of two powerful factors (i) the attraction of Empiricist conceptions of learning and (ii) the fascination with language diversity.

The first is a horse that I have hobbied on many times before. If you think that acquisition is largely inductive then universals without clear surface reflexes are a challenging concept. Being Eish with a taste for universals leads one to naturally erroneously understand Chomsky universals as Greenbrg universals (as Greenberg universals are the only ones that Eism tolerates).

The second force leading to the confusion between Greenberg and Chomsky universals comes from a fascination with linguistic variation (clearly something that is at the center of CHY). FL/UG rests on the idea that underlyingly there is very little real G variation. If one’s interest is in variation, then this notion of UG will seem way off track. Just look at all the differences! To be told that this just surface morphology will seem unhelpful at best and hostile at worst. The natural response is to look for helpful typological universals and these, not surprisingly will be Greenbergian. Here the generalizations concern surface patterns, as do typological differences if Chomsky’s conception of FL/UG is on the right track. Typological interests do not require embracing a Greenberg conception of universals (unlike a commitment to Eism, which does). However, it is, I believe, a constant temptation. The fact is that a Chomskyan conception of UG is consistent with the view that there are very few (if any) robust typological (i.e. surface true) universals. UG in Chomsky's sense doesn’t need them. It just needs a way of mapping overt forms to underlying forms. In other words, UG needs to be coupled with a theory of acquisition, but this theory does not require that there be surface true universals. Of course, there may be some but they are not conceptually required.

So that’s it. CHY’s conclusions only follow from a flawed understanding of what a universal is and what UG enjoins. The argument is not very good. Sadly, it might well be influential, which is why I spent so much effort trying to dismember it. It appeared in an influential cog sci journal. It will be read as undermining the notion of UG and the relevance of PoS reasoning. It will do so not because the arguments are sound but because this is a welcome conclusion to many. I strongly suggest that GGers educate their psycho counterparts and explain to them why Cognition has once again failed to understand what Chomskyan linguistics is all about. I also suggest that understanding a PoS argument be placed at the center of the field’s pedagogical concerns. It really helps to know how to construct one.


[1] I have no idea why this assumption is so robust among linguists. I don’t believe that anyone ever argued the case and many argued that it was not. So, for example, Chomsky explicitly denies this that one could operationalize grammaticality in terms of (categorical, or otherwise) judgments of acceptability (see, e.g. Current Issues: 7-9 and chapter 3). In fact, there is little reason to believe that there can be operational criteria of FL/UG notions of grammaticality, as holds true for any interesting abstract scientific notions (See Current Issues: 56-7).  If this is correct, then systematic preference judgments might be just as revealing of underlying grammatical form as categorical judgments. At any rate, the assumption that preferences exclusively reflect extra grammatical factors is tendentious. It really depends.
[2] I return on a moment to explicate this term.
[3] It is actually harder to come up with a good story for variable island effects, though as I’ve mentioned before I believe that Kush’s ambiguity hypothesis is likely on the right track.
[4] As in: why does this morphine put you to sleep? In virtue of its dormitive powers. 
I should add that CHY needs to offer an account of what the process consists in at pains of undermining its main argument. The argument is that M-1 and M-2 are sensitive to the same distinction. But this distinction cannot be a Gish one because Gish ones result on categorical judgments. But the M-2 judgments are not categorical therefore the distinction cannot be grammatical. How then to explain m-1 judgments? Well they are categorical because the non G distinction in M-2 has been grammaticalized in M-1. This invites the obvious question: what’s the output of the process of grammaticalization? It sounds like the end product is to render the distinction a grammatical one. But if this is so, then the premise that the distinction is the same in both M-1 and M-2 fails for what is a non-grammatical distinction in M-2 is a grammatical distinction in M-1. The only way to explicate what is going on in the argument is to specify what grammaticalization is and what it does. CHY does not do this.