Comments

Tuesday, October 23, 2012

Empiricism, Rationalism and Generative Grammar



 Here is my first rule of research:

 Those things not worth doing are not worth doing well. 

So, qua generative linguist, what’s worth doing? We get a handle on this by asking what are the central questions driving the Generative Enterprise? Two immediately come to mind: (1) What’s in UG and (2) Why is it there?  Taking these two queries as dispositive identifies the object of inquiry as the structure of UG.  The point of generative research is (or should be) to limn its fine structure and research is (or should be) evaluated by whether it helps us achieve this end.

This is a relatively focused conception of the goal(s) of Generative Grammar and it relegates many kinds of inquiry (e.g. what’s the structure of the Japanese DP?, how do kids use language to order chicken fingers in restaurants? Does language L allow multiple case checking? Are relative clauses islands in Swedish?, Are bound pronouns spelled out traces, etc.) to (at most) a subordinate status.  So why do I focus on (1) and (2)? Here’s one reason why.

I grew up in a philosophy department and learned about Chomsky’s work in linguistics through a series of philo debates in the early to mid 70s, most especially  “the innateness controversy.” In those days, linguistics was cutting edge as it was the major battle-ground on which two great philosophical traditions -Empiricism and Rationalism- met and disputed.  For me and my friends, Chomsky stood toe to toe with Plato, Descartes, Leibinz, and Kant and led the charge against the forces of darkness viz. Skinner and Quine, heading the party of Aristotle, Locke and Hume. Unlike Vegas, what happened in linguistics did not stay there, it leached out into the wider intellectual world and had big consequences, or so it felt to us (especially over beers, on Friday nights, in downtown Montreal, in our early 20s). The stakes were high, nothing less than the nature of mind and its relation to the external world.  Linguistics was the leading edge of the cognitive revolution, the best case against the empiricist conception of mind and a model for the emerging cognitive sciences.  How did linguistics manage this?  Here’s a potted reconstruction.

Empiricism, a species of environmentalism (natural selection being another), holds that minds are structured by the environments in which they are situated. The leading metaphor is the mind as soft perfectly receptive wax tablet (or empty cupboard) which the external world shapes (or fills) via sensory input.  The leading slogan, borrowed from the medievals, is “nothing in the intellect that is not first in the senses.”  The mind, at its best, faithfully records the external world’s patterns through the windows of sensation.

Rationalists have a different animating picture. Leibniz, for example, opposed the wax tablet metaphor with another: ideas are in the mind in the way that a figure is implicit in the veins of a piece of marble.  The sculptor cuts along the marble’s grain to reveal the figures that are inchoately there.  In this picture, the environment is the sculptor, the veined marble the mind.  The image highlights two main differences with the empiricist picture. First, minds come to environments structured. They have a natural grain, allowing some figures (ideas) to easily emerge while preventing or slowing the realization of others. Second, whereas a hot wax imprint of an object mirrors the contours of the imprinting object, there is no resemblance between the whacks of the chisel and the forms that such whackings bring to life.  Rationalists allow minds to represent external reality but deny that they do so in virtue of some sort of similarity obtaining between the sensory perceptions and the ideas they prompt. Thus, whereas Rationalists postulated causal connections between mental content and environmental input they denied that environments shape those contents.  The distinction between triggering and shaping was an important one.

Associationism is the modern avatar of empiricism.  The technology is more sophisticated, neural nets and stimulus-response schedules replacing wax tablets and empty cupboards, but the guiding intuition is the same. Minds are pattern matchers able with sufficient exposure to the patterns around them to tune themselves to the patterns impinging on them. What made Chomsky’s ideas about Generative Grammar so exciting was that they showed that this empiricist picture could not be right.  To account for a native speaker’s linguistic competence requires that humans come equipped with special purpose mental procedures and this is inconsistent with empiricisms associationist psychology.  Two features of linguistic competence were of particular importance: first that the competence emerges relatively rapidly, without the learning being guided and despite data that is far from perfect. Second, much of what speakers know about their language is not attested at all in the data they have access to and use.  No data, no possible associationist route to the mind. Ergo: the mind must be structured. Point to the Rationalists. 

In retrospect, it is hard to see why we were so surprised and animated by Chomsky’s arguments.  Indeed, considered naively the idea that humans come equipped with a species specific dedicated linguistic capacity is the ‘duh’-position.  Based on simple observations it’s clear that nothing else learns language as we do. Indeed, nothing else comes anywhere close. Only a sophisticate could conclude otherwise.[1] However, so widespread was the empiricist perspective that the cognoscenti took it to be simple common sense. Hence, it required a sustained frontal attack to displace it.  That’s why watching Chomsky topple empiricism was so exciting and why work in generative grammar reached beyond linguistics to influence thinking in cognitive science, philosophy and computer science as well.

In sum what made linguistics exciting (and still makes it exciting) is that it provides an easy way to plug into a very great long-lived debate about the structure of the mind. All you need to do to participate is the following: next time you read or hear a paper, ask yourself (or the lecturer) “what does this tell us about the structure of UG?”











[1] Come to think of it, maybe this is what made empiricism so attractive to the learned; it challenges the obvious and so has the sheen of scientific sophistication? 

Thursday, October 18, 2012

‘I’ before ‘E’: unambiguity


In my next post, I’ll discuss the I-language/E-language distinction that Chomsky introduced in Knowledge of Language. The history of this distinction is illuminating, and it helps explain what the ‘I-’ means. But for today, let me stipulate that an I-language is a generative procedure that connects articulations of some kind (say, sounds or gestures) with meanings of some kind. Let ‘E-language’ be a covering term for anything else—a set of word strings, a social practice, or whatever—that might be called a language.

That’s already enough to make it clear that many alleged “debates” about whether kids learn the languages they acquire (and whether such languages are transformational) aren’t really debates. One “side” argues that humans naturally acquire I-languages that connect articulations with meanings in accord with logically contingent constraints that are not learned. The other “side” shows how a certain kind of learner could acquire an E-language that is like a human I-language in respects other than the ones highlighted by the evidence.

To take a much discussed kind of example, the word-strings indicated with (1-3)
(1)  The guest was fed waffles?
(2)  The guest fed the parking meter?
(3)  The guest who was fed waffles fed the parking meter?
can be used—with rising intonation—to ask yes/no questions, with the corresponding declaratives indicating affirmative answers. With regard to (1), the question can also be asked with (4). But (5) is not another way of asking question (3).
                        (4)  Was the guest fed waffles?
(5)  Was the guest who fed waffles fed the parking meter?
On the contrary, (5) is understood as the bizarre question indicated with (6).
(6)  The guest who fed waffles was fed the parking meter?
Put another way, (5) is unambiguous: it has the meaning indicated with (6), and it fails to have the meaning indicated (3). This “negative” fact is of interest. One can easily imagine a generative procedure that connects the pronunciation of (5) with both meanings, or just the meaning of (3). But (5) can only be understood as the bizarre question, even though (3) is the more likely question, given what we know about guests and waffles. Similarly, (7)
                        (7)  Was the hiker who lost kept walking in circles?
is understood as indicated with (7a) and not (7b).
(7a)  The hiker who lost was kept walking in circles?
(7b)  The hiker who was lost kept walking in circles?
So in acquiring English, one acquires an I-language that connects articulations with meanings in a way that makes (5) and (7) unambiguous.
           
Now perhaps kids somehow learn that their parents and peers use I-languages that
connect articulations with meanings in this constrained way. I doubt it, for reasons that have been reviewed often. (I’ve done my time on such reviews.) But in principle, I can imagine replies of the following form: show how kids could start with a more permissive generative procedure—or a strategy for acquiring I-languages that would support acquisition of I-languages in which (5) is ambiguous—and then use available experience to figure out that the “local” I-languages are more constrained. I have not, however, encountered such replies. What I have encountered (see the reviews just mentioned) are descriptions of machines that can learn to classify strings like (8) as defective, while classifying strings like (9-11) as undefective.
                        (8)  Was the guest who hungry was tired?
                        (9)   The guest was hungry?
                        (10)  Was the guest hungry?
                        (11)  Was the guest who was hungry tired?
Such machines can, in effect, learn to put an asterisk on (8) while leaving (9-11) unmarked.

But that’s beside the point. The phenomenon illustrated with (1-7) is not that kids acquire hard to learn procedures for putting asterisks on strings. The point is that kids acquire I-languages (procedures that connect articulations with meanings) that are constrained in certain ways. Of course, any biologically implemented procedure will be constrained in ways that are unlearned. But the interesting nativist claim—not rebutted by inventing learnable procedures for classifying strings as defective—is that particular constraints (e.g., those characterized in terms of constraints on displacement) are unlearned.

Linguists, or their informants, might mark the oddness of (12) as shown below.
(12)  *The guest who fed waffles was fed the parking meter.
But like (7a), (12) can be understood as an English sentence that expresses a crazy thought. Another day, I’ll talk about contrasts with (13) and (14).
(13)  *Colorless green ideas sleep furiously.
                        (14)  *I might been have there.
There are complications, and not only because acceptability differs from grammaticality. But whatever we say about asterisks, a child who acquires English acquires an I-language that connects the pronunciation of (12) with the corresponding meaning. Such a child also ends up knowing that (12) is bizarre thing to say. But it’s bizarre because of what (12) means. Likewise for (5). And (5) wouldn’t be bizarre if it could have the meaning of (3).
(5)  Was the guest who fed waffles fed the parking meter?
(3)  The guest who was fed waffles fed the parking meter?
That raises the question of why of kids don’t acquire I-languages that are more semantically permissive. Building machines that can learn to put an asterisk on (5) doesn’t address this question, much less suggest that kids learn that strings like (5) are unambiguous. One can try to build a machine that classifies (5) as a “generable but deviant” string and (14) as “ungenerable.” But the question remains: why does (5) have one meaning rather than two? Similar remarks apply to (15), which has the meaning of (16) and not (17).
                        (15)  Can pigs that fly talk?
                        (16)  Pigs that fly can talk?                         
(17)  Pigs that can fly talk?
If we want to understand the human capacity to acquire I-languages—procedures that connect articulations with meanings in certain ways—then it’s hard to see the point of inventing machines that learn to classify strings as generable or not. There are, of course, other goals. But to have a debate about I-languages, both sides have to talk about them. 

‘I’ before ‘E’: hello


The management has kindly invited me to post, from time to time, on matters related to the faculty of language. Complaining has been encouraged. But where to start?

I do think that the I-language/E-language distinction is insufficiently appreciated. OK, that’s understatement. I think that failure to appreciate this distinction fosters many diseases that currently plague the field: a tendency to ignore the strongest evidence against empiricist conceptions of language acquisition; confusion about the data adduced in “poverty of stimulus” arguments; related confusion about what the asterisk means in examples like ‘*I might been have there’; extensional conceptions of meaning; the practice of representing intensions—and worse, intentions—with sets of possible worlds; misguided conceptions of how grammatical competence is related to comprehension; misunderstandings of how “algorithmic” levels of description (as in Marr-style theories of vision) are related to “functional” and “implementational” levels. The list could, and probably will, go on.

Eventually, I may get around to complaining about something other than failures to appreciate the I-language/E-language distinction. But at least for a while, I’ll focus on one big idea that many people profess to accept: when kids acquire a language, they don’t simply acquire an infinite set of word strings, whatever that would mean; rather, each kid acquires at least one generative procedure that somehow connects (boundlessly many) meanings of some kind with (boundlessly many) articulations of some kind.

In saying that this is a big idea, I don’t mean that it is surprising, much less that it ought to be controversial. On the contrary, I think that in retrospect, it ought to seem nearly truistic. But sometimes, a near truism can be theoretically fruitful by drawing attention to phenomena that call for explanation, and suggesting a useful conception of the basic target(s) of inquiry. (Think about the claim that heritable variation in fitness leads to evolution, and its relation to the bolder idea that all life on earth descended from a common source.) It may be obvious that in acquiring a language, a child acquires a procedure that somehow generates articulation-meaning pairs. But the literature suggests that the implications of this obvious point have not been absorbed. Or so I’ll be saying, more than twice, in the weeks ahead.

Tuesday, October 16, 2012

How to Play the Game


Imagine the following not uncommon scenario: Theory T claims that explaining the particulars of phenomenon P requires assumptions A1….An. Someone, S, doesn’t like one or all of these assumptions for a variety of reasons (never discount the causal efficacy of dyspepsia) and decides that T is wrong.  Here’s the question: what is S obliged to do? Answer: It depends.

There is no moral, religious, legal, or social obligation that S do anything at all. You can think anything you want, and say anything you feel like saying (as my daughter used to say: “Nobody is the boss of me!”).  But, if you want to play the explanation game, the “science” game, then you are obliged to do more, a lot more.  You are obliged to explain why you think the assumptions are faulty and (usually, though there are some exceptions) you are obliged to offer an (at least sketchy) non-trivial question begging account of P.  S cannot simply note that s/he thinks that T is really really wrong, or that T is unappealing and makes her/him feel ill, or that s/he wished T were wrong for some unspecified, no doubt, humanitarian reason.  Doing this is just deciding not to play the game.  Sadly, many critics of Generative Grammar have decided that they don’t like the theory, passionately articulate their dissent but do not follow the rules.  To repeat: nobody needs to play, but unless you follow the rules nobody should take your views (prejudices?) particularly seriously. Adherence to the rules of the game is the price of admission to the discussion.

You’ve probably guessed why I mention this. A lot of people want their views to be taken seriously despite not playing by the rules. For some odd reason they think they should be exempt because it’s just so clear that they are right and generative grammarians (especially Chomsky and his intoxicated minions) are wrong.  But, though this might sustain a warm glow of amour propre it does not admit you to the game.  Lest you think that I have descended into caricature, consider a recent short paper by David Adger - Constructionsand grammatical explanation – that does the heavy lifting exposing how far from serious certain well-known forms of construction grammar are. Those interested in the agnotology of science will enjoy this spirited well-aimed take-down. Here’s a teaser quote to whet your appetite:

…CxG [Construction Grammar, NH] proponents have to provide a theory of how learning takes place so as to give rise to a constructional hierarchy, but even book length studies on this, such as Tomasello (2003), provide no theory beyond analogy combined with vague pragmatic principles.

Let me leave you with a simple piece of advice that has served me well: when you hear the word ‘analogy’ reach for your wallet.

Monday, October 15, 2012

The Chomsky Problem?



I am an unabashed and irritatingly vocal admirer of Chomsky’s many intellectual contributions. I read everything he writes and has written in Linguistics and Philosophy (ok, almost everything: I came to linguistics from philosophy so I have refrained from dipping into the pleasures of SPE and a few of Chomsky’s other phonologically targeted products) and a very large chunk of his political work.  So it came to me as quite a surprise to find out about the “Chomsky Problem” in the August 29th 2012 issue of the TLS.  I have run into many “problems” over the years: The “Maria Problem” in the Sound of Music (singing novitiates can be bothersome), the “three body problem” (something you want to avoid if you value simple calculations), the “Mind-Body Problem,” the “Problem of Other Minds,” and the “Problem of Induction,” (these have occupied philosophers for a long time and will no doubt be hot topics for a while still), Das Adam Smith Probleme (Germans puzzle over how one guy could have written two books that appear to be quite complementary (who can guess what puzzles German intellectuals!), Theory of Moral Sentiments and The Wealth of Nations) among others.  But never the “Chomsky Problem.” David Hawkes explains it as follows: Chomsky’s writings in theoretical linguistics and his political commentary “appear to contradict each other.” What’s the contradiction? Hawkes believes (and he suggests that he is not alone) that (1) and (2) cannot both be coherently entertained.  

(1)           UG is a feature of human brains and is hard wired into our genes
(2)           Conservative forms of social organization are neither immutable nor natural

The problem seems to be that (1) “can be easily characterized as reactionary” because it “diminishes the influence of the environment on human behavior,” which apparently implies that that those forms of social organization that do exist must exist as a matter of biology. It doesn’t take a lot of effort to see that whatever the relation between (1) and (2) might be, “contradiction” is not one of them. The truth of (1) has no implications whatsoever regarding that of (2), a position that Chomsky adopts, as Hawkes observes.

Independent of (1)’s relation to (2), Hawkes’ claims concerning (1) are very confused.  He seems to identify genetic coding with immutability and immunity from environmental influence (viz. puberty is genetically programmed but environmental factors can accelerate or delay its onset) . However, very few genetically determined characteristics are so isolated from all environmental impact (think diet and height). Certainly, as far as language goes, even if UG is genetically coded, which particular language a speaker acquires is acutely sensitive to her linguistic environment. There is no plausible logical route from the fact that all languages have certain formal structural similarities to the conclusion that they are all identical in every particular, or, even more of a stretch, that because all languages have a common form, anything at all follows about social organization. Indeed, even if every particular about a given language was coded in the genes (a position that nobody entertains), it is hard to see what this would imply for the large social and political concerns Hawkes is worried about. After all, face recognition and pitch perception have genetic components but neither Hawkes nor anyone else has suggested that this has any political implications.

I conclude that Hawkes cannot literally mean what he says. Rather, Hawkes is tempting us with the following slippery slope inference: If any feature of human behavior has an “immutable” biological basis, every one does.  As Hawkes sees it there is an “affinity” which inclines those that adopt (1) to embrace conservative forms of political and social organization. Is there any truth to this? 

Not on the face of it. As he notes, Chomsky himself rejects any connection.  But let’s for a moment take the supposition seriously, for there is something decidedly odd about Hawkes’ views for they seem to assume that non-conservative forms of social organization require that we assume that human nature has no genetic roots.  And this is very problematic.  Why so?

First, it is unlikely to be true.  Daily, we find evidence that many of our cognitive and affective characteristics are elaborated on biologically given foundations.  It’s a bad idea to tie opposition to oppressive forms of social and political organization on the extreme view that biology has nothing to do with human nature.

Second, it is unnecessary to take such an extreme position. Imagine (as some have proposed) that humans come equipped with a kind of UG for ethics and morality (Rawls, Kant) or have a natural communal instinct (Aristotle, Bukharin) or a built in capacity for sympathy (Hutchinson). On this view, some forms of moral judgment are more “natural” than others, at least for humans. What follows from this? As regards what we ought to do, not much, if one distinguishes (as I do, at least for rough and ready purposes) what is the case from what should be the case, i.e. facts from values. At most what these kinds of considerations talk to is the feasibility of one or another form of social organization given the human propensities on which they can be founded. For example, social fraternity (we are all in this together) can live on the natural sociability of humans and the sympathy the feel for one another’s circumstances. Similarly the concern for justice can exploit our sense of sympathy (Hume) or our common faculty of reason (Kant). 

Third, nothing (absolutely nothing) that we know from biology, psychology or neuroscience precludes any of the forms of social organization that Hawkes or Chomsky or me cares about. This does not mean to say that there don’t exist good reasons for preferring some social arrangements to others. Rather, the relevant reasons for one preference or another have little to do with discoveries in biology, psychology or neuroscience. The fact is that one can learn more about human nature that is polictically or socially relevant from a good novel (or even a bad one) than from the “deep” insights science has provided (hint: one should always treat NYT headlines in the Tuesday science section with considerable skepticism). No doubt many bad arguments have been provided to buttress one or another execrable policy (think back to the IQ debates of yore or the speculations concerning the genetic inability of females to do math). But bad arguments of this kind are not bad because they are based in biology, they are bad because they are bad science, bad philosophy and bad public policy and should be opposed as such.

Fourth, the anything-can-be-human-nature view is also susceptible to abuse.  After all if human nature is so malleable that it can tolerate any form of social and political organization why prefer one to another? Skinner, a very pure environmentalist, argued in Beyond Freedom and Dignity that freedom and dignity were illusions that stood in the way of a utopia based on behavior modification implemented by wise psychologist kings. There is more than enough fodder here to argue that extreme environmentalism has nasty consequences for democracy. Or, put positively, the belief that humans as such share certain basic cognitive and affective features (i.e. share them just in virtue of being humans) can ground (and has grounded) ideals of equality, freedom and justice with important social and political implications for how societies ought to be organized. An old enlightenment theme that keeps cropping up in Chomsky’s political writings concerns each human’s creative potential (think Kant, Rousseau, von Humboldt), an everyday manifestation of which is the creative use of language. Thinking of people in this way suggests forms of social organization where this creative potential is not stifled and is allowed to flower.  This doesn’t follow logically, but as Hawkes might say, there is an affinity here.  An old prof of mine, Harry Bracken, once pointed out that rationalist conceptions of human nature provide “mild conceptual brakes” against invidiously distinguishing among humans (aka: racism) and provide a basis for treating all equitably and with common respect.  As Bracken noted, this is not an apodictic truth, but such conceptions of human nature might provide a slight nudge in a decent direction.

So, contra Hawkes, Chomsky is certainly right that nativist conceptions of language are logically independent of questions of social and political organization.  Indeed, science has yet to uncover a conception of human nature capable of having any interesting moral or political consequences, nor do I believe that such discoveries are on the horizon.  In sum, right now the general form of the “Chomsky Problem” (what does science tell us about how to organize society) is not worth answering for it rests on the faulty presupposition that science has something interesting and original to add to the conversation. Science cannot tell us that we ought to treat one another decently, and anyone who needs science to “confirm” this moral precept won’t believe the evidence anyway.

Thursday, October 11, 2012

The Gallistel-King Conjecture



Ever since Randy Gallistel came to UMD to give a series of lectures that eventually became his book (with Adam King) Memory and the Computational Brain Bill Idsardi and I have been discussing his deliberately provocative thesis that current neuroscience has fundamentally misidentified basic brain architecture. The book argues that the current conception of the brain in which learning is understood as rewiring of a plastic brain via changes of synaptic conductance (what fires together wires together) cannot be correct (connectionism is an expression of this neuronal worldview). The argument is elaborate and I will leave a more detailed discussion of its main argument for another post. However, what I want to very briefly mention here is a Gallistel and King (G&K) conjecture and some recent work that appears to relate to it.  I say “appears” because, believe me, I am no expert (in fact I am not even knowledgeable enough to be a novice) and so readers should take what follows as essentially an impressionistic riff based on chapter 16 of G&K’s book and a recent broadcast on Science Friday; What Your Genes Can Tell You About YourMemory. So with this caveat lector firmly before you, my conscience is clear and my riff begins.

What is G&K’s conjecture?  They argue in the book that brains must have the architecture of classical computers. Concretely this means that brains must be able to put things in memory, retrieve things from memory and compute over those things stored in memory.  As computation requires applying functions to arguments, we need a way of coding variables and operations that bind and value them.  These are all operations characteristic of a classical machine with a Turing-von Neumann (TvN) architecture.  As most of cognition consists of operations that redeem symbols from memory, computations over these symbols and subsequent storage, the brain, the organ that secretes cognition, must have a TvN architecture.  This is the conclusion. The argument is detailed, very pedagogical and a must read. 

G&K contrast the TvM conception with the currently common view of brains. In this view, brains do not have a TvM structure but are more like neural nets, which, G&K argue is an inadequate physical basis for cognitive computations and so must be wrong. This is very much a minority view. Why? Because brains don’t “look like” computers and they do look like neural nets. Ok, it’s probably more complicated than that, but certainly part of it as anyone who has had the misfortune to hear a connectionist talk knows. So if G&K are right that the common view is wrong and the brain has TvM structure then how does the brain do this? The G&K conjecture is that the requisite structure already exists within neurons, in its molecular structure (169):

…the genome contains complex data structures, just as does the memory of a computer, and they are encoded in both cases through the use of the same architecture employed in the same way: an addressable memory in which many of the memories addressed themselves generate probes for addresses. [There is a] close parallel between the functional structure of computer memory and the functional structure of the molecular machinery that carries inherited information forward in time for use in the construction and maintenance of organic structure…

Their conjecture is that the physical platform for the kinds of computational mechanisms we need to understand cognition exploits this same molecular structure.  DNA, RNA and proteins constitute (part of) the cognitive code in addition tos being the chemical realization of the genetic code.

G&K are very careful to moot this suggestion with all the caution that it deserves. In fact to call it a ‘suggestion’ is already too grandiose, let’s say a hunch or a guess.  This is where the Science Friday segment comes in. It appears that research is discovering that memories are in fact molecularly coded.  There are epigenetic mechanisms that code specific memories in various brain regions by laying down the proteins of the right kind using basically the same DNA mechanisms that code for genetic inheritance and development.  Combined with the G&K conjecture, this discovery might be the tip of a pretty exciting cognitive-neuroscience iceberg and has the potential of overturning a good deal of conventional wisdom, as the scientists interviewed hint at.

This has serious implications for linguists, if true (and recall this is all an impressionistic riff). There exists a very bad argument that the representations that linguists know and love cannot be “psychologically real” because they are not implementable in brain architecture, i.e. neural nets.  In my view, this has always been a weak argument, but one that seems to have quite a bit of suasive power to non-generative grammarians.  If the G&K conjecture is on the right track, however, there is no reason to think that brains cannot code for the kinds of representations we regularly use to account for grammatical competence.  The question will shift from whether they do to how they do.

Let’s end with a little parable based on some G&K remarks (p.281). They make an interesting observation about the history of modern biochemistry and consider its implications for current neuroscience. Watson and Crick’s (W&C) great accomplishment was to find a way to chemically incarnate the classical gene. Until they did their work, the gene was considered a nice computing device, but many biologists believed that it was not “biologically real.”  W&C proved otherwise and this entirely changed biochemistry. The field after W&C was entirely different from the field before W&C.  But what of classical genetics, how much did it change? In contrast to biochemistry, the basics remained essentially as they were before W&C.  Biochemistry had to “catch up” to genetics, not the other way around.  This story has a moral: substitute neuroscience for biochemistry and cognition/linguistics for genetics.  There is no reason a priori to think that “hard” (and expensive) neuroscience occupies the intellectual high ground to which “soft” (and cheap) cognition/linguistics must accommodate itself. Matters might well be the reverse, as they were once before when “hard” biochemistry ended up conforming to “soft” genetics.  The discoveries discussed on Science Friday make the G&K conjecture a little less farfetched and tentatively suggests that the analogy with biochemistry/genetics is prophetic. We may be getting ready to say bye-bye to all those inane connectionist models. Yeah!!!

Friday, October 5, 2012

Three psychologists walk into a bar…



Linguists should applaud recent efforts by the Royal Society to inject humor into research on language.  Clearly, the Royals believe that linguists have taken their work far too seriously and in a welcome reversal of a previous policy of benign neglect, they have published a near pitch perfect parody (“How Hierarchical is language use”) by the trio of Frank, Bod and Christiansen (hereafter FBC). These authors breezily describe a research program aimed at unseating a fixed point in research on natural language that has been virtually uncontested for the last several hundred years: that sentences and phrases have both linear and hierarchical dimensions. Like all good parodists, they are undaunted by the obvious; in this case the fact that anyone who has ever examined language has concluded that words combine into phrases that combine into sentences. No, this dynamic trio, in the grand tradition of J. Swift and W. Allen, suggests that all we need are bi-grams and tri-grams.  Using the powerful methods of neuroscience, computational linguistics and behavioral psychology they propose that it’s all just one damn word after another, no hierarchy needed. Moreover, FBC do this without ever letting down their parodic cloak.  Indeed, the paper is so successfully crafted (it has such a vivid sense of seriousness) that as a public service I believe it is necessary to affix to it the right warning label just in case those with arthritic funny bones are misled, take this exquisite lampoon too seriously, and thereby wander into an intellectual desert.

The paper, like all good parody has a central rhetorical thread.  The first strand is that linguistic hierarchy poses a problem for evolution due to its biological sui genericity.  The second, as FBC note, is that linguistic use cares about sentential sequence. The third combines the two to conclude that because linguistic hierarchy is evolutionarily problematic (in contrast to sequence information) it couldn’t possibly have evolved (i.e. arisen in humans) and so linguistic systems can’t really have any.  Conclusion; linguistic use is sensitive exclusively to sequence information, as it must be. The satire requires that you inadvertently slip to the conclusion that if language use exploits linear properties of sentences and phrases then hierarchy is dispensable for all linguistic analysis (after all: if you can hop on one leg who needs two legs?). Occam and his razor are invoked to make sure that after you slip to their desired conclusion you don’t jump back up incredulous.  It’s all neatly done and very amusing.

Let’s pull the conceit apart to better admire its artistry. FBC know that linguists from the thirteenth century grammarians, through Bloomfield and Harris in the mid 1950s to Chomsky today have all taken it as obvious that natural language grammars are hierarchically organized, labeled brackets (or parse trees) being the instrument of choice to display this. The reader who is in on their send-up knows why. There are centuries of data in its favor.  Here’s a taste.

Such bracketing allows one to distinguish the two readings of ‘old men and women,’ the ambiguity of ‘I photographed a woman with a camera,’ and the three readings in ‘I saw the girl sitting on the stoop.’ The several readings can be easily coaxed from these sentences, even if some jump to the ear faster than others.  So, if one’s interest is in accounting for how the same string of words can carry several readings labeled brackets (or equivalently parse trees) are very very handy.  And not only for this. Once one takes even a modestly serious look at language (something no satirist should do as too much seriousness wrong foots the parodic muse), one finds an inexhaustible number of intra-sentential relations that seem to supervene on hierarchical rather than linear properties of phrases and sentences.  This even has a name; grammatical rules are structure dependent.  Here’s a short (and not exhaustive) list of phenomena that advert to hierarchical structure: Aux fronting, WH question formation, topicalization, focus movement, VP fronting, VP ellipsis, reflexive binding, pronoun obviation, passivization, raising, negative concord, sluicing, parasitic gap licensing, donkey anaphora, island effects, …etc., etc., etc. In fact, it is almost impossible to find a syntactic phenomenon that fails to exploit hierarchical relations.

Knowing all of this, what does a good satirist do? Ignore and misrepresent. So for example, FBC find evidence from neuroscience and psychology that there are linear effects in language use.  Yup, those neuro guys using their big expensive fMRIs are able to finally show that Broca’s area lights up (I hope the pictures of Broca’s area are in purple, as red, blue and green are so last year) when both unimpaired and aphasic humans parse a sentence left to right. Even more astounding, Broca’s area also lights up in processing music! Wow, call the Nobel committee. This is a breathtaking discovery. Who would have thought that order left/right a difference might understanding in a sentence make!! Thank goodness for “repetitive transcranial magnetic stimulation” techniques for without these the relevance of linear order to parsing sentences would have surely remained hidden from view. You gotta love this.  Said with a straight pen this is very funny stuff.

But wait, there’s more. FBC (no doubt channeling Anthony Trollope who made a similar observation in his 1883 autobiography) note that linear order can affect the recognition of agreement dependencies.  So we say and hear approvingly The coat with the ripped cuffs were hanging in the closet rather than was hanging in the closet when the linearly nearest nominal is plural rather than singular; an interesting and curious effect that implicates linear proximity in assessing agreement. This demonstrates that those linguists who natter on about hierarchy being relevant for coding agreement effects are just obtuse.  Of course, the budding satirist should take note here and learn how to use data to misdirect.  Quietly kick under a nearby rug the fact that The doors in this compound were/*is closed does not pattern like the master example or that speakers when asked to assess the acceptability of the first sentence above with was in place of were rate it no worse than were, or that for many sentences proximity is irrelevant (e.g. The book that the boys liked was/*were long). No good parody lingers over complications: if linear order rules in one example then it does so everywhere, every time. Disagree and Occam (Sweeny Todd style) will gut you with his razor.  Anyone with aspirations to satire has to love FBC senseis' demonstrated artistry.

Consider one last wonderful manoeuver.  Chomsky fathered one of the canonical arguments for structure dependence based on aux inversion in Yes/No (Y/N) question.  The argument is as follows. Consider how Y/N questions are formed in English. Here are some examples:
(1)       a. Can John come
b. Will Mary sing
c. Is Frank kissing Sue
Using (1a-c) as “data” what kind of rule would one form? Easy: take the aux (i.e. “helping verb”) and move it to the front.  Question: which aux? Easy to answer in the examples in (1) as there is but one.  So let’s consider more complex forms.  For example, what is the correct form of the Y/N question taking (2a) as an answer?  It is clearly (2b) not (2c):
(2)       a. John is saying that Mat can swim
            b. Is John saying that Mat can swim
            c. *Can John is saying that Mat swim
Ok, it seems that the rule needs to specify which helping verb to take when there is more than one.  Here are two possible answers: take the one linearly closest to the front, i.e. the left most one. That works in correctly singling out (2b) from (2c). But, and there is always a ‘but’ isn’t there, what of yet more complex forms.  What do we do in (3a,b)?
(3)       a. The fact that John is sleeping should surprise Sue
            b. The man who Bill is talking to will surprise Sue
            c. That John was asleep all day might irritate Mary
Here if we move the leftmost helping verb, the one linearly closest to the front we get rather delightful word salad:
(4)       a. *Is the fact that John sleeping should surprise Mary
            b. *Is the man who Bill talking to will surprise Mary
            c. *Was that John asleep all day might irritate Mary
This indicates that the rule cannot be framed in terms of moving the linearly leftmost helping verb. So what’s the right restriction? Well it seems that we need to move the “highest” one, the one next to the subject.  The subject in (3a) is the fact that Bill is sleeping, in (3b) it is the man who Bill is talking to and in (3c) that John was asleep all day. So the right auxiliaries to move to form the unimpeachable (5a,b,c) are should, will and might respectively:
(5)       a. Should the fact that John is sleeping surprise Mary
            b. Will the man who Bill is talking to surprise Mary
            c. Might that John was asleep all day irritate Mary
Note, to get this right we invoke hierarchical notions: we need to treat several words (in fact the linear size of the subject is unbounded) as a single unit, i.e. the “subject.” Not surprisingly, the same rule works in the other cases as well. 

This short survey of some facts is, of course, not the whole story. However, it gives one a good flavor for the kind of problem that has convinced grammarians that hierarchical structure matters. And that humans are predisposed to exploit such structure in forming the rules of grammar.  And that this predisposition is built into the powers that humans bring to the task of acquiring and using language.  The reasoning is simple: the data that will tell the child to choose the hierarchical specification over the linear one is only available in examples like (4) and (5) and given that such sentences are virtually absent in the data available to the child learning the grammar it cannot be that the predisposition towards the hierarchical condition is data driven. 

Champions of the linear have analyzed this simple example to death (as Jerry Fodor once observed: this should teach Chomsky never to give a simple uncomplicated illustrative example!).  They have worked mightily to construct algorithms using bi- and tri-grams that can distinguish the relevant good and bad cases all the while eschewing hierarchical structure. FBC note this and get rid of all problems of hierarchy with the wave of reference or two.  Moreover, being first-rate parodists they keep hidden the fact that all these algorithms have a common problem. Were the facts opposite to those cited the relevant algorithms could “learn” these as well. So, it is only an accident that that there is no natural language that has a rule analogous to the one that generates (4). For FBC such a natural language would be no less humanly accessible than the ones we happen to find.  We even know what this linguistic “gap” would look like ((4) good, (5) bad).  Thus the attested linguistic holes on this view are purely accidental.  It is of course perfectly conceivable that the absence of well-formed structures like (4) is an accidental gap and that children could learn to form Y/N questions in this way.  Accidents do happen.  And some people can be sold famous bridges.  Such niceties however would clog up a good parody, and so FBC wisely put them aside.

There are other titivations that FBC artfully use to ornament their piece, to make it seem like they are really serious about the overall proposal. Good satire demands detail and a straight face.  However, I don’t want to ruin your pleasure in finding these bonbons for yourself.  What I would like to do is end with an appreciation of what I take the real tour de force of their masterpiece to be. As noted at the outset, FBC wrap the whole discussion in a delightful Darwinian wrapping. It seems that natural selection cannot digest the kinds of hierarchy found in grammars. Thus, linear relations it is or you are a creationist!! Like any good parody, this is more suggested than stated. But the evolutionary angle adds a welcome frisson to the discussion. So, not only is this paper really funny, but it smacks of the monumental. Just think, hierarchy:God:religious fanatic, linearity:Darwin:hard scientist.  Onto the ramparts! It’s time for some culture war.

Let me end with another round of kudos. This paper is a must read. Until I got through it I thought that the art of the academic lampoon was dead. FBC have proved me wrong. There are levels of silliness, stupidity and obtuseness left to plumb. Thanks to FBC and the Royal Society for demonstrating that parody and satire are still possible.