Comments

Showing posts with label Autonomy of syntax. Show all posts
Showing posts with label Autonomy of syntax. Show all posts

Monday, October 14, 2013

Syntax in the brain


It’s not news that syntactic structure is perceptible in the absence of meaningfulness. After all, even though the slithy toves did gyre and gimble in the mabe, still and all colorless green ideas do sleep furiously. Not news maybe, but still able to instruct, as I found out in getting ready for this years Baggett Lectures (see here). The invited speaker this year is Stanislas Dehaene and to get ready I’ve read (under the guidance of my own personal Virgil, Ellen Lau, guide to the several nether circles of neuroscience) some recent papers by him on how the brain does syntax. One (here) was both short (a highly commendable property) and absorbing. Like many neuroscience efforts, this one is collaborative, brought to you by the French team of Pallier, Devauchelle and Dehaene (PDD). I found the PDD results (and methods to the degree that I could be made to appreciate them) thought provoking. Here’s why.

The paper starts from the conventional grammatical assumption that “sentences are not mere strings of words but possess a hierarchical structure with constituents nested inside each other” (2522).[1] PDD’s task is to find out where, if anywhere, the brains tracks/builds this hierarchy. PDD construct a very clever model that allows them to use fMRI techniques to zero in on those regions sensitive to hierarchical structure.

Before proceeding, it’s worth noting that this takes quite a bit of work. Part of what makes the paper fun, is the little model that allows PDD to index hierarchy to differential blood flows (the BOLD (blood-oxygen-level-dependent) response, which is what an fMRI tracks). It predicts a linear relationship between the BOLD response and phrasal size (roughly indexed to word length (a “useful approximation”) and they use this relationship to probe ROIs (i.e. regions of interest (Damn I love these acronyms)) that respond to this predicted relationship using two different kinds of linguistic probes. The first are strings of words containing phrases ranging from 1 to 12 words long (actually 12/6/4/3/2/1 e.g. 12: I believe that you should accept the proposal of your new associate, 4: mayor of the city he hates this color they read their names). The second are jabberwocky strings with the same structure (e.g. I tosieve that you should begept the tropufal of your tew viroate). Here’s what they found:

1.     They found brain regions that responded to these probes in the predicted linear manner. Four in the STS region, and two in the left inferior gyrus (IFGtri and IFGorb).
2.     The regions responded differentially to the two kinds of linguistic probes. Thus, all four regions responded to the first kind of probe (“normal prose”) while jabberwocky only elicited responses in IFGorb (with some response with a lower statistical threshold in left posterior STS and IFGtri).

In sum, different brain regions light up exclusively to phrasal syntax independent of content words. Thus, the brain seems to distinguish contentful morphemes from more functional ones and it does so by processing the relevant information in different regions.
And this is interesting, I believe. Why?

Well, think of what PDD could have found. One possibility is that all words are effectively the same; the only difference between content words and functional vocabulary residing in their statistical frequency, the closed class content words being far more common than the open class contentful vocab. Were this correct, then we should expect no regional segregation of the two kinds of vocab, just, say, a bigger or smaller response based on the lexical richness of the input. Thus, we might have expected all regions to respond equally to all of the inputs though the size of the response would have differed. But this is not what PDD found. What they found is differential activity across various regions with one group of sites responding exclusively to structural input even in the absence of meaningful content. This sure smells a lot like the wetware analogue of the autonomy of syntax thesis (yes, I was not surprised, and yes, I was nonetheless chuffed). PDD (2526) make exactly this autonomy of syntax point in noting that their results underline “the relative independence of syntax from lexico-semantic features.”

Second, the results might have implications for grammar lexicalization (GL) (an idea that Thomas has been posting about recently). From what I can tell, GL sees grammatical dependencies as the byproduct of restrictions coded as features on lexical terminals. Grammatical dependencies on the GL view just are the sum total of lexical dependencies. If this is correct, then a question arises: what does this kind of position lead us to expect about grammatical structure in the absence of such lexical information? I assume that Jabberwocky vocab is not part of our lexicon and so a grammar that exclusively builds structure based on info coded in lexical terminals will not have access to (at least some) grammatically relevant information in Jabberwocky input (e.g. how to combine the and slithy toves in the absence of features on the latter?). Does this mean that we should not be able to build syntactic structure in the absence of the relevant terminals or that we will react differently to input composed of “real” lexemes vs Jabberwocky vocab? Behaviorally, we know that we can distinguish well- from ill- formed Jabberwocky. So we know that the absence of a lot of lexical vocab does not impede the construction of syntactic structure. PDD further shows that neurally there are parts of the brain that respond to syntactic structure regardless of the presence of featurally marked terminals (and recall, that this need not have been the case). A non-GL view of grammar has no problems with this, as grammatical structure is not a by-product of lexical features[2] (at least not in general, only the functional vocab is possibly relevant).[3] I cannot tell if this is a puzzle if one takes a fully GL view of things, but it certainly indicates that parts of the brain seem tuned to structure without much apparent lexical content. Indeed, it suggests (at least to me) that GL might have things backwards: it’s not that grammatical structure arises from lexical specifications but that lexical specifications are by-products of enjoying certain syntactic relations. Formally, this might be a distinction without a difference, but psychologically and neurally, these two ways of viewing the ontogeny of structure look like they might (note the weasel word here) have very different empirical consequences.

I am sure that I have over-interpreted the PDD results. The structure they probe is very simple (just right branching phrase structure) and the results rely on some rather radical simplifications. However, I am told, that this is the current state of the fMRI art and even with these caveats, PDD is interesting and thought provoking, and, who knows, maybe more than a little relevant to how we should think about the etiology of grammar. It seems that brains, or at least parts of brains, are sensitive to structure regardless of what that structure contains. This is an old idea, one perfectly expected given a pretty orthodox conception of the autonomy of syntax. It’s nice to see that this same conception is leading to new and intriguing work investigating how brains go about building structure.



[1] This truism, sadly, is not always acknowledged. For example, it is still possible for psychologists to find a publishing outlet of considerable repute for papers that deny this (see here). Fortunately, flat-earthism seems to be loosing its allure in neuroscience.
[2] Note that this does not entail that grammatical information might not be lexicalized for tasks like parsing. It only implies that grammatical dependencies are more primitive than lexical ones. The former determine the latter, not vice versa.
[3] I say ‘possibly’ as the functional vocab might just be the morphological outward manifestations of grammatical dependencies rather than the elements from which such dependencies are formed. There is not obvious reason for thinking that structure is there because of the functional vocab, though there is reason for assuming that functional vocab and syntactic structure are correlated, sometimes strongly.

Friday, January 11, 2013

Darwin's Problem


Humans are uniquely linguistically facile. This raises an interesting evolutionary question, an abstract version of which Minimalists have taken very much to heart: How did this linguistic capacity arise in the species? Following Cedric Boeckx, let’s dub this “Darwin’s Problem.” Answers to this problem have two separable parts: (i) an account of how that which made language a cognitive option became mentally available, (ii) an account of how the available option became fixed in the species.

Most give a “miracle theory” account of (i). What I mean is that it is regularly assumed that some kind of adventitious genetic change/mutation occurred that, when added to the cognitive apparatus already there, combined with it to allow for the emergence of a mental faculty with the key features of FL.  The “miracle” means to mark the observation that this change “just happened,” it’s a brute fact. Minimalists try to (abstractly) characterize the nature of this change (what was added), but that there is no attempt to explain why the change occurred. It just did. What’s up for grabs is the nature of the change (adding Merge being the currently favored candidate, though there have been other proposals, (one by yours truly)) and the number of these. Given the logic of the case (short time span etc.), one miracle is acceptable, two maybe barely tolerable, three fougetaboutit! At any rate, a miracle occurred sometime in the last (roughly) 100,000 years in at least one member of the species.  This brings us to (ii).

Once the miracle occurs, it must be fixed in the population, presumably by giving its bearers some selective advantage (this is the Darwin part).  With respect to language there are basically two possible sources for this advantage, which correspond with two classical views about the utility of language; language as vehicle for communication and language as vehicle for thought. Pinker and Bloom are perhaps the most famous advocates of the first conception. Chomsky is a well-known advocate of the second.

There are two main problems with the communication view.

First, it requires double the number of miracles. Note, this follows from two observations: that it takes at least two to communicate and mutations (the required miracle) originate in individuals and spread to populations via the reproductive success of the favored individuals. Thus, as improbable is it is for Merge, say, to pop into an individual ape mind once the chances of it doing so twice in two different (assuming that communication is between at least two individuals) proximate (if not near each other than the capacity to communicate won’t be realized) individuals, is much much more improbable still. Indeed if the events are independent, then it’s the square of the probability of the unique event.

Second, we need a story about why the particular form of communication, a communication system based on Merge like grammars, is so much more advantageous than a simpler system would be.  Here’s what I mean. Consider a simple linear N-V-(N) grammar with a vocabulary of 500 verbs and 1,000 nouns. This can support roughly 500,000,000 different messages. That’s a good number of messages, all without hierarchical recursion.  We know that animal communication doesn’t require recursion. The evolutionary question then  is what communicative advantage does the miracle promote that would be particularly advantageous?

Considerations like these led many to conclude that the main selective advantage of language was its enrichment of thought rather than its communicative efficacy. Here’s Francois Jacob’s take:

…the role of language as a communication system between individuals would have come about secondarily…Its primary function would rather have been, as with earlier evolutionary steps in mammals, the representation of a finer and “richer” reality,” a way of handling more efficiently a greater amount of information. As exemplified throughout the whole animal kingdom, communication can be easily established between individual organisms.  Even among hominids which had to hunt and live in community, most of the information to be shared with others and concerning immediate features of life could be handled by means of rather simple codes.  In contrast, to translate a visual and auditory world so that objects and events can be precisely labeled and recognized weeks or years later requires a much more elaborate coding system. The quality of language that makes it unique does not seem to be so much its role in communicating directives for action as its role in symbolizing, in invoking cognitive images. We mold our “reality” with our words and our sentences in the same way as we mold it with our vision and our hearing.  And the versatility of human language also makes it a unique tool for the development of imagination. It allows infinite combinations of symbols and, therefore, mental creation of possible worlds (58).[1]

Thus, the proposal is that grammatical structures enhance the class of entertainable and easily retrievable of thoughts. It allows for the imagination of alternatives, thereby, one might suppose, enhancing planning and action (as well as making dawdling that much more enjoyable!).  At any rate, were this so, then it is not hard to imagine how a miracle that enabled this would immediately endow its individual bearer with the kinds of advantages that natural selection cares about and how, therefore, this miracle could go forth and multiply via its bearers going forth and multiplying.

Before going on, we should appreciate that all of this is speculative.  As Lewontin has made clear, there is a big, perhaps ultimately insurmountable, step between this and a serious scientifically grounded selective explanation. As he demonstrates in detail (c.f. his "The Evolution of Cognition" in volume 4 of The Invitation to Cognitive Science), it’s extremely hard to move beyond just-so stories and provide empirically justified evolutionary accounts of cognitive capacities.

This said, there are tantalizing hints and what I want to point to one.  I have just reread some fascinating work (from 1999) by Hermer-Vazquez, Spelke and Katsnelson (H-VSK) that bears on these questions. They provide evidence for the kind of scenario that Jacob describes above.  Here’s what they found (from the abstract):

Under many circumstances, children and rats reorient themselves through a process which operates only on information about the shape of the environment... In contrast, human adults relocate themselves more flexibly by conjoining geometric and non-geometric information to specify their position. The present experiments used a dual-task method to investigate the processes that underlie the flexible conjunction of information…Together the experiments suggest that humans’ flexible spatial memory depends on the ability to combine divers information sources rapidly into unitary representations and that this ability, in turn, depends on natural language.

The experiments all involve disorienting children and adults in a rectangular room. The task is to find something in a prescribed corner. Sometimes the indicated corner abuts a wall with a certain color, thereby distinguishing it from the geometrically analogous opposite corner. Adults are able to exploit the additional color information to locate themselves and thus to identify the right corner (i.e. color serves to disambiguate the geometrical information). Prelinguistically capable kids cannot. Nor can rats.  More interesting still, H-VSK found a way of stopping adults from using the color information by having them engage in a language task while reorienting themselves. Presto, the adults start acting like kids and rats.  Importantly, engaging in additional non-linguistic tasks during reorientation does not stop successful identification of the correct corner.  This strongly implicates language use in facilitating spatial orientation.

I hope I have piqued your interest. The experiments are a delight to read (so do so) and the implications for Jacob’s (and Chomsky’s) evolutionary scenario very suggestive. Here we have a case where linguistic facility directly enhances something as basic as spatial orientation, a capacity that it does not take much imagination to suppose would be useful to our hunter-gatherer ancestors and would endow selective advantage in a wide range of plausibly relevant environments.

How exactly does language help?  H-VSK speculate that language constitutes a kind of interlingua allowing diverse information from separately encapsulated cognitive modules to combine into single thoughts.  The capacity to so combine diverse concepts allows for more complex thoughts and thereby allows, in Jacob’s words, for “the representation of a finer and “richer” reality.” In sum, were the Jacob-Chomsky speculation on the right track we might expect to find cognitive enhancement for selectionistically valuable traits, and this seems to be what H-VSK have found. Wow!

Need I say that this is still very speculative?  However, though a first step, it is very interesting and fits well with certain other assumptions out there minimalists are sure to find congenial.

First, standard Minimalist theory proposes a strong asymmetry between the two interfaces.  Rather than syntax being a pairing of sound (AP) and meaning (CI) (the standard view since Aristotle), it is more accurately thought of as a relation between structure and meaning with sound as an add-on (Chomsky has strongly pushed this line of late).  The derivation from lexicon to CI is clean and well designed (e.g. it meets Inclusiveness, Extension and Full Interpretation). The mapping to sound is considerably messier (e.g. does not conform to Inclusiveness). This fits well with the Jacob-Chomsky conception which presumes that the real biological action starts with the generation of complex thoughts that grammar makes available, not spoken outputs, which are a later accretion. 

Second, Generative Syntax endorses the autonomy of syntax thesis (AOS). Though AOS has often been misunderstood to assert that there is no relation between the grammar and meaning, it actually means that the primitives and operations of the grammar are independent of the contents of what they are used to express. In particular, syntactic categories, principles and operations to not reduce to semantic ones. Many have taken this to be a serious defect. However, in the context H-VSK’s results it looks like a great design feature.  Precisely because the syntax is autonomous it is able to combine information from different encapsulated modules. In other words, autonomy is just the flip side of not being modularly restricted.  The intra modular primitives and operations cannot do this, which is what makes it impossible for rats, young kids and linguistically distracted adults from combining different kinds of information (i.e. predicates from different modules). From the present perspective, a more revealing term for the autonomy of syntax might be the inter-modularity of syntax, autonomy being precisely the property we want in a tool required to combine diverse types of thoughts and concepts, ones otherwise confined to specialized cognitively encapsulated modules.

Last, consider hierarchy.  The kind of combination H-VSK’s tasks require is one that allows for diverse kinds of information to work together to produce finer and finer descriptions. In other words, we want the capacity to modify, viz. stack adverbs, specify events, combine nouns and adjectives, use sentences to cut down possibilities (e.g. as  relativization does) etc.  This is the conceptual value added that syntactic hierarchy provides, and it does so in spades. 

In sum, the Jacob-Chomsky “conjecture” when combined with a generative syntax with a minimalist flavor has a suggestive tang: it links what is special (viz. recursive hierarchy) with what is plausibly beneficial (viz. the capacity to entertain new and useful thoughts).

One of the novelties of the Minimalist Program has been the elevation of Darwin’s Problem to prominence along side Plato’s.  Interestingly, it appears that the empirical just-so stories of yore might finally graduate to empirically so-so stories and, one day maybe even to thus-so stories. Wouldn’t that be nice? Can’t blame a person for dreaming. In the meantime, take a look at H-VSK. It’s a great paper.


[1] From The Possible and the Actual (1982). University of Washington Press; Seattle.

Thursday, November 8, 2012

The Autonomy of Syntax



Barbara Partee gave her first lecture today (which was very interesting btw) and discussed three versions of the autonomy of syntax thesis (AoS) that I would like to quickly comment on here. Non- generativists regularly advert to the AoS citing it as one of Generative Grammar’s more obvious absurdities (e.g. see here).  The thesis is interpreted as asserting that syntax is independent of meaning, understanding this to assert that the syntax of some expression has no consequences on what it means and vice versa. This view is rightly taken to be absurd. Fortunately, nobody has ever held this position.  So what then is the AoS?

Barbara identified three different interpretations:

First, a negative thesis: syntax is not reducible to operations of the interface systems.  This means that syntactic structures, syntactic rules and syntactic primitives are not reducible to semantic structures, rules and primitives (nor for that matter phonological ones, though few have been tempted with this kind of reduction). Of course, the converse is also true; the properties of the interpretive systems (meaning and “sound”) are not reducible to properties of the syntax.  There are systematic interconnections, but no reduction.  The classic illustration of this comes from thinking about the sentence “colorless green ideas sleep furiously” and comparing it to “green sleep colorless furiously ideas.” As observed long ago, neither sentence is particularly meaningful but the first is syntactically regular (viz. conforms to a standard English pattern as in “clear limpid prose persuades easily”), hence its greater acceptability when compared to the word salad of the second example despite the semantic anemia of both.  The same point can be made by considering unacceptable sentences like “The girl seems sleeping.” This sentence has a perfectly clear meaning (viz. “the girl seems to be sleeping” and not “the girl seems sleepy”) even though the sentence is syntactically odd. To speak somewhat inexactly, ‘meaningful’ and ‘grammatical’ dissociate and hence there can be no reduction of one to the other.[1]

The second interpretation of the AoS is more theory internal: it asserts that the application of syntactic operations is not conditional on semantic factors. Here is a useful formulation stolen form a handout of Fritz Newmeyer.

The syntactic rules and principles of a language are formulated without reference to meaning, discourse, or language use
This is stronger than the first version of the AoS for even if syntactic operations are semantically irreducible, the syntactic generation of a sentence could (logically) depend on the interpretation that the generated sentence has or the context in which it is used or the communicative intent of the speaker. Generative theories have generally respected this version of the AoS as well, though there have been suggested principles that violate it (at least in spirit). The Fox-Reinhart thesis is a contemporary example. It sanctions movement just in case the movement has an effect on meaning (Chomsky’s version adds a strong feature to the derivation just in case it has an effect on interpretation). Both appear to violate this second version of the AoS.

Just a caveat before moving on: the AoS should not be confused with the syntax first thesis mentioned in the psycho-ling literature. The latter is taken to be a principle of parsing which states that an expression must be assigned a complete syntactic structure before it is semantically evaluated.  The AoS commits no hostages to the temporal dynamics involved in parsing.  Parsing could interleave syntactic and semantic rules without the latter conditioning the former.

Barbara dubs the third interpretation explanatory autonomy.  It is a methodological principle regulating what counts as a valid motivation for a syntactic proposal.  Barbara interprets this to mean that semantic “facts” cannot justify postulating syntactic structure, only syntactic “facts” can.  I find this version of the AoS problematic and am unsure if it ever had much of a hold on syntactic practice, though if it did it shouldn’t have. First, it is not clear that facts come labeled ‘syntactic’ and ‘semantic,’ and if not the regulative proposal is contentless (I believe that Chomsky once made a similar point but damn if I can recall where).  Second, the intuitive version of this interpretation of the AoS has been regularly violated in practice from the earliest days of generative grammar. For example, what we now call thematic considerations have regularly been invoked in postulating common underlying structures for sentences related by movement, e.g. one good reason for thinking that actives (John kissed Mary) and passives (Mary was kissed (by John)) derive from a common underling structure is it allows the structural configurations of theta role assignment to be streamlined (e.g. in both cases ‘Mary’ is the underlying object).  Indeed, this is so for all cases of movement. 

Perhaps this version of the AoS had greater purchase in the Aspects era when the Katz-Postal hypothesis (KPH) was widely adopted.  In KPH theories the only input to semantic interpretation is Deep Structure. If correct, transformations have to be meaning preserving as their outputs, Surface Structures, do not feed semantic interpretation.  In such a context semantic concerns cannot motivate a particular transformation. If one identifies “transformational” with “syntactic” (a mistake, as Deep Structure is also a syntactic level in an Aspects style theory) the avoidance of “semantic” considerations would make sense.  However, to repeat, there is no principled way of distinguishing a syntactic form a semantic fact and so the methodological dictum is hollow.

Let me end with one more quick distinction: The AoS should also be distinguished from another important claim, which I will dub the Primacy of Syntax thesis (PoSt).  PoSt is a substantive claim asserting that syntax is where generativity lives.  More specifically  syntax is the recursive engine in natural language, the interfaces being “interpretive” rather than generative. No generative syntax, no complex thoughts, no unbounded structures.  The syntax generates the structures that the semantics (and phonology) interprets. But, the semantics and phonology as such have no generative powers.  If PoSt holds true then the first interpretation of the AoS follows.  This said, the theses are different and are usefully distinguished.



[1] This is speaking inexactly for ‘grammatical’ contrasts with ‘meaningful’ in being a technical term.  The predicate for observables is ‘acceptable.’ What the AoS examples above observe is that a notion of syntactic well-formedness is required in addition to meaningfulness if relative acceptability is to be accounted for.