Comments

Showing posts with label Dehaene. Show all posts
Showing posts with label Dehaene. Show all posts

Wednesday, October 5, 2016

Form and function; the sources of structure

I just read a fascinating paper and excellent comment thereupon in Nature Neuroscience (thx to Pierre Pica for sending them along) (here and here). The papers make the interesting point that, as I will argue below, illuminate two very different views of what structure is and where it comes from.  The two views have names that should now be familiar to you: Rationalism (R) and Empiricism (E). What is interesting about the two papers discussed below is that they indicate that R and E are contrasting philosophical conceptions that have important empirical consequences for very concrete research. In other words, R and E are philosophical in the best sense, leading to different conceptions with testable empirical (though not Empiricist) consequences. Or, to put this another way, R and E are, broadly speaking, research programs pointing to different conceptions of what structure is and how it arises.[1]

Before getting into this more contentious larger theme, let’s review the basic findings. The main paper is written by the conglomerate of Saygin, Osher, Norton, Youssoufian, Beach, Feather, Gaab, Gabrielli and Kanwisher (Henceforth Saygin et al). The comment is written by Dehaene and Dehaene-Lambertz (DDL). The principle finding is, as the title makes admirably clear, that “connectivity precedes function in the development of the visual word form area.” What’s this mean?

Saygin et al observes that the brain is divided up into different functional regions and that these are “found in approximately the same anatomical location in virtually every normal adult”.[2] The question is how this organization arises: “how does a particular cortical location become earmarked”? (Saygin et al:1250). There are two possibilities: (i) the connectivity follows the function or (ii) the function follows the connectivity. Let’s expand a bit.

(i) is the idea that in virtue of what a region of brain does it wires up with another region of brain because of what it does at roughly the same time. This is roughly the Hebbian idea that regions that fire together wire together (FTWT). So, a region that is sensitive to certain kinds of visual features (e.g. Visual Word From Area (VFWA)) hooks up with an area where “language processing is often found” (DDL:1193) to deliver a system that undergirds reading (coding a dependency between “sounds” and “letters”/”words”). 

(ii) reverses the causal flow. Rather than intrinsic functional properties of the different regions driving their connectivity (via concurrent firing), the extrinsic connectivity patterns of the regions drives their functional differentiation. To coin a phrase: areas that are wired together fire together (WTFT). This is what Saygin et al finds :

This tight relationship between function and connectivity across the cortex suggests a developmental hypothesis: patterns of extrinsic connectivity (or connectivity fingerprints) may arise early in development, instructing subsequent functional devel­opment.

The may is redeemed to a does by following young kids before and after they learn to read. As DDL summarizes it (1193):

To genuinely test the hypothesis that the VWFA owes its specializa­tion to a pre-existing connectivity pattern, it was necessary to measure brain connectivity in children before they learned to read. This is what Saygin et al. now report. They acquired diffusion-weighted images in children around the age of 5 and used them to reconstruct the approximate trajectory of anatomical fiber tracts in their brain. For every voxel in the ven­tral visual cortex, they obtained a signature pro­file of its quantitative connectivity with 81 other brain regions. They then examined whether a machine-learning algorithm could be trained to predict, from this connectivity profile, whether or not a voxel would become selective to written words 3 years later, once the children had become literate. Finally, they tested their algo­rithm on a child whose data had not been used for training. And it worked: prior connectivity predicted subsequent function (my bold, NH). Although many children did not yet have a VWFA at the age of 5, the connections that were already in place could be used to anticipate where the VWFA would appear once they learned to read.

I’ve bolded the conclusion: WTFT and not FTWT. What makes the Saygin et al results particularly interesting is their precision. Saygin et al is able to predict the “precise location of the VWFA” in each kid based on “the connectivity of this region even before the functional specialization for orthography in the VWFA exists” (1254). So voxels that are not sensitive to words and letters before kids learn to read, become so in virtue of prior (non functionally based) connections to language regions.

Some remarks before getting into the philosophical issues.

First, getting to this result requires lots of work, both neuro imaging work and good behavioral work. This paper is a nice model for how the two can be integrated to provide a really big and juicy result.

Second, this appeares in a really fancy journal (Nature Neurosceince) and one can hope that it will help set a standard for good cog-neuro work, work that emphasizes both the cognition and the neuroscience. Saygin et al does a lot of good cog work to show that in non-readers VWFA is not differentially sensitive to letters/words even though it comes to be so sensitive after kids have learned to read.

Third, DDL points out (1192-3) that whatever VWFA is sensitive to it is not simply visual features (i.e. a bias for certain kinds of letter like shapes).  Why not? Because (i) the region is sensitive to letters and not numerals despite letters and numerals being formed using the same basic shapes and (ii) VWFA is located in the same place in blind subjects and non-blind ones so long as the blond ones can read braille or letters converted into “synthetic spatiotemporal sound patterns.” As DDL cooly puts it:

This finding seems to rule out any explanation based on visual features: the so-called ‘visual’ cortex must, in fact, possess abstract properties that make it appropriate to recognize the ‘shapes’ of letters, numbers or other objects regardless of input modality.

So, it appears that what VWFA takes as a “shape” is itself influenced by what the language area would deem shapely. It’s not just two perceptual domains with their independently specifiable features getting in sync, for what even counts as a shape depends on what an area is wired to. VWFA treats a “shape” as letter-like if it tags a “shape” that is languagy.

Ok, finally time for the sermon: at the broadest level E and R differ in their views of where structure comes from and its relation to function.

For Es, function is the causal driver and structure subserves it. Want to understand the properties of language, look at its communicative function. Want to understand animal genomes, look at the evolutionarily successful phenotypic expressions of these genomes. Want to understand brain architecture, look at how regions function in response to external stimuli and apply Hebbian FTWT algorithms.  For Es, structure follows function. Indeed, structure just is a convenient summary of functionally useful episodes. Cognitive structures reflect shaping effects of the environmental inputs of value to the relevant mind. Laws of nature are just summaries of what natural objects “do.” Brain architectures are reflections of how sensory sensitive brain regions wire up when concurrently activated. Structure is a summary of what happens. In short, form follows (useful) function.

Rs beg to differ. Es understand structure as a precondition of function. It doesn’t follow function but precedes it (put Descartes before the functional horse!). Function is causally constrained by form, which is causally prior. For Rs, the laws of nature reflect the underlying structure of an invisible real substrate. Mental organization causally enables various kinds of cognitive activity. Linguistic competence (the structure of FL/UG and the structure of individual Gs) allows for linguistic performances, including communication and language acquisition. Genomic structure channels selection.  In other words, function follows form. The latter is causally prior. Structure  “instructs” (Saygin et al’sterm) subsequent functional development.

For a very long time, the neurosciences have been in an Empiricist grip. Saygin et al provides a strong argument that the E vision has things exactly backwards and that the Hebbian Esih connectionist conception is likely the wrong way of understanding the neural and functional structure of the brain.[3] Brains come with a lot of extrinsic structure and this structure casually determines how it organizes itself functionally. Moreover, at least in the case of the VWFA, Darwinian selection pressures (another kind of functional “cause”) will not explain the underlying connectivity. Why not? Because as DDL notes (1192) alphabets are around 3800 years old and “those times are far too short for Darwinian evolution to have shaped our genome for reading.” That means that Saygin et al’s results will have no “deeper” functional explanations, at least as concerns the VWFA. Nope, it’s functionally inexplicable structure all the very bottom. Connectivity is the causal key. Function follows. Saygin et al speculate that what is right for VWFA will hold for brain organization more generally. Is the speculation correct? Dunno. But being a card carrying R you know where I’d lay my bets.


[1] I develop this theme in an article here.
[2] This seems like a pretty big deal to me and argues against any simple minded view of brain plasticity, I would imagine. Maybe any part the brain can perform any possible computation, but the fact that brains regularly organize themselves in pretty much the same way seems to indicate that this organization is not entirely haphazard and that there is method behind it. So, if it is true that the brain is perfectly plastic (which I really don’t believe) then this suggested that it is not the computational differences responsible for its large scale functional architecture. Saygin et al suggest another causal mechanism.
[3] In this it seems to be reprising the history of immunology which moved from a theory in which the environment instructed the immune system to one in which the structure of the immune system took causal priority. See here for a history.

Monday, September 19, 2016

Brain mechanisms and minimalism

I just read a very interesting shortish paper by Dehaene and associates (Dehaene, Meyniel, Wacongne, Wang and Pallier (DMWWP) that appeared in Neuron. I did not find an open source link, but you can use this one if you are university affiliated. I recommend it highly, not the least reason being that Neuron is a very fancy journal and GG gets very good press there. There is a rumor running around that Cog Neuro types have dismissed the findings of GG as of little interest or consequence to brain research. DMWWP puts paid to this and notes, quite rightly, that the problem lies less with GG than with the current state of brain science. This is a decidedly Gallistel inspired theme (i.e. the cog part of cog-neuro is in many domains (e.g. language) healthier and more compelling than the neuro part and it is time for the neuro types to pay attention and try to find mechanisms adequate for dealing with the well grounded cog stuff that has been discovered rather than think it msut be false because the inadequate and primitive neuro models (i.e. neural net/connectionist) don’t have ways of dealing with it) and the more places it gets said the greater the likelihood that CN types will pay attention. So, this is a very good piece for the likes of us (or at least me).

The goal of the paper is to get Cog-Neuro Science (CNS) people to start taking the integration of behavioral, computational and neural as CNS’s main central concern. Here is the abstract:

A sequence of images, sounds, or words can be stored at several levels of detail, from specific items and their timing to abstract structure. We propose a taxonomy of five distinct cerebral mechanisms for sequence coding: transitions and timing knowledge, chunking, ordinal knowledge, algebraic patterns, and nested tree structures. In each case, we review the available experimental paradigms and list the behavioral and neural signatures of the systems involved. Tree structures require a specific recursive neural code, as yet unidentified by electrophysiology, possibly unique to humans, and which may explain the singularity of human language and cognition.

I found the paper interesting in at least three ways.

First, it focuses on mechanisms, not phenomena. So, the paper identifies five kinds of basic operations that reasonably underlies a variety of mental phenomena and takes the aim of CNS to (i) find where in the brain these operations are executed, (ii) provide descriptions of circuits/computational operations that could execute such operations and (iii) investigate how these circuits might be/are neutrally realized.

Second, it shows how phenomena can be and have been used to probe the structure of these mechanisms. This is very well done for the first three kinds of mechanisms: (i) approximate timing of one item relative to the proceeding one, (ii) chunking items into larger units, and (iii) the ordinal ranking of items. Things get more speculative (in a good way, I might add) for the more “abstract” operations: the coding of “algebraic” patterns and nested generated structures.

Third, it gives you a good sense of the kinds of things that CNS types want from linguistics and why minimalism is such a good fit for these desires.

Let me say a word about each.

The review of the literature on coding time relations is a useful pedagogical case. DMWWP reviews the kind of evidence used to show that organisms “maintain internal representations of elapsed time” (3). It then look for “a characteristic signature” of this representation and the “killer” data that supports the representational claim. It then reviews the various brain locations that respond to these signature properties and review the kind of circuit that could code this kind of representation, arguing that “predictive coding” (i.e. ones that “form an internal model of input sequences”) is the right one in that it alone accommodates the basic behavioral facts (4) (basically minsmatched negativity effects without an overt mismatch). Next, it discusses a specific “spiking neuron model” of predictive coding (4) that “requires a neurophysiological mechanism of “time stamp” neurons that are tuned to specific temporal intervals,”  which have, in fact, been found in various parts of the brain. So, in this case we get the full Monte: a task that implicates signature properties of the mechanism, that demands certain kinds of computational circuits, realized by specific neuronal models, realized in neurons of a particular kind, found in different parts of the brain. It is not quite the Barn Owl (see here), but it is very very good.

DMWWP do this more or less again for chunking, though in this case “the precise neural mechanisms of chunk formulation remain unknown” (6). And then again for ordinal representations. Here there are models for how this kind of information might be neutrally coded in terms of “conjunctive cells jointly sensitive to ordinal information and stimulus identity” (8). These kinds of conjunctive neurons seem to be all over the place, with potential application, DMWWP suggests, as neuronal mechanisms for thematic saturation.

The last two kinds of mechanisms, those that would be required to represent algebraic patterns and hierarchical tree-like structures are behaviorally very well-established but currently pose very serious challenges on the neuro side. DMWWP observes that humans, even very young ones, demonstrate amazing facility in tracking such patterns. Monkeys also appear able to exploit similar abstract structures, though DMWWP suggests that their algebraic representations are not quite like ours (9). DMWWP further correctly notes that these sorts of patterns and the neural mechanisms underlying them are of “great interest” as “language, music and mathematics” are replete with such. So, it is clear that humans can deploy algebraic patters which “abstract away from the specific identity and timing of the sequence patterns and to grasp their underlying pattern,” and maybe other animals can too. However, to date there is “no accepted neural network mechanism to accomplish this and it looks like “all current neural network models seem too limited to account for abstract rule-extraction abilities” (9). So, the problem for CNS is that it is absolutely clear that human (and maybe monkey) brains have algebraic competence though it is completely unclear how to model this in wet ware. Now, that is the right way to put matters!

This last reiterates conclusions that Gallistel and Marcus have made in great detail elsewhere. Algebraic knowledge requires the capacity to distinguish variables from values of variables. This is easy to do in standard computer architectures but is not at all trivial in connectionist/neural net frameworks (as Gallistel has argued at length (e.g. see here)). Indeed, one of Gallistel’s main arguments with such neural architectures is their inability to distinguish variables from their values, and to store them separately and call them as needed. Neural nets don’t do this well (e.g. they cannot store a value and later retrieve it), and that is the problem because we do and we do it a lot and easily. DMWWP basically endorses this position.

The last mechanism required is one sufficient to code the dependencies in a nested tree.[1] One of the nice things about DMWWP is that it recognizes that linguistics has demonstrated that the brain codes for these kinds of data structures. This is obvious to us, but the position is not common in the CNS community and the fact that DMWWP is making this case in Neuron is a big deal. As in the case of algebraic patterns, there is no good models of how these kinds of (unbounded) hierarchical dependencies might be neurally coded. The DMWWP conclusion? The CNS community should start working on the problem. To repeat, this is very different from the standard CNS reaction to these facts, which is to dismiss the linguistic data because there are no known mechanisms for dealing with it.

Before ending I want to make a couple of observations.

First, this kind of approach, looking for basic computational mechanisms that are implicated in a variety of behaviors, fits well with the aims of the minimalist program (MP). How so? Well, IMO, MP has two immediate theoretical goals: to show that the standard kinds of dependencies characteristic of linguistic competence are all different manifestations of the same underlying mechanism (e.g. are all instances of Merge). Were it possible to unify the various modules (binding, movement, control, selection, case, theta, etc) as different faces of the same Merge relation and were we able to find the neural “merge” circuit then we would have found the neural basis for linguistic competence. So if all grammatical relations are really just ones built out of merges, then CNSers of language could look for these and thereby discover the neural basis for syntax. In this sense, MP is the kind of theory that CNSers of language should hope is correct. Find one circuit and you’ve solved the basic problem. DMWWP clearly has bought into this hope.

Second, it suggests what GGers with cognitive ambitions should be looking for theoretically. We should be trying to extract basic operations from our grammatical analyses as these will be what CNSers will be interested in trying to find. In other words, the interesting result from a CNS perspective is not a specification of how a complicated set of interactions work, but isolating the core mechanisms that are doing the interacting. And this implies, I believe, trying to unify the various kinds of operations and modules and entities we find (e.g. in a theory like GB) to a very small number of core operations (in the best case just one). DMWWP’s program aims at this level of grain, as does MP and that is why they look like a good fit.

Third, as any MPer knows, FL is not just Merge. There are other operations. It is useful to consider how we might analyze linguistic phenomena that are Merge recalcitrant in these terms. Feature checking and algebraic structures seem made for each other. Maybe memory limitations could undergird something like phases (see DMWWP discussion of a Marcus suggestion on p. 11 that something like phases chunk large trees into “overlapping but incompletely bound subtrees”). At any rate, getting comfortable with the kinds of mental mechanisms extant in other parts of cognition and perception might help linguists focus on the central MP question: what basic operations are linguistically proprietary? One answer is: those operations required in addition to those that other animals have (e.g. time interval determination, ordinal sequencing, chunking, etc.).

This is a good paper, especially so because of where it appears (a very leading brain journal) and because it treats linguistic work as obviously relevant to the CNS of language. The project is basically Marr’s, and unlike so much CNS work, it does not try to shoehorn cognition (including language) into some predetermined conception of neural mechanism which effectively pretends that what we have discovered over the last 60 years does not exist.



[1] DMWWP notes that the real problem is dependencies in an unbounded nested tree. It is not merely the hierarchy, but the unboundedness (i.e. recursion) as well.

Monday, October 14, 2013

Syntax in the brain


It’s not news that syntactic structure is perceptible in the absence of meaningfulness. After all, even though the slithy toves did gyre and gimble in the mabe, still and all colorless green ideas do sleep furiously. Not news maybe, but still able to instruct, as I found out in getting ready for this years Baggett Lectures (see here). The invited speaker this year is Stanislas Dehaene and to get ready I’ve read (under the guidance of my own personal Virgil, Ellen Lau, guide to the several nether circles of neuroscience) some recent papers by him on how the brain does syntax. One (here) was both short (a highly commendable property) and absorbing. Like many neuroscience efforts, this one is collaborative, brought to you by the French team of Pallier, Devauchelle and Dehaene (PDD). I found the PDD results (and methods to the degree that I could be made to appreciate them) thought provoking. Here’s why.

The paper starts from the conventional grammatical assumption that “sentences are not mere strings of words but possess a hierarchical structure with constituents nested inside each other” (2522).[1] PDD’s task is to find out where, if anywhere, the brains tracks/builds this hierarchy. PDD construct a very clever model that allows them to use fMRI techniques to zero in on those regions sensitive to hierarchical structure.

Before proceeding, it’s worth noting that this takes quite a bit of work. Part of what makes the paper fun, is the little model that allows PDD to index hierarchy to differential blood flows (the BOLD (blood-oxygen-level-dependent) response, which is what an fMRI tracks). It predicts a linear relationship between the BOLD response and phrasal size (roughly indexed to word length (a “useful approximation”) and they use this relationship to probe ROIs (i.e. regions of interest (Damn I love these acronyms)) that respond to this predicted relationship using two different kinds of linguistic probes. The first are strings of words containing phrases ranging from 1 to 12 words long (actually 12/6/4/3/2/1 e.g. 12: I believe that you should accept the proposal of your new associate, 4: mayor of the city he hates this color they read their names). The second are jabberwocky strings with the same structure (e.g. I tosieve that you should begept the tropufal of your tew viroate). Here’s what they found:

1.     They found brain regions that responded to these probes in the predicted linear manner. Four in the STS region, and two in the left inferior gyrus (IFGtri and IFGorb).
2.     The regions responded differentially to the two kinds of linguistic probes. Thus, all four regions responded to the first kind of probe (“normal prose”) while jabberwocky only elicited responses in IFGorb (with some response with a lower statistical threshold in left posterior STS and IFGtri).

In sum, different brain regions light up exclusively to phrasal syntax independent of content words. Thus, the brain seems to distinguish contentful morphemes from more functional ones and it does so by processing the relevant information in different regions.
And this is interesting, I believe. Why?

Well, think of what PDD could have found. One possibility is that all words are effectively the same; the only difference between content words and functional vocabulary residing in their statistical frequency, the closed class content words being far more common than the open class contentful vocab. Were this correct, then we should expect no regional segregation of the two kinds of vocab, just, say, a bigger or smaller response based on the lexical richness of the input. Thus, we might have expected all regions to respond equally to all of the inputs though the size of the response would have differed. But this is not what PDD found. What they found is differential activity across various regions with one group of sites responding exclusively to structural input even in the absence of meaningful content. This sure smells a lot like the wetware analogue of the autonomy of syntax thesis (yes, I was not surprised, and yes, I was nonetheless chuffed). PDD (2526) make exactly this autonomy of syntax point in noting that their results underline “the relative independence of syntax from lexico-semantic features.”

Second, the results might have implications for grammar lexicalization (GL) (an idea that Thomas has been posting about recently). From what I can tell, GL sees grammatical dependencies as the byproduct of restrictions coded as features on lexical terminals. Grammatical dependencies on the GL view just are the sum total of lexical dependencies. If this is correct, then a question arises: what does this kind of position lead us to expect about grammatical structure in the absence of such lexical information? I assume that Jabberwocky vocab is not part of our lexicon and so a grammar that exclusively builds structure based on info coded in lexical terminals will not have access to (at least some) grammatically relevant information in Jabberwocky input (e.g. how to combine the and slithy toves in the absence of features on the latter?). Does this mean that we should not be able to build syntactic structure in the absence of the relevant terminals or that we will react differently to input composed of “real” lexemes vs Jabberwocky vocab? Behaviorally, we know that we can distinguish well- from ill- formed Jabberwocky. So we know that the absence of a lot of lexical vocab does not impede the construction of syntactic structure. PDD further shows that neurally there are parts of the brain that respond to syntactic structure regardless of the presence of featurally marked terminals (and recall, that this need not have been the case). A non-GL view of grammar has no problems with this, as grammatical structure is not a by-product of lexical features[2] (at least not in general, only the functional vocab is possibly relevant).[3] I cannot tell if this is a puzzle if one takes a fully GL view of things, but it certainly indicates that parts of the brain seem tuned to structure without much apparent lexical content. Indeed, it suggests (at least to me) that GL might have things backwards: it’s not that grammatical structure arises from lexical specifications but that lexical specifications are by-products of enjoying certain syntactic relations. Formally, this might be a distinction without a difference, but psychologically and neurally, these two ways of viewing the ontogeny of structure look like they might (note the weasel word here) have very different empirical consequences.

I am sure that I have over-interpreted the PDD results. The structure they probe is very simple (just right branching phrase structure) and the results rely on some rather radical simplifications. However, I am told, that this is the current state of the fMRI art and even with these caveats, PDD is interesting and thought provoking, and, who knows, maybe more than a little relevant to how we should think about the etiology of grammar. It seems that brains, or at least parts of brains, are sensitive to structure regardless of what that structure contains. This is an old idea, one perfectly expected given a pretty orthodox conception of the autonomy of syntax. It’s nice to see that this same conception is leading to new and intriguing work investigating how brains go about building structure.



[1] This truism, sadly, is not always acknowledged. For example, it is still possible for psychologists to find a publishing outlet of considerable repute for papers that deny this (see here). Fortunately, flat-earthism seems to be loosing its allure in neuroscience.
[2] Note that this does not entail that grammatical information might not be lexicalized for tasks like parsing. It only implies that grammatical dependencies are more primitive than lexical ones. The former determine the latter, not vice versa.
[3] I say ‘possibly’ as the functional vocab might just be the morphological outward manifestations of grammatical dependencies rather than the elements from which such dependencies are formed. There is not obvious reason for thinking that structure is there because of the functional vocab, though there is reason for assuming that functional vocab and syntactic structure are correlated, sometimes strongly.

Monday, August 26, 2013

Got Culture?


In the last chapter of Dehaene’s Reading the Brain he speculates about one of the really big human questions: whence culture? The books big thesis, concentrating on reading and writing as vehicles for cultural transmission, is the Neuronal Recycling Thesis (NRT). The idea is simple; culture supervenes on neuronal mechanisms that arose to serve other ends. Think exaptation as applied to culture.  Thus, reading and writing are underpinned by proto letters, which themselves live on ecologically natural patterns useful for object recognition.  So too, the hope goes, for the rest of what we think of as culture. However, as Dehaene quickly notes, if this is the source, and “we share most, if not all of these processors [i.e. recycled structures NH] with other primates, why are we the only species to have generated immense and well-developed cultures” (loc 4999). Dehaene has little patience for those who fail to see a qualitative difference between human cultural achievements and those of our ape cousins.

…the scarcity of animal cultures and the paucity of their contents stand in sharp contrast to the immense list of cultural traditions that even the smallest human groups develop spontaneously. (loc 4999)

Dehaene specifically points to the absence of “graphic invention” in primates as “not due to any trivial visual or motor limitation” or to a lack of interest in drawing, apparently (loc 5020). He puts the problem nicely:

If cultural invention stems from the recycling of brain mechanisms that humans share with other primates, the immense discrepancy between the cultural skills of human beings and chimpanzees needs to be explained. (loc 5020)

He also surveys several putative answers, and finds them wanting. His remarks on Tomasello (loc 5046-5067) seem to me quite correct, noting that though Tomasello’s mind reading account might explain how culture might spread and its achievements retained cross generationally:[1]

…it says little…about the initial spark that triggers cultural invention. No doubt the human species is particularly gifted at spreading culture – but it is also the only species to create culture in the first place. (loc 5067, his emphasis)

So what’s Dehaene’s proposal?

My own view is that another singular change was needed - the capacity to arrive at new combinations of ideas and the elaboration of a conscious mental synthesis (loc 5067).

This is quite a mouthful, and so far as I can see, what Dehaene means by this is that our frontal lobe got bigger and that this provided a “”neuronal workspace” whose main function is to assemble, confront, recombine, and synthesize knowledge” (loc 5089).

I don’t find this particularly enlightening. It’s neuro-speak for something happened, relevant somethings always involving the brain (wouldn’t it be refreshing if every once in a while the kidney, liver or heart were implicated!). In other words, the brain got bigger and we got culture. Hmm. This might be a bit unfair. Dehaene does say more.

He notes that the primate cortex, in contrast to ours, is largely modular, with “its own specific inputs, internal structure, and outputs.” Our prefrontal cortex in contrast “emit and receive much more diverse cortical signals” and so “tend to be less specialized.” In addition, the our brains are less “modular” and have greater “bandwidth.” This works to prevent “the division of data and allows out behavior to be guided by any combination of information from past or present experience.” (loc 5089)

Broken down to its essentials, Dehaene is here identifying the demodularization of thought as the key ingredient to the emergence of culture. As he notes (loc 5168), in this he agrees with Liz Spelke (and others) who has argued that the general ability to integrate information across modules is what spices up our thinking beyond what we find in other primates.  Interestingly for my purposes here, Spelke ties this capacity for cross module integration to the development of linguistic facility (see here).

This assumption, that language is a necessary condition for the emergence of the kind of culture we see in humans is consistent with the hypothesis Minimalists have been assuming (following people like Tatersall (here)) that the anthropological “big bang,” which occurred in the last 25-50,000 years, piggy backed on the emergence of FL in the last 50-100,000 years. Moreover, it’s language as module buster that gets the whole amazing culture show on the road.

But what features of language make it a module buster?  What allows grammar to “assemble and recombine” otherwise modular information? What’s the secret linguistic sauce?

Sadly, neither Dehaene nor Spelke say.  Which is too bad as me and my lunch buddies (thx Paul, Bill) have discussed this question off and on for several years now, without a lot to show for it. However, let me try to suggest a key characteristic that we (aka I) believe is implicated. The key is syntax!

The idea is that FL provides a general-purpose syntax for combining information trapped within modules.  Syntax is key here, for I am assuming (almost certainly wrongly, so feel free to jump in at any point) what makes information modular is some feature of the module internal representations that make it difficult for them to “combine” with extra-modular information. I say syntax for once information trapped within a module can combine with information in another module it appears that, more often than not, the combination can be interpreted. Thus, it’s not that the combination of modularly segregated concepts is semantically undigestible, rather the problem seems to be getting the concepts to talk to one another in the first place, and, I take this to mean, to syntactically combine. So module busting will amount of figuring out how to treat otherwise distinct expressions in the same way. We need some kind of abstract feature that, when attached to an arbitrary expression, allows it to combine with any other expression from any other module.  What we need, in effect, is, what Chomsky called, an “edge-feature,”  (EF) a thingamajig that allows expressions to freely combine.

Now, if you are like me, you will not find this proposal a big step forward for it seems to more name a solution than provide one. After all, what can EFs be such that they possess such powers?  I am not sure, but I am pretty confident that whatever this power is it’s purely syntactic. It is an intrinsic property of lexical atoms and it is an inherited property of congeries of such (i.e. outputs of Merge).  I have suggested (here) that EFs are, in fact, labels, which function to close Merge in the domain of the lexical items (LIs). In the same place I proposed that labeling is the distinctively linguistic operation, which in concert with other cognitively recycled operations, allowed for the emergence of FL.

How might labels do this?  Good question. An answer will require addressing a more basic question: what are labels?  We know what they must do: they must license the combination both of lexical atoms and complexes of such.  Atomic LIs are labels.  Complexes of LIs are labeled in virtue of containing atomic ones. The $64,000 question (doesn’t sound like much of a prize anymore, does it?) is how to characterize this.  Stay tuned.

So, culture supervenes on language and language is the recycling of more primitive cognitive operations spiced with a bit of labeling. Need I say that this is a very “personal” (read “extremely idiosyncratic and not currently fashionable”) view?  Current MP accounts are very label-phobic.  However, the question Dehaene raises is a good one, especially for theories like MP that presuppose lots of cognitive recycling.[2]  It’s not one whose detailed answer is anywhere on the horizon. But like all good questions, I suspect that it will have lots of staying power and will provide lots of opportunities for fun conversations.



[1] It’s good to see that Tomasello is capable of begging the interesting question regardless of where he puts his efforts.
[2] See discussion in the comments I had with Jan Koster about this my previous post (here).