Comments

Showing posts with label Frankland and Greene. Show all posts
Showing posts with label Frankland and Greene. Show all posts

Sunday, November 1, 2015

Brains, grammars and hype (part 2)

In the previous post (here), I showed how Frankland and Greene identifies a role sensitive region of cortex and sub-areas within that region that are differentially sensitive to the doer and done-to roles. In other words, if correct, F&G offers a hypothesis about where roles like doer and done-to get coded. Finding a region sensitive to thematic parameters would be a useful contribution given our vast ignorance concerning the brain bases of anything (see here discussed here). Let me repeat this loudly lest I not be heard: FINDING A REGION SENSITIVE TO THEMATIC PARAMETERS WOULD BE A USEFUL CONTRIBUTION GIVEN OUR VAST IGNORANCE CONCERNING THE BRAIN BASES OF ANYTHING. However, F&G claims to do a whole lot more than this. Here I want to consider if it does do more. So the question for what follows: does F&G explain how the brain codes thematic information as it appears to claim to do?

No. Not really. The paper may have identified a region that correlates to role information but F&H’s claim that it explains how brains code such information seems to me quite overblown.[1] Here’s what I mean.

What would it mean to show how brains code such information? F&G tells us. In the abstract, it takes its discovered empirical results to support the following claim:

At a high level, these regions may function like topographically defined data registers, encoding the fluctuating values of abstract semantic variables. This functional architecture, which in key respects resembles that of a classical computer, may play a critical role in enabling humans to flexibly generate complex thoughts.

What’s this mean? Those familiar with earlier critiques of connectionism should recognize the allusions. People like Fodor and Pylyshyn, Marcus, and Gallistel argued that brains had a Turing rather than a connectionist architecture. They provided various arguments for this, including observations about the systematicity of cognition (in particular in language), which makes perfect sense if one assumes that brains embodied read/write memories with variables and valuation of variables, being key elements.  Most of the arguments provided were behavioral (though see Gallistel for more direct arguments that brains cannot be connectionist either). F&G is clearly pointing to these claims in the abstract above (indeed, Fodor and Pylyshyn, Marcus and Pinker are noted in the bibliography in relation to this). So, F&G clearly intends its results to be an argument in favor of Turing architectures and a challenge for connectionist architectures. However, if this is the intent, I don’t see that F&G’s argument adds anything to the earlier behavioral arguments. Why not?

F&G notes that its results are consistent with Turing architectures, but then so are most connectionist models so far as I can tell. There is nothing in these models that prevents the hidden layers (appropriately tuned) from isolating doer and done-to roles. Indeed, this is regularly done in such models for other abstract categories. So, if F&G intends to use its results to argue for classical architectures, then it is unclear to me what it has actually added to the arguments advanced by Fodor & Pylyshyn, Marcus or Gallistel. Note, I have nothing against the conclusion that connectionist architectures are bad neural models (less coyly: I am pretty confident that connectionist architectures suck). What I don’t see is that F&G adds anything to the previous arguments. I would go further (as you probably knew I would). The concluding discussion section of F&G notes that there is a “class of models that use matrix operations to combine spatially distributed representations into conjunctive representations…that could potentially be augmented…[to] encode conjunctive representations for distinct semantic roles” (11737). For the uninitiated, this is connectionist speak. In other words, as F&G notes, its results do not argue against a connectionist conception in favor of a more classical Turing view. Or more correctly, the F&G results do not add anything to the earlier (completely compelling arguments) arguments. So, if F&G intends its “how” contribution to consist in an argument for a classical architecture and against a connectionist one, then, by its own admission, it fails.[2]

What else could the “how” mean? Another possible contrast is between the kinds of codes the brain uses to track information; in particular does the brain use a place code or a rate code to track doers and done-tos. Let me expand a bit.

One line of thinking (that F&G says its results endorse) exploits geography to code information: “functional segregation corresponding to spatial segregation” and binding of variables to values executed by bringing the two into spatial proximity. This contrasts with another view wherein binding is signaled through temporal proximity (synchronization) rather than spatial. F&G claims that its results (my emphasis)

suggest that such temporal correlations may be unnecessary in this case because the bindings may instead be encoded through the instantiation of distributed patterns of activity in spatially dissociable patches of cortex devoted to representing distinct semantic variables” (11736). 

However as the paper notes, and the highlighted mealy-mouthed modals indicate, this conclusion is not particularly well supported by their experiments. Or, more correctly, F&G’s tools preclude a strong choice between the two.  As F&G notes, the hunt was conducted using fMRI and because these have limited temporal resolution (on the order of 1000 ms) fMRI probes cannot generally “see” rate codes. The best that F&G can conclude is that because it was able to localize roles in geographically proximate yet distinct locals this suggests that a place coding of roles might be right, though not to the exclusion of rate codes. The logic is that place codes require segregated (proximate?) geography and this was found. Hence the finding supports the claim that for role information the brain uses a place code. But this conclusion does not follow. To establish it firmly one needs the inverse: if segregated regions then place code. But this is not obviously true. Moreover, and here I am asking, do neuro people believe that anytime they can localize functions in different (nearby) places that this is evidence for place codes? Sounds wrong to me, but, hey, I don’t do this.[3]

I should add that the second experiment is the crucial one for this conclusion, and it is less robust than the first as F&G notes. The bifurcation of lmSTC into doer and done-to areas is quite subtle empirically and some of the participants in the UMD discussion thought that the data here was quite brittle. Again, this is beyond my pay grade.

F&G, then, really says very little (if anything) about the how question. In fact, it never really addresses it except tangentially. “Where?,” not “how?”, is what F&G addresses.  Let me squawk about this for a moment.

IMO, neuro types often confuse how does X work with where is X located. Why they think answering one answers the other I do not know. I don’t object to the claim that knowing where things are in the brain might be/is likely to be a good first step in figuring out how the brain does what it does. But reading F&G (and this paper is hardly unique) leads me to think that CNers can’t tell the difference between where and how.  And this is a problem.

One consequence of the confusion is that it denigrates the cognitive work that it presupposes. F&G relies on an unanalyzed conception of thematic roles. In fact, it relies on a truism: that sentences like John saw Mary do not mean the same as Mary saw John and that the difference has something to do with the fact that what sentences say about John/Mary in the first sentence is effectively reverses what the second sentence says about them.  This is a truism, or as close to one as might be imagined.  However, as any linguist knows, there are many different theories to explain how this truism is true. Some exploit theta roles, some grammatical roles, some the internal/external distinction, some first vs second merge, some predicate argument structure with 1st and 2nd argument positions of a predicate, some Deep Structures, some kernel sentences, etc. When a linguist asks how is thematic information represented, s/he means how can we distinguish between these apparently different conceptions all of which code/represent the observed doer/done-to difference. F&G cannot tell us which of these is right, nor does it intend to. This “how?” question is beyond the technical reach of current neuro apparatus.  That’s not a criticism. Here is the criticism: by confusing where with how, F&G continues the tradition of treating distinctions beyond the range of its probes as non-questions, rather than as questions beyond the resolution of its methods. The fact is that cognitive probes into the structure of brains is right now far more powerful than the currently most fashionable technology in neuro-science. fMRI might generate pretty pictures, but it's a pretty coarse technology. Right now, behavioral methods generally allow us to probe brain structure in a far more refined way than neuro methods do. That CN technology cannot usefully probe well motivated behaviorally based claims is what we should expect, and is what we find.

A second feature of the where/how confusion is that it leads one to abstract away from the most serious question in the neuro-sciences. Call it Gallistel’s question: how do brains embody mental constructs?  For example, how does wetware code for a variable or a value thereof? How do brains read and write to memory, bind a variable, distinguish between types and tokens?  Nobody knows. In fact, as Gallistel has observed, most CNers don’t even understand that this is the “how?” question that needs addressing (see here for discussion). The cognitive literature, including that in linguistics, has shown that we need these notions. Much of current neuroscience assumes that brain architectures that cannot do any of this (indeed that apparently deny, if Gallistel is right, that brains ever do this) are serviceable. This is partly abetted by the fact that current thinking fails to distinguish where from how. F&G is another example of this wider confusion.

I could go on, but I won’t. F&G makes a contribution: it identifies one possible place for where role information in some sense (however it is represented and whether it is specifically linguistic or not) might live. Given the current state of neuroscience, this is not nothing. However, the paper’s rhetoric (BS really) is way over the top. The introduction and conclusion motivate the investigation by pointing to really big issues (in particular recursion and Turing architecture). It purports to address these issues but in truth it can’t. The results are neutral wrt them. In the process, F&G sows lots of confusion and makes lots of simple errors thereby makind it hard to find the useful kernel in the morass. This leads me to one final observation.

I have heard it argued that without the overstatement and the BS the paper could never have been published. This is sometimes said in apparent justification of the BS and hype. If so neuroscience is in really bad shape. Moreover, I am skeptical that the hype is necessary, though I am sure that even if it is, it is odious to sling it nonetheless. Let me vent.

First, I doubt that a more measured presentation would have prevented publication. The result is not trivial and could have been presented as relevant to finding where linguistically/conceptually important concepts live in brain tissue.

Second, wanting to get published is no excuse for BS. This is not show business. BS goes against the fundamental values of the scientific enterprise and should not be tolerated, even if it might be useful career-wise.[4] The big problem is that such BS is fast becoming part of standard practice.  And like all S it greases a slippery slope: BS facilitates publication, we become more indulgent towards it and this will serve to further BSify research and publication. There is no excuse for this, or at least not one that should pass the smell test (and BS does smell). Whatever, F&G has told us about brains, it is mired in overstatement and self promotion. That’s the main reason many have reacted so strongly, and rightly so.[5] And that’s too bad because F&G does have something to tell us of interest.



[1] Steve Pinker’s tweet highlights these F&G ambitions as well. It reads: “The most important paper in cognitive neuroscience in many years: How does the brain represent who did what to whom.” Note the “how.” I wonder if the tweet would have had the same impact if we replaced ‘how’ with ‘where.’ I can’t tell, though I think that the howish version sounds far more interesting. And this is exactly the problem.
[2] In the discussion section, F&G observes relations between its results and some previous findings in the literature. An interesting one relates to deficit studies that identify insult to the lmSTC results in “who did what to whom” problems for stimuli presented aurally and visually. This suggests the possibility, as F&G note, that this area is not linguistically dedicated. In other words, this area might be part of an “amodal language of thought.” If this is so, it might be interesting to see if analogous areas in non-linguistically endowed animals can similarly discriminate doers from done-tos.  This might even have some interesting linguistic significance concerning the theoretical utility of theta roles as discussed here. F&G leaves the linguistic status of lmSTC for future research. Hope it gets done.
[3] Also, how important is the proximity? Say that doers were found in one area and done-tos were found several sulci away. Would this be a problem for place codes? I don’t know. At any rate, the relation between being localizable and being place coded strikes me as looser than F&G suggests. In fact, I could imagine that even were rate codes employed to code some functional feature the sources generating the relevant rates might nonetheless localize somewhat.  I don’t know that this is so, but nothing F&G says leads me to think that this is impossible or even false. So a question to cognoscenti: is this inference from localizable to place code legit?
[4] IMO, BS is the most corrosive feature of much current research. As Frankfurt has argued, it might be even worse than lying for unlike the latter it has no regard for truth whatsoever. Stan Dehaene was the editor for the paper and he should really have removed this BS from the paper. He knows better.
[5] BTW, F&G does not get its BS right either. See the box marked “Significance” on the first page of the paper. It suggests that the problem of theta roles is the same as the problem of recursion. This is false. The roles that F&G addresses have nothing to do with Humboldt’s making infinite use of finite means. Here we have a finite set of possible sentences templatically specifiable wrt roles of two arguments. Recursion gives you sentences with many doers and many done tos, in fact unboundedly many. F&G has nothing to say about where the brain codes this.

Sunday, October 25, 2015

Brains, grammars and hype (part 1)

This recent neuro-ling paper by Frankland and Greene (F&G) in PNAS has generated a lot of critical comment by linguists, and rightly so.  The paper demonstrates several unfortunate flaws in the thinking (and moral character (I return to this)) of cog-neuro (CN) types. The most startling, intellectually speaking, is CNers deep seated dualism. It appears (if judged by their writing rather than their occasional denunciation of ghosts and souls) that CNers do not believe in the identity thesis (i.e. minds and brains are the same thing). In fact, they seem to believe that behavioral evidence, no matter how subtle or empirically well grounded or replicable or statistically significant or robust or of large effect size or…is inherently incapable of doing much to advance our understanding about how brains are organized. The only real evidence, on this view, comes from fMRI/MEG/ etc. brain studies. Thus, mentalistic investigations of brains (behavioral, cognitive, psychological) cannot possibly inform us about how brains are structured. Talk about dualism! Not even Descartes would have been caught dead saying such things. But for many CNers it’s either show them the meat (literally) or take a hike.

This is quite evident in the F&G paper, which is why it has generated such push back from linguists. Angelika Kratzer’s reaction (here) is right on the mark. Let me quote her:

One quote (attributed to Steven Frankland) in the Harvard Gazette article may point to the source of the communication problem: “This [the systematic representation of agents and themes/patients, A.K.] has been a central theoretical discussion in cognitive science for a long time, and although it has seemed like a pretty good bet that the brain works this way, there’s been little direct empirical evidence for it.” This quote makes it appear as if the idea that the human mind systematically represents agents and themes/patients has had the mere status of a bet before the distinction could be actually localized in the brain. That the distinction is systematically represented in all languages of the world is not given the status of a fact in this quote - it doesn't count as  "empirical evidence". It's like denying your pulse the status of a fact before we can localize the mechanisms that regulate it in the brain.

Yup.[1] As Angelika notes, this quote presupposes that behavioral work, in this case linguistic work, no matter how extensive, does not even rise to the level of “evidence” about brain structure. It seems that minds are one thing and brains another. Dualism anyone?

At any rate, take a look at Angelika’s piece. It makes this point well. Consequently, I thought I would try to do something uncharacteristic in what follows. Instead of zeroing in on the inanities (which, to repeat, are many, and which I will be unable to refrain from mentioning from time to time), I would like to zero in on the substantive contribution F&G makes to our understanding of the brain bases of language competence. But read this for yourself. I am no expert in these matters. However, I am channeling the wisdom of others in what follows. The UMD ling dept congregated recently to discuss the paper and I left convinced that despite the over blown rhetoric (I will return to this) and false advertising (I return to this too) the paper does make a modest contribution. And in the spirit of babies and bathwaters I will try to outline what this might be.  It goes without saying, but I will say it nonetheless, that I am no expert in these matters and I am relying on the knowledge of others here, and I might have screwed things up in translation (and if so, I hope others will chime in), but with all these caveats, here’s why I think that the paper is not just wrongheaded (though it is that too), but makes a possible contribution to our understanding of a very difficult topic. [2]

F&G is a fishing expedition of a kind that we have seen before. For example, Pallier, Devauchelle and Dehaene (discussed here) do something similar in hunting for where the Merge operation lives in the brain. F&G are hunting for brain correlates of “thematic” (yes these are scare quotes, I return to this) information; specifically, “whether and how” (p.11732) the brain codes the “who did what to whom” information that sentence’s express.

The “whether” part of the question, linguists rightly believe, has been already well-established. The most generous reading of F&G is that it agrees but notes that this still leaves open three questions: (1) can we find more neuro based indices of this well established cognitive fact (i.e. fMRI or MEG/EEG or lesion data), (2) can we localize these “thematic” brain effects and (3) what might such localization tell us about how brains code this information. IMO, the most interesting features of F&G regards the first two questions, for which it provides tentative answers. What are these?

F&G conducts two kinds of experiments. The first “identifies a broad region,” the left medial Superior Temporal Cortex (aka: lmSTC) that is able to reliably distinguish sentences that express the same theta information. For example, it can distinguish the sentence pairs John kicked Mary/Mary was kicked by John from Mary kicked John/John was kicked by Mary. By “averaging” over the active and passive pairs, the experiment zeros in on the doers and done-tos and abstracts away from surface syntax.  At the least the experiment shows that this region is not sensitive to just the words involved as these are held constant in the contrasting pairs. What matters is the “thematic” structure.

How well does lmSTC do in distinguishing these contrasts? It succeeds about 57% of the time. By ling standards this is really not enough. After all, humans succeed about 100% of the time. Thus, we need to explain why a region that fails 43% of the time to correctly distinguish what speakers never fail to distinguish (and this is the kind of thing that speakers are virtually perfect at) nonetheless is the brain basis of this overt behavioral capacity.

And there is a ready possible account: the resolution of the fMRI probe is not good enough to eliminate interfering noise and this noise is what reduces discrimination to a mere 57%. However, as the area discriminates above chance then it is a reasonable guess that it is not only sensitive to thematic distinctions, but is where these distinctions get coded and we would see this yet more clearly were we able to get an even finer probe.[3]

So, lmSTC tracks thematic information.  Before going on, I should add that F&G identifies another area that tracks this information at with roughly the same accuracy (the right posterior insula/extreme capsule region (call this region R-2)). However, the paper treats the response of this region as (at best) secondary and most likely not relevant. Why? Because whereas the lmSTC predicts further downstream brain responses, the second area does not.  The neuro crowd at our little UMD discussion got all hot about this wrinkle, so let me tell you a bit about it.

F&G shows that there is a correlation between responses to sentences in lsSTC and the amygdala, where affective responses are apparently evoked. The paper shows that the amygdala gets all hyped up to sentence pairs like The grandfather kicked the baby/the baby was kicked by the grandfather but not to The baby kicked the grandfather/the grandfather was kicked by the baby. Why the differential response? Because the amygdala doesn’t like it when babies are badly treated by wicked granddads but thinks that babies kicking old folks is not really very bad (after all how much harm could a baby kick do?). F&G interprets this, reasonably enough, as showing that the info extracted in lmSTC is used by the amygdala in responding. R-2 shows no such correlation. So whatever is going on there does not correlate with downstream amygdala responses. F&G concludes that R-2 “failed to meet additional minimal functional criteria for encoding sentence meaning” (11733). From what I can tell, the only criteria it failed to code is this downstream impact. Make of this what you will. From where I sit, it does not imply that the same thematic distinctions are not coded in R-2 as in lmSTC, only that the active use of this information pipelines directly from the latter to the amygdala but not the former. But, the CNers really liked this, so I offer it to you for your appreciation.

The second experiment zeros in on the fine structure of lmSTC. In particular, it aims to see whether and where in this already identified region doers and done-tos are coded. F&G does this, again, by seeing how the region’s discrimination powers generalize. How do the subparts of the region react to doers and done-tos as such. Here’s how F&G describes the procedure for isolating doers and done-tos as such in lmSTC :

For our principal searchlight analyses, four-way classifiers were trained to identify the agent or patient using data generated by four out of five verbs. The classifiers were then tested on data from sentences containing the withheld verb. For example, the classifiers were tested using patterns generated by “th dog chased the man,” having never previously encountered patterns generated by sentences involving “chased,” but having been trained to identify “dog” as the agent and “man” as the patient” in other verb contexts. …Thus, this analysis targets regions that instantiate consistent patterns of activity for (for example) “dog as agent” across verb contexts, discriminable from “man as agent” …A region that carries this information therefore encodes “who did it?” across nouns and verb contexts tested. (11734).

It turns out that two proximate yet distinct parts of lmSTC seem to discriminate doers from done-tos. Actually, the results for done-tos is cleaner. The borders of the doer region is muddier (moreover, active and passives of the same roles don’t function quite in parallel).[4] At any rate, F&G conclude that the lmSTC spatially bifurcates the two roles, and, at least to me (and more importantly the neuro-psycho people in the UMD discussion group), this conclusion seems reasonable given the data.

That’s what F&G shows: if correct, it identifies a role sensitive region and areas within that region differentially sensitive to the doer and done-to roles. In other words, if correct, F&G offers a hypothesis about where roles get coded.  But F&G claims to do a whole lot more. Does it? We return to this in the next post.



[1] Lest you think that Frankland indulged in hyperbole for the delectation of the newshounds alone, the same sentiment permeates the PNAS piece.
[2] Thanks particularly to Ellen Lau for organizing this and to Allyson Ettinger for a vigorous defense of the paper. I have shamelessly stolen all that I could from the excellent discussion.
[3] Note that this “guess” is quite a bit more precarious than the “bet” that who did what to whom info is represented in brains. There is no doubt that brains code for such information, and this is much more solid than the proposal that it gets coded in lmSTC. This just reiterates Angelika’s apt remarks above.
[4] F&G discusses why not, but the discussion is pretty inconclusive. If lmSTC exclusively tracks “thematic” information then this results is clearly unexpected.d