Comments

Showing posts with label rationalism. Show all posts
Showing posts with label rationalism. Show all posts

Monday, May 23, 2016

The return of behaviorism

There is a resurgence of vulgar Empiricism (E). It’s rampant now, but be patient, it will soon die out as the groundless extravagant claims made on its behalf will soon be seen to, yet again, prove sterile. But it is back and getting airing in the popular press.

Of the above, the only part that is likely difficult to understand is what I intend by ‘vulgar.’ I am not a big fan of the E-weltanschauung, but even within Empiricism there are more and less sophisticated versions. The least sophisticated in the mental sciences is some version of behaviorism (B). What marks it out as particularly vulgar? Its complete repudiation of mental representations (MR). Most of the famous E philosophers (Locke and Hume for example) were not averse to MRs. They had no problem believing that the world, through the senses, produces representations in the mind and that these representations are causally implicated in much of cognitive behavior. What differentiates classical E from classical Rationalism (R) is not MRs but the degree to which MRs are structured by experience alone. For E, the MR structure pretty closely tracks the structure of the environmental input as sampled by the senses. For R, the structure of MRs reflects innate properties of the mind in combination with what the senses provide of the environmental landscape. This is what the debate about blank/wax tablets is all about. Not whether the mind has MRs but whether the properties of the MRs we have reduce to sensory properties (statistical or otherwise) of the environment. Es say ‘yes,’ Rs ‘no.’

Actually this is a bit of a caricature. Everyone believes that the brain/mind brings something to the table. Thus, nobody thinks that the brain/mind is unstructured as such brains/minds cannot generalize and everyone believes that brains/minds that do not generalize cannot acquire/learn anything. The question then is really how structured is the brain/mind. For Es the mind/brain is largely a near perfect absorber of environmental information with some statistical smoothing techniques thrown in. For R extracting useful information from sensory input requires a whole lot of given/innate structure to support the inductions required. Thus, for Es the gap between what you perceive and what you acquire is pretty slim, while for Rs the gap is quite wide and bridging this gap requires a lot of pre-packaged knowledge. So everyone is a nativist. The debate is what kinds of native structure is imputed.

If this is right, the logical conclusion of E is B. In particular, in the limit, the mind brings nothing but the capacity to perfectly reflect environmental input to cognition. And if this is so, then all talk of MRs is just a convenient way of coding environmental input and its statistical regularities. And if so, MRs are actually dispensable and so we can (and should) dump reference to them. This was Skinner’s gambit. B takes all the E talk of MRs as theoretically nugatory given that all MRs do is recapitulate the structure of the environment as sampled by the senses. MRs, on this view, are just summaries of experience and are explanatorily eliminable. The logical conclusion, the one that B endorses, is to dump the representational middlemen (i.e. MRs) that stand between the environment and behavior. All the brain is, on this view, is a way of mapping between stimulus inputs and behavior, all the talk of MRs just being misleading ways of talking about the history of stimuli. Or, we don’t need this talk of minds and the MR talk it suggests, we can just think of the brain as a giant I/O device that “somehow” maps stimuli to behaviors.

Note, that without representations there is no real place for information processing and the computer picture of the mind. Indeed, this is exactly the point that critics of E and B have long made (e.g. Chomsky, Fodor, and Gallistel to name three of my favorites). But, of course the argument can be aimed in the reverse direction (as Jerry Fodor sagely noted someone’s modus ponens can be someone else’s modus tollens): ‘If B then the brain does not process information’ (i.e. the opposite of ‘If the brain processes info then not B’). And this is what I mean by the resurgence of vulgar E. B is back, and getting popular press.

 Aeon has a recent piece against the view of the brain as an information processing device (here). The author is Robert Epstein. The view is B through and through. The brain is just a vehicle for pairing inputs with behaviors based on reward (no, I am not kidding).  Here is the relevant quote (13) :

As we navigate through the world, we are changed by a variety of experiences. Of special note are experiences of three types: (1) we observe what is happening around us (other people behaving, sounds of music, instructions directed at us, words on pages, images on screens); (2) we are exposed to the pairing of unimportant stimuli (such as sirens) with important stimuli (such as the appearance of police cars); (3) we are punished or rewarded for behaving in certain ways.

No MRs mediate input and output. I/O is all there is. 

Misleading headlines notwithstanding, no one really has the slightest idea how the brain changes after we have learned to sing a song or recite a poem. But neither the song nor the poem has been ‘stored’ in it. The brain has simply changed in an orderly way that now allows us to sing the song or recite the poem under certain conditions. When called on to perform, neither the song nor the poem is in any sense ‘retrieved’ from anywhere in the brain, any more than my finger movements are ‘retrieved’ when I tap my finger on my desk. We simply sing or recite – no retrieval necessary (14).

No need for memory banks or MRs. All we need is “the brain to change in an orderly way as a result of our experiences” (17). So sensory inputs, rewards, behavioral outputs. And the brain? That organ that mediates this process. Skinner must be schepping nachas!

Let me end with a couple of references and observations.

First, there are several very good long detailed critiques of this Epstein piece out there (Thx to Bill Idsardi for sending them my way). Here and here are two useful ones. I take heart in these quick replies for it seems that this time around there are a large number of people who appreciate just how vulgar B conceptions of the brain are. Aeon, which published this piece, is, I have concluded, a serious source of scientific disinformation. Anything printed therein should be treated with the utmost care, and, if it is on cog-neuro topics, the presumption must be that it is junk. Recall that Vyvyan Evans found a home here too. And talk about junk!

Second, there is something logically pleasing about articles like Epstein’s; they do take an idea to its logical conclusion. B really is the natural endpoint of E. Intellectually, it’s vulgarity is a virtue for it displays what much E succeeds in hiding. Critics of E (especially Randy and Jerry) have noted its lack of fit with the leading ideas of computational approaches to neuro-cognition. In an odd way, the Epstein piece agrees with these critiques. It agrees that the logical terminus of E (i.e. B) is inimical with the information processing view of the brain. If this is right, the brain has no intrinsic structure. It is “empty,” a mere bit of meat serving as physiological venue for combining experience and reward with an eye towards behavior. Randy and Jerry and Noam (and moi!) could not agree more. On this behaviorist view of things the brain is empty and pretty simple. And that’s the problem with this view. The Epstein piece has the logic right, it just doesn’t recognize a reductio, no matter how glaring.

Third, the piece identifies B’s fellow travellers. So, not surprisingly embodied cognition makes an appearance and the piece is more than a bit redolent of connectionist obfuscation. In the old days, connectionists liked to make holistic pronouncements about the opacity of the inner workings of the neural nets. This gave it a nice anti-reductionist feel and legislated questions about how the innards of the system worked unaskable. It gave the whole theory a kind of new age, post-modern gloss with an Aquarian appeal. Well, the Epstein piece assembles the same cast of characters in roughly the same way.


Last observation: the critiques I linked to above both dwell on how misinformed this piece is. I agree. There is very little argumentation and what there is, is amazingly thin. I am not surprised, really. It is hard to make a good case for E in general and B in particular. Chomsky’s justly famous review of Skinner’s Verbal Behavior demonstrated this in detail. Nonetheless, E is back. If this be so, for my money, I prefer the vulgar forms, the ones that flaunt the basic flaws. And if you are looking for a good version of a really bad set of Eish ideas, the Epstein article is the one for you.

Monday, March 14, 2016

The deep difference between acceptable and grammatical

Up comes a linguist in the street interviewer and asks: “So NH, what would grad students at UMD find to be one of your more annoying habits?” I would answer: my unrelenting obsession with forever banishing from the linguistics lexicon the phrase “grammaticality judgment.” As I never tire of making clear, usually in a flurry of red ball-point scribbles and exclamation marks, the correct term is “acceptability judgment,” at least when used, as it almost invariably is, to describe how speakers rate some bit of data. “Acceptability” is the name of the scale along which such speaker judgments array. “Grammaticality” is how linguists explain (or partly explain) these acceptability judgments. Linguists make grammaticality judgments when advancing one or another analysis of some bit of acceptability data. I doubt that there is an interesting scale for such theoretical assessments.

Why the disregard for this crucial difference among practicing linguists? Here’s a benign proposal. A sentence’s acceptability is prima facie evidence that it is grammatical and that a descriptively adequate G should generate it. A sentence’s unacceptability is prima facie evidence that a descriptively adequate G should not generate it. Given this, using the terms interchangeably is no big deal. Of course, not all facies are prima and we recognize that there are unacceptable sentences that an adequate G should generate and that some things that are judged acceptable nonetheless should not be generated. We thus both recognize the difference between the two notions, despite their intimate intercourse, and interchange them guilt free.

On this benign view, my OCD behavior is simple pedantry, a sign of my inexorable aging and decline. However, I recently read a paper by Katz and Bever (K&B) that vindicates my sensitivities (see here), which, of course, I like very much and am writing to recommend to you (I would nominate it for classic status).[1] Of relevance here, K&B argues that the distinction between grammaticality and acceptability is an important one and that blurring it often reflects the baleful influence of that most pernicious intellectual habit of mind, EMPIRICISM! I have come to believe that K&B is right about this (as well, I should add, about many other things, though not all). So before getting into the argument regarding acceptability and Empiricism, let me recommend it to you again. Like an earlier paper by Bever that I posted about recently (here), this is a Whig History of an interesting period of GG research. There is a sustained critical discussion of early Generative Semantics that is worth looking at, especially given the recent rise of interest in these kinds of ideas. But this is not what I want to discuss here. For the remainder, let me zero in on one or two particular points in K&B that got me thinking.

Let’s start with the acceptability vs grammaticality distinction. K&B spend a lot of time contrasting Chomsky’s understanding of Gs and Transformations with Zellig Harris’s. For Harris, Gs were seen as compact ways to cataloguing linguistic corpora. Here is K&B (15):

 …grammars came to be viewed as efficient data catalogues of linguistic corpora, and linguistic theory took the form of a mechanical discovery procedure for cataloguing linguistic data.

Bloomfieldian structuralism concentrated on analyzing phonology and morphology in these terms. Harris’s contribution was to propose a way of extending these methods to syntax (16):

Harris’s particular achievement was to find a way of setting up substitution frames for sentences so that sentences could be grouped according to the environments they share, similar to the way that phonemes or morphemes were grouped by shared environments…Discourse analysis was…the product of this attempt to extend the range of taxanomic analysis beyond the level of immediate constituents.

Harris proposed two important conceptual innovations to extend Structuralist taxonomic techniques to sentences; kernel sentences and transformations. Kernels are a small “well-defined set of forms” and transformations, when applied to kernels, “yields all the sentence constructions of the language” (17). The coocurrence restrictions that lie at the heart of the taxanomy are stated at the level of kernel sentences. Transformations of a given kernel define an equivalence class of sentences that share the same discourse “constituency.” K&B put this nicely, albeit in a footnote (16:#3):

In discourse analysis, transformations serve as the means of normalizing texts, that is of converting the sentences of the text into a standard form so that they can be compared and intersentence properties [viz. their coocuurences, NH] discovered.

So, for Harris, kernel sentences and transformations are ways of compressing a text’s distributional regularities (i.e. “cataloguing the data of a corpus” (12)).[2]

This is entirely unlike the modern GG conception due to Chomsky, as you all know. But in case you need a refresher, for modern GG, Gs are mental objects internalized in brains of native speakers and which underlie their ability to produce and understand an effectively unbounded number of sentences, most of which have never before been encountered (aka; linguistic creativity). Transformations are a species of rule these mental Gs contain that map meaning relevant levels of G information to sound relevant (or articulator relevant) levels of G information. Importantly, on this view, Gs are not ways of characterizing the distributional properties of texts or speech. They are (intended) descriptions of mental structures.

As K&B note, so understood, much of the structure of Gs is not surface visible. The consequence?

The input to the language acquisition process no longer seems rich enough and the output no longer simple enough for the child to obtain its knowledge of the latter by inductive inferences that generalize the distributional regularities found in speech. For now the important properties of the language lie hidden beneath the surface form of sentences and the grammatical structure to be acquired is seen as an extremely complex system of highly intricate rules relating the underlying levels of sentences to their surface phonetic form. (12)

In other words, once one treats Gs as mental constructs the possibility of an Empiricist understanding of what lies behind human linguistic facility disappears as a reasonable prospect and is replaced by a Rationalist conception of mind. This is what made Chomsky’s early writings on language so important. They served to discredit empiricism in the behavioral sciences (though ‘discredit’ is too weak a word for what happened). Or as K&B nicely summarize matters (12):

From the general intellectual viewpoint, the most significant aspect of the transformationalist revolution is that it is a decisive defeat of empiricism in an influential social science.  The natural position for an empiricist to adopt on the question of the nature of grammars is the structuralist theory of taxanomic grammar, since on this theory every property essential to a language is characterizable on the basis of observable features of the surface form of its sentences. Hence, everything that must be acquired in gaining mastery of a language is “out in the open”; moreover, it can be learned on the basis of procedures for segmenting and classifying speech that presupposes only inductive generalizations from observable distributional regularities. On the structuralist theory of taxanomic grammar, the environmental input to language acquisition is rich enough, relative to the presumed richness of the grammatical structure of the language, for this acquisition process to take place without the help of innate principles about the universal structure of language…

Give up the idea that Gs are just generalizations of the surface properties of speech, and the plausibility of Empiricism rapidly fades. Thus, enter Chomsky and Rationalism, exit taxonomy and Empiricism.

The shift from the Harris Structuralist, to the Chomsky mentalist, conception of Gs naturally shifts interest to the kinds of rules that Gs contain and to the generative properties of these rules. And importantly, from a rule-based perspective it is possible to define a notion of ‘grammaticality’ that is purely formal: a sentence is grammatical iff it is generated by the grammar. This, K&B note is not dependent on the distribution of forms in a corpus. It is a purely formal notion, which, given the Chomsky understanding of Gs, is central to understanding human linguistic facility. Moreover, it allows for several conceptions of well-formedness (phonological, syntactic, semantic etc.) that together contribute along with other factors to a notion of acceptability, but are not reducible to it. So, given the rationalist conception, it is easy and natural to distinguish various ingredients of acceptability.

A view that takes grammaticality to just be a representation of acceptability, the Harris view, finds this to be artificial at best and ill-founded at worst (see K&B quote of Harris p. 20). On a corpus-based view of Gs, sentences are expected to vary in acceptability along a cline (reflecting, for example, how likely they are to be found in a certain text environment). After all, Gs are just compact representations of precisely such facts. And this runs together all sorts factors that appear diverse from the standard GG perspective. As K&B put it (21):

Statements of the likelihood of new forms occurring under certain conditions must express every feature of the situation that exerts an influence on likelihood of occurrence. This means that all sorts of grammatically extraneous features are reflected on a par with genuine grammatical constraints. For example, complexity of constituent structure, length of sentences, social mores, and so on often exerts a real influence on the probability that a certain n-tuple of morphemes will occur in the corpus.

Or to put this another way: a Rationalist conception of G allows for a “sharp and absolute distinction between the grammatical and the ungrammatical, and between the competence principles that determine the grammatical and anything else that combines with them to produce performance” (29). So, a Rationalist conception understands linguistic performance to be a complex interaction effect of discrete interacting systems. Grammaticality does not track the linguistic environment. Linguistic experience is gradient. It does not reflect the algebraic nature of the underlying sub-systems. ‘Acceptability’ tracks the gradiance, ‘grammaticality’ the discrete algebra. Confusing the two threatens a return to structuralism and its attendant Empiricism.

Let me mention one other point that K&B makes that I found very helpful. They outline what a Structuralist discovery procedure (DP) is (15). It is “explicit procedures for segmenting and classifying utterances that would automatically apply to a corpus to organize it in a form that meets” four conditions:

1.     The G is a hierarchy of classes; lower units being temporal segments of speech event, the higher are classes or sequences of classes.
2.     The elements of each level are determined by their distributional features together with their representations at the immediately lower level.
3.     Information in the construction of a G flows “upward” from level to level, i.e. no information at a higher level can be used to determine an analysis at a lower level.
4.     The main distributional principles for determining class memberships at level Li are complementary distribution and free variation at level Li-1.

Noting the structure of a discovery procedure (DP) in (1-4) allows us to appreciate why Chomsky stressed the autonomy of levels in his early work. If, for example, the syntactic level is autonomous (i.e. not inferable from the distributional properties of other levels) then the idea that DPs could be adequate accounts of language learning evaporates.[3] And once one focuses on the rules relating articulation and interpretation the plausibility of a DP for language with the properties in (1-4) becomes very implausible, or, as K&B nicely put it (33):

Given that actual speech is so messy, heterogeneous, fuzzy and filled with one or another performance error, the empiricist’s explanation of Chomskyan rules, as having been learned as a purely inductive generalization of a sample of actual speech is hard to take seriously to say the very least.[4]

So, if one is an Empiricist, then one will have to deny the idea of a G as a rule based system of the GG variety. Thus, it is no surprise that Empiricists discussing language like to emphasize the acceptability gradients characteristic of actual speech. Or, to put this in terms relevant to the discussion above, why Empiricists will understand ‘grammaticality’ as the limiting case of ‘acceptability.’

Ok, this post, once again, is far too long. Look at the paper. It’s really good and useful. It also is a useful prophylactic against recurring Empiricism and, unfortunately, we cannot have too much of that.



[1] Sadly, pages 18-19 are missing from the online version. It would be nice to repair this sometime in the future. If there is a student of Tom’s at U of Arizona reading this, maybe you can fix it.
[2] I would note that in this context corpus linguistics makes sense as an enterprise. It is entirely unclear whether it makes any sense once one gives up this structuralist perspective and adopts a Chomsky view of Gs and transformations. Furthermore, I am very skeptical that there exist Harris-like regularities over texts, even if normalized to kernel sentences. Chomsky’s observation that sentences are not “stimulus bound” if accurate (and IMO they are) undermine the view that we can say anything at all about the distributions of sentences in texts. We cannot predict with any reliability what someone will say next (unless, of course, it is your mother), and even if we could in some stylized texts, it would tell us nothing about how the sentence could be felicitously used. In other words, there would be precious little generalizations across texts. Thus, I doubt that there is any interesting statistical regularities regarding the distribution of sentences in texts (at least understood as stretches of discourse).
            Btw, there is some evidence for this. It is well known that machines trained on one kind of corpus do a piss poor job of generalizing to a different kind of corpus. This is quite unexpected if figuring out the distribution of sentences in one text gave you a good idea of what would take place in others. Understanding how to order in a restaurant or make airline reservations does not carry over well to a discussion of Trump’s (and the rest of the GOP’s) execrable politics, for example. A long time ago, in a galaxy far far away I very polemically discussed these issues in a paper with Elan Dresher that still gives me chuckles when I read it. See Cognition 1976, 4, pp.32l‑398.
[3] That this schema looks so much like those characteristic of Deep Learning suggests that it cannot be a correct general theory of language acquisition. It just won’t work, and we know this because it was tried before.
[4] That speech is messy is a problem, but not the only problem. The bigger one is that there is virtually no evidence in the PLD for may of G properties (e.g. ECP effects, island effects, binding effects etc.). Thus the data is both degenerate and deficient.

Tuesday, March 8, 2016

Bever's Whig History of GG

I love Whig History (WH). I have even tried my hand at it (here, here, here, here). What sets them apart from actual history is that they abstract away from the accidents of history and, if successful, reveal the “inner logic” of historical events. A Whig History’s conceit is that it outlines how events should have unfolded had they been guided by rational considerations. We all know that these are never all that goes on, but scientists hope that this goes on often enough, if not at the individual level, then at the level of the discipline as a whole. Even this might be debated (see here), but to the degree that we can fashion a WH, to that degree we can rationally guide inquiry by learning from our past mistakes and accomplishments. It’s a noble hope, and I am a fervent believer.

Given this, I am always on the lookout for good rational reconstructions of linguistic history. I recently came across a very good one by Tom Bever that I want to make available (here). Let me mark a few of my personal favorite highlights.

1.     The paper starts with a nice contrast between Behaviorist (B) vs R methodologies:

The behaviorist prescribes possible adult structures in terms of a theory of what can be learned; the rationalist explores adult structures in order to find out what a developmental theory must explain.

Two comments: First, throughout the paper, TB contrasts B with R. However, the right contrast is E with R, B being a particularly pernicious species of E. Es take mental structures to be inductive products of sensory inputs. Bs repudiate mental structures altogether, favoring direct correlation with environmental stimuli. So whereas Es allow for mental structures which are reducible to environmental parameters, Bs eschew even these.[1]  Chomsky’s anti-E arguments were not confined to the B version of E. It extends to all Associationist conceptions.

Second, TB’s observation regarding the contrasting directions of explanation for Es and Rs exposes E’s unscientific a priorism. Es start with an unfounded theory of learning and infer from this what is and is not acquirable/learnable. This relies on the (incorrect) assumption that the learning theory is well-grounded and so can be used to legislate acquisition’s limits.

Why such confidence in the learning theory? I am not entirely certain. In part, I suspect that this is because Es confuse two different issues: they run together the pretty obvious correct observation that belief fixation causally requires stimulus input (e.g. I speak west island Montreal English because I was raised in an English speaking community of west-island Montrealers) with the general conception that all beliefs can be logically reduced to inductions over observational (viz. sensational) inputs. Rs can (and do) accept the first truistic part while rejecting the second much stronger conception (e.g. the autonomy of syntax thesis just is the claim that syntactic categories and processes cannot be reduced to either semantic or phonetic (i.e. observational) inputs). Here’s where Rs introduce the notion of an environmental “trigger.” Stimuli can trigger the emergence of beliefs. They do not shape them. Beliefs are more than congeries of stimuli. They have properties of their own not reducible to (inductive) properties of the observational inputs.

Rs reverse the E direction of inquiry. Rs start with a description of the beliefs attained and then ask what kind of acquisition mechanism is required to fix the beliefs so described. In short, Rs argue from facts describable in (relatively) neutral theoretical terms and then look for cognitive theories able to derive these data. If this looks like standard scientific practice, it’s because it is. Theories that ascribe a priori knowledge to the acquisition system (as R accounts typically do) need not themselves suffer from methodological a priorism (as E theories of learning typically do). These points have often been confused. Why? R has suffered from a branding problem. The morphological connection between ‘empiricism’ and ‘empirical’ has misled many onto thinking that Es care about the data while Rs don’t. False. If anything, the reverse is closer to the truth, for Rs do no put unfounded a priori restrictions on the class of admissible explananda.

2.     Empiricism in linguistics had a particular theoretical face: the discovery procedure (DP), understood as follows (115):

Language was to be described in a hierarchy of levels of learned units such that the units at each level can be expressed as a grouping of units at an intuitively lower level. The lowest level was necessarily composed of physically definable units.

This conception has a very modern ring. It’s the intuition that lies behind Deep Learning (DL) (see here). DL exploits a simple idea: that learning not only induces from the observational input but that outputs of prior inductions can serve as inputs to later (more ”abstract”) ones. In contrast to turtles, its inductions all the way up. DL, then, is just the rediscovery of DPs, this time with slightly fancier machines and algorithms. DL is now very much in vogue. It informs the work of psychologists like Elisa Newport, among others. However, whatever its technological virtues, GGers know it to an inadequate theory of language acquisition. How do we know this? Because we’ve run around this track before. DL is a gussied up DP and all the new surface embroidery does not make it any more adequate as an acquisition model for language. Why not? Because higher levels are not just inductive generalizations over lower ones. Levels have their own distinctive properties, and this we have known for at least 60 years.

TB’s discussion of DPs and their empirical failures is very informative (especially Harris’s contribution to the structuralist DP enterprise). It also makes clear why the notion of “levels” and, in particular, their “autonomy” is such a big deal. If levels enjoy autonomy then they cannot be reduced to generalizations over information at earlier levels. There can, of course, be mapping relations between levels, but reduction is impossible. Furthermore, in contrast to DP (and DL) there is no asymmetry to the permissible information flow: lower levels can speak to higher ones and vice versa. Given the contemporary scene, there is a certain déjà vu quality to TBs history, and the lessons learned 60 years ago have, unfortunately, been largely unlearned. In other words, TB’s discussion is, sadly, very relevant still.

3.     Linguistics and Psycholinguistics

The bulk of TB’s paper is a discussion of how early theories of GG mixed with the ambitions of psychologists. GG is a theory of competence. We investigate this competence by examining native speaker judgments under “reflective equilibrium.” Such judgments abstract away from the baleful effects of resource limitations such as memory restrictions or inattention and (it is hoped) this allows for a clear inspection of the system of linguistic knowledge as such. As TB notes, very early on there was an interesting interaction between GG so understood and theories of linguistic behavior (122):

Linguistics made a firm point of insisting that, at most, a grammar was a model of “competence” – what the speaker knows. This was distinguished form “performance” – how the speaker implements this knowledge. But, despite this distinction, the syntactic model had great appeal as a model of the processes we carry out when we talk and listen. It offered a precise answer to the question of what we know when they know the sentences in their language: we know the different coherent levels or representation and the linguistic rules that interrelate those levels. It was tempting to postulate that the theory of what we know is a theory of what we do…This knowledge is linked to behavior in such a way that every syntactic operation corresponds to a psychological process…

Testing the hypothesis that there is a one-to-one relation between grammatical rules/levels and psychological processes and structures was described as investigating the “psychological reality” of linguistic structures/operations in ongoing behavior. In other words, how well does linguistic theory accommodate behavioral measures (confusability, production time, processing time, memorizability, priming) of language use in real time? TB reviews this history, and it is fascinating.

A couple of comments: First, the use of the term “psychological reality” was unfortunate. It implied that what GG studied was not a part of psychology. However, this, if TB is right, was not the intent. Rather, the aim was to see if the notions that GGers used to great effect in describing linguistic knowledge could be extended to directly explain occurrent linguistic behavior. TB’s review suggests that the answer is in part “yes!” (see TB’s discussion of the click experiments, especially as regards deep structure on 127). However, there were problems as well, at least as regards early theories. Curiously, IMO, one interesting feature of TB’s discussion is that the problems cited for the “identification thesis” (IT) are far less obvious from the vantage point of today’s Gs then those of yesteryear.

Let me put this another way: one thing that theorists like to ask experimentalists is what the latter bring to the theoretical table. There is a constant demand that psycholinguistic results have implications for theories of competence. Now, I am not one who believes that the goal of psycholinguistic research should be to answer the questions that most amuse me. There are other questions of linguistic interest. However, the early history that TB reviews provides potentially interesting examples of how psycholinguistic results would have been useful for theoreticians to consider. In particular TB offers examples in which the psycholinguistic results of this period pointed towards more modern theories earlier than purely linguistic considerations did (e.g. see the discussion of particle movement (125) or dative shift (124)). Thus, this period offers examples of what many keep asking for, and so they are worth thinking about.

Second, TB argues that the “psychological reality” considerations had mixed results. The consensus was that there is lots of evidence for the “reality” of linguistic levels but less evidence that G rules and psychological processes are in a one-to-one relation. In other words, there is consensus that the Derivational Theory of Complexity (DTC) is wrong.

For what it’s worth, my own view is that this conclusion is overstated. IMO it’s hard to see how the DTC could be wrong (see here). Of course, this does not mean that we yet understand how it is right.  Nonetheless, a reasonable research program is to see how far we can get in assuming that there is a very high level of transparency between the operations and structures of our best competence theories and those of our best performance theories. At least as a regulative ideal, this looks like a good assumption, and it has produced some very interesting work (e.g. see here).

Let’s end. Tom Bever has written a very useful paper on a fascinating period of GG history. It’s a very good read, with lessons of great contemporary relevance. I wish that were not so, but it is. So take a look.


[1] If internal representations map perfectly onto environmental variables, then the advantages of the former are unclear. However, eschewing representations altogether is not a hallmark of classical Eism.

Friday, January 9, 2015

More on intra-nuron computation

I here linked to a Youtube video of a talk that Randy Gallistel gave in Boston that goes over his argument for an intra-neuron based conception of brain computation. This recapitulates much of what he argues for in his paper that I posted (here), although it only goes over one of three recent experiments that supports this view in detail (btw, the other two are also pretty neat).  It is well worth watching for the presentation is very clear and easy to follow.

It also has an added bonus: a skeptical commentary by John Lisman from Brandeis. Lisman argues that the results that Randy points to are not inconsistent with a more classical inter-neuron circuit based conception of brain computation. He reviews some of his own work in this regard, which is interesting. 

I personally found a few other things interesting as well. First, Lisman has two main arguments, both promisory. The first is that though he agrees with Randy that the LTP/D evidence does not support the required timing for learning that it has been pressed to serve, this does not entail that some other version of the theory might not be serviceable. The second is that he has recently shown how to block/erase LTP/D accretions and hopes to run an experiment showing that this suffices to also erase a memory. He notes that should he be able to do this it would provide evidence for the view that the inter-neuron LTP/D mechanism underlies memory.  This experiment has not yet been run, hence my description as ‘promisory.’

Randy notes in his talk and paper that the LTP/D theory of memory as strengthened connections among neurons in a net is a venerable theory. It’s been around for a very long time. Lisman seems to agree. I thus found it interesting that Lisman’s retort did not proceed by citing chapter and verse of results in favor of the view but was largely defensive: a “that this is wrong does not mean that some version may be right” strategy for defending it. Second, I found it interesting that Lisman did not seem to understand Randy’s request to outlined how numerical information could be stored in a synapse in a way that it could be used. Randy wanted a description of a general mechanism for doing this independently of how this information was then put to use. Lisman kept retorting how this or that particular phenomenon might be modeled. In Randy’s view memory is the capacity to code information for storage that can later be retrieved. There is no mention as to how this information will be used. Presumably stored information can be used in multiple ways. For Lisman memory is not something that is use neutral but a link between one behavior and another. Lisman, in effect, does not seem to understand what Randy is asking him to provide, or, to be more charitable (as I should be) he rejects the idea that memory is disconnected from the uses it is put to.

I mention this for it reflects a point that is a central feature of Empiricism: identifying what something is as what that thing does. I’ve discussed this before (here), but I liked this particular example of the difference in action.  Let me say a bit more.

Many (me included) have tended to take the salient property of Empiricism to be its penchant for associationism. Randy points to this in his paper and lecture and he is, of course, correct to note the strong relationship. However, I now think that the deeper feature of Empiricism is its identification of the “powers/nature” of a thing (these are Cartwright’s terms, see above link) with what it does. Rationalists reject this. To specify the powers or nature of something is to provide an abstract description of its properties. What it does is a complex interaction of these properties with other things.  So when Randy asks for a physical basis for memory he wants an account of how to store information in a retrievable form independent of how this information might be later used and manifested. The mechanism is general. The occasion of use is just one manifestation of it. For Empiricists what something is just is a summary (perhaps statistical) of its effects. Not so for the Rationalist.[1]

Last point: Lisman takes Randy as arguing that brain computation must supervene on DNA/RNA structure. Randy points out both in the paper and in the talk that this is not a central feature of his view. We know how information is coded in DNA/RNA and so it provides a proof of concept for intra-neuronal computation that DNA/RNA machinery could provide the physical mechanisms required for such computation. But, as Randy notes that the cognitive machinery be DNA/RNA is not necessary to his main point (see p. 6). Other kinds of molecules can serve to undergird such computation (the most relevant feature, it seems, having two stable states thus allowing the molecule to serve as a switch coding 1/0). This said, Randy notes in the video that we currently do not know much about how information is coded within the neuron. Thus, if Randy is correct, these very important details remain to be developed. Randy’s argument is that this is a reasonable place to look for the computational bases of cognition given that the apparent failure of the inter-neuron conception and the evidence that is amassing that individual cells (rather than networks) do in fact store acquired information in usable form.

To end: As I’ve said before, I find this to be really exciting stuff. So, watch the video, it’s really fun.



[1] On this view, the competence/performance distinction is itself an important Rationalist one.