Comments

Showing posts with label Tom Bever. Show all posts
Showing posts with label Tom Bever. Show all posts

Tuesday, March 8, 2016

Bever's Whig History of GG

I love Whig History (WH). I have even tried my hand at it (here, here, here, here). What sets them apart from actual history is that they abstract away from the accidents of history and, if successful, reveal the “inner logic” of historical events. A Whig History’s conceit is that it outlines how events should have unfolded had they been guided by rational considerations. We all know that these are never all that goes on, but scientists hope that this goes on often enough, if not at the individual level, then at the level of the discipline as a whole. Even this might be debated (see here), but to the degree that we can fashion a WH, to that degree we can rationally guide inquiry by learning from our past mistakes and accomplishments. It’s a noble hope, and I am a fervent believer.

Given this, I am always on the lookout for good rational reconstructions of linguistic history. I recently came across a very good one by Tom Bever that I want to make available (here). Let me mark a few of my personal favorite highlights.

1.     The paper starts with a nice contrast between Behaviorist (B) vs R methodologies:

The behaviorist prescribes possible adult structures in terms of a theory of what can be learned; the rationalist explores adult structures in order to find out what a developmental theory must explain.

Two comments: First, throughout the paper, TB contrasts B with R. However, the right contrast is E with R, B being a particularly pernicious species of E. Es take mental structures to be inductive products of sensory inputs. Bs repudiate mental structures altogether, favoring direct correlation with environmental stimuli. So whereas Es allow for mental structures which are reducible to environmental parameters, Bs eschew even these.[1]  Chomsky’s anti-E arguments were not confined to the B version of E. It extends to all Associationist conceptions.

Second, TB’s observation regarding the contrasting directions of explanation for Es and Rs exposes E’s unscientific a priorism. Es start with an unfounded theory of learning and infer from this what is and is not acquirable/learnable. This relies on the (incorrect) assumption that the learning theory is well-grounded and so can be used to legislate acquisition’s limits.

Why such confidence in the learning theory? I am not entirely certain. In part, I suspect that this is because Es confuse two different issues: they run together the pretty obvious correct observation that belief fixation causally requires stimulus input (e.g. I speak west island Montreal English because I was raised in an English speaking community of west-island Montrealers) with the general conception that all beliefs can be logically reduced to inductions over observational (viz. sensational) inputs. Rs can (and do) accept the first truistic part while rejecting the second much stronger conception (e.g. the autonomy of syntax thesis just is the claim that syntactic categories and processes cannot be reduced to either semantic or phonetic (i.e. observational) inputs). Here’s where Rs introduce the notion of an environmental “trigger.” Stimuli can trigger the emergence of beliefs. They do not shape them. Beliefs are more than congeries of stimuli. They have properties of their own not reducible to (inductive) properties of the observational inputs.

Rs reverse the E direction of inquiry. Rs start with a description of the beliefs attained and then ask what kind of acquisition mechanism is required to fix the beliefs so described. In short, Rs argue from facts describable in (relatively) neutral theoretical terms and then look for cognitive theories able to derive these data. If this looks like standard scientific practice, it’s because it is. Theories that ascribe a priori knowledge to the acquisition system (as R accounts typically do) need not themselves suffer from methodological a priorism (as E theories of learning typically do). These points have often been confused. Why? R has suffered from a branding problem. The morphological connection between ‘empiricism’ and ‘empirical’ has misled many onto thinking that Es care about the data while Rs don’t. False. If anything, the reverse is closer to the truth, for Rs do no put unfounded a priori restrictions on the class of admissible explananda.

2.     Empiricism in linguistics had a particular theoretical face: the discovery procedure (DP), understood as follows (115):

Language was to be described in a hierarchy of levels of learned units such that the units at each level can be expressed as a grouping of units at an intuitively lower level. The lowest level was necessarily composed of physically definable units.

This conception has a very modern ring. It’s the intuition that lies behind Deep Learning (DL) (see here). DL exploits a simple idea: that learning not only induces from the observational input but that outputs of prior inductions can serve as inputs to later (more ”abstract”) ones. In contrast to turtles, its inductions all the way up. DL, then, is just the rediscovery of DPs, this time with slightly fancier machines and algorithms. DL is now very much in vogue. It informs the work of psychologists like Elisa Newport, among others. However, whatever its technological virtues, GGers know it to an inadequate theory of language acquisition. How do we know this? Because we’ve run around this track before. DL is a gussied up DP and all the new surface embroidery does not make it any more adequate as an acquisition model for language. Why not? Because higher levels are not just inductive generalizations over lower ones. Levels have their own distinctive properties, and this we have known for at least 60 years.

TB’s discussion of DPs and their empirical failures is very informative (especially Harris’s contribution to the structuralist DP enterprise). It also makes clear why the notion of “levels” and, in particular, their “autonomy” is such a big deal. If levels enjoy autonomy then they cannot be reduced to generalizations over information at earlier levels. There can, of course, be mapping relations between levels, but reduction is impossible. Furthermore, in contrast to DP (and DL) there is no asymmetry to the permissible information flow: lower levels can speak to higher ones and vice versa. Given the contemporary scene, there is a certain déjà vu quality to TBs history, and the lessons learned 60 years ago have, unfortunately, been largely unlearned. In other words, TB’s discussion is, sadly, very relevant still.

3.     Linguistics and Psycholinguistics

The bulk of TB’s paper is a discussion of how early theories of GG mixed with the ambitions of psychologists. GG is a theory of competence. We investigate this competence by examining native speaker judgments under “reflective equilibrium.” Such judgments abstract away from the baleful effects of resource limitations such as memory restrictions or inattention and (it is hoped) this allows for a clear inspection of the system of linguistic knowledge as such. As TB notes, very early on there was an interesting interaction between GG so understood and theories of linguistic behavior (122):

Linguistics made a firm point of insisting that, at most, a grammar was a model of “competence” – what the speaker knows. This was distinguished form “performance” – how the speaker implements this knowledge. But, despite this distinction, the syntactic model had great appeal as a model of the processes we carry out when we talk and listen. It offered a precise answer to the question of what we know when they know the sentences in their language: we know the different coherent levels or representation and the linguistic rules that interrelate those levels. It was tempting to postulate that the theory of what we know is a theory of what we do…This knowledge is linked to behavior in such a way that every syntactic operation corresponds to a psychological process…

Testing the hypothesis that there is a one-to-one relation between grammatical rules/levels and psychological processes and structures was described as investigating the “psychological reality” of linguistic structures/operations in ongoing behavior. In other words, how well does linguistic theory accommodate behavioral measures (confusability, production time, processing time, memorizability, priming) of language use in real time? TB reviews this history, and it is fascinating.

A couple of comments: First, the use of the term “psychological reality” was unfortunate. It implied that what GG studied was not a part of psychology. However, this, if TB is right, was not the intent. Rather, the aim was to see if the notions that GGers used to great effect in describing linguistic knowledge could be extended to directly explain occurrent linguistic behavior. TB’s review suggests that the answer is in part “yes!” (see TB’s discussion of the click experiments, especially as regards deep structure on 127). However, there were problems as well, at least as regards early theories. Curiously, IMO, one interesting feature of TB’s discussion is that the problems cited for the “identification thesis” (IT) are far less obvious from the vantage point of today’s Gs then those of yesteryear.

Let me put this another way: one thing that theorists like to ask experimentalists is what the latter bring to the theoretical table. There is a constant demand that psycholinguistic results have implications for theories of competence. Now, I am not one who believes that the goal of psycholinguistic research should be to answer the questions that most amuse me. There are other questions of linguistic interest. However, the early history that TB reviews provides potentially interesting examples of how psycholinguistic results would have been useful for theoreticians to consider. In particular TB offers examples in which the psycholinguistic results of this period pointed towards more modern theories earlier than purely linguistic considerations did (e.g. see the discussion of particle movement (125) or dative shift (124)). Thus, this period offers examples of what many keep asking for, and so they are worth thinking about.

Second, TB argues that the “psychological reality” considerations had mixed results. The consensus was that there is lots of evidence for the “reality” of linguistic levels but less evidence that G rules and psychological processes are in a one-to-one relation. In other words, there is consensus that the Derivational Theory of Complexity (DTC) is wrong.

For what it’s worth, my own view is that this conclusion is overstated. IMO it’s hard to see how the DTC could be wrong (see here). Of course, this does not mean that we yet understand how it is right.  Nonetheless, a reasonable research program is to see how far we can get in assuming that there is a very high level of transparency between the operations and structures of our best competence theories and those of our best performance theories. At least as a regulative ideal, this looks like a good assumption, and it has produced some very interesting work (e.g. see here).

Let’s end. Tom Bever has written a very useful paper on a fascinating period of GG history. It’s a very good read, with lessons of great contemporary relevance. I wish that were not so, but it is. So take a look.


[1] If internal representations map perfectly onto environmental variables, then the advantages of the former are unclear. However, eschewing representations altogether is not a hallmark of classical Eism.

Saturday, March 7, 2015

How to make $1,000,000

My mother once told me about an easy way to become a millionaire: start with $10 million. This seems to be advice that second generations are particularly good at following. And not only as regards inter-generational wealth transfer. Like families, journals also enjoy life cycles, with founders giving way to a next generation. And as in families, regression to the mean (i.e. the headlong rush to average) seems to be an inexorable force. However, in contrast to the rise and decline of wealth in families, the move from provocative to staid in journals is rarely catalogued. Rarely, but not never. Here is a paper (by Priva and Austerweil (P&A)) that charts the change in intellectual focus of Cognition, a (once?) very important cogsci journal. What does P&A show? Two things: (i) It shows that the mix of papers in the journal has substantially changed. Whereas in the beginning, there was a fair mix of theory and experimental papers (theory papers predominating), since the mid 2000s the mix has dramatically changed, with experimental papers forming the bulk of academic product. Theory papers have not entirely disappeared, but they have been substantially overtaken by their experimental kin. (ii) That papers on language and development have gone from central topics of interest to a somewhat marginal presence.[1]

How surprising is this?  Let me start by discussing (i), the decline of “theory,” a favorite obsession of mine (see here). Well, first off, from one common perspective, some decline might be expected. We all know the Kuhnian trope; “revolutionary” periods of scientific inquiry where paradigms are contested, big ideas are born and old ones die out (one old fogey at a time in Plank time) give way to periods of “normal science” where the solid majestic wall of scientific accomplishment is carefully and industriously built brick by careful empirical brick. The picture offered is one in which the basic framework ideas get hashed out and then their implications are empirically fleshed out. I never really liked this way of conceptualizing matters (there is a lot of hashing and fleshing going on all the time in serious work), but I think that this picture has some appeal descriptively. Sadly, it also seems to have normative attractions, especially to next generation editors. Here’s what I mean.

Editing a journal is a lot of work. Much of it thankless. So before I go off the deep end here in a minute, let me personally thank those that take this task on, for their work and commitment is invaluable and what we think of as science could not succeed without such effort. That said, precisely because of how hard it is to do, you need to be driven (nuts?) to start a journal. What drives you? The feeling that there is something new to say but that there is no good place to say it. Moreover, not only is that something new, it must be important and new. And it is not possible to say these new important things in the current journals because the new ideas cut things up in new ways or approach problems from premises that don’t fit into the existing journalistic matrix.[2] So, at the very least, the extant venues are not congenial places to publish, and in some cases are outright hostile.

The emergence of cogsci (something that happened when I was growing up intellectually) had this feel to it. There was a self-conscious cognitive revolution, with very self-conscious revolutionaries. Furthermore, this revolution was fought on several fronts: linguistics, psychology, computer science and philosophy being the four main ones. Indeed, for a while, it was not clear where one left off and the other began. Linguists read philosophy and psychology papers, psychologists knew about Transformational Grammar and could Locke from Descartes, philosophers debated what to make of innate ideas, representations and rule following based on work in linguistic and psychology and computer scientists (especially in AI) worried about the computational properties of mental representations (think Marr and Marcus for example).  Cogsci lived at the intersection of these many disciplines, was nurtured by their cross disciplinary discussions and, for someone like me, cogsci became identified as the investigation of the structures of minds (and one day brains) using the techniques and methods of thought that each discipline brought to the feast. Boy was this exciting. Not surprisingly, the premiere journal for the advancement of this vision was Cognition. Why not surprisingly? Because the founding editors, Jacques Mehler and Tom Bever, were two people that thoroughly embodied this new combined intellectual vision (and were and are two of its leading lights) and they built Cognition to reflect it.[3]

A nice way of seeing this is to read Mehler’s “farewell remarks” here. It is very explicit about what gap the journal was intended to fill:

Our aim was to change the publishing landscape in psychology and related disciplines that became part of “Cognitive Science.” …[P]sychology had turned almost exclusively into an experimental discipline with an overt disdain for theory…Linguistics had become a descriptive discipline often favoring normative or purely descriptive over theoretical approaches. Professional journals in line with this outlook generally obliged contributors to write their papers in standard format that privileged the shortest possible introductions and conclusions, methods and procedures used in experiments. Papers by non-experimental scientists say, philosophers of mind or theoretical linguists, were rarely even accepted…. (p. 7)

In service of this, the journal was the venue of lots of BIG debates concerning connectionism, the representational theory of mind, compositionality, AI models of mind, prototypes, domain specificity, computational complexity, core knowledge and much much more. In fact, Cognition did something almost miraculous: It became a truly inter-disciplinary journal, something that administrators and science bureaucrats (including publishers) love to talk about (but, it seems, often fail appreciate when it happens).

P&A records that this Cognition now seems to be largely gone. It is no longer the journal its editors founded. There is little philosophy and little linguistics or linguistically based psychology. Nor does it seem to any longer be the venue where big ideas are thrashed out. Three illustrations: (i) the critical discussions concerning Bayesian methods in psychology have not occurred in the pages of Cognition,[4] (ii) nor have the Gallistel-like critiques of connectionist neuro-science gotten much of an airing, (iii) nor have extensive critiques of resurgent “language” empiricism (e.g. Tomasello) made an appearance. These have gotten play elsewhere, and that is a good thing, but these dogs have not barked in Cognition, and their absence is a good indicator of how much Cognition has changed. Moreover, this change is no accident. It was policy.

How so? Well, in the same issue that Mehler penned his farewell the new incoming editor Gerry Altmann gave his inaugural editorial (here). It’s really worth reading the Mehler and Altmann pieces side by side, if nothing else as an exercise in the sociology of science. I’ve rarely read anything that so embodies (and embraces) the Kuhnian distinction between revolutionary vs normal science. Altmann’s editorial is six pages long. After some standard boilerplate thanking Mehler & Co. for its path-breaking efforts, Altmann sets out his vision of the future. It comes in two parts.

First the ideal paper:

To be published in Cognition, articles must be robust in respect of the fit between the theory, the data and the literature in which the work is grounded. They should have a breadth to them that enables the specific research they describe to make contact with more general issues in cognition; the more explicit this contact, the greater the impact of the research beyond the confines of the specialized research community. (2)

It’s worth contrasting this ideal with the more expansive one provided by Mehler above. In Altmann’s, there is already an emphasis on “data” that was missing from Mehler’s discussion. In other words, Altmann’s ideal has an up front experimental tilt. Data’s the lede. The vision thing is filler. To see this, read the two sentences in reverse order. The sense of what is important changes. In the actual quoted order what matters is data fit then idea quality. Reverse the sentences and we get first idea quality and then data fit. Moreover, unlike Mehler’s pitch, what’s clear here is that Altmann does not envision papers that might be good and worthwhile even were they bereft of data to fit. It more or less assumes that the conceptual issues that were at the foundation of the cogsci revolution have all been thoroughly investigated and understood (or were largely irrelevant to begin with (dare I say, maybe even BS?)). More charitably, it assumes that if something new does rise under the cognitive sun, it will arise from the carefully fitted data. In short, the main job of the cogscientist is see how the theories fit the facts (or vice versa). Theory alone is so your grandparent’s cognition.

The second part of the editorial reinforces this reading. The last 3 pages (i.e. half the editorial), section 3, concerns “the appropriate analyses of data” (4). It’s a long discussion of what stats to use and how to use them. There is no equally long section discussing thos hot topics/problems, what issues are worth addressing and why. This reinforces the conclusion that what Cognition will henceforth worry most about is data fit and experimental procedure. Sounds like the kind of journal that Mehler and Bever had hoped that Cognition would displace. Indeed, prior to Cognition’s founding, psychology had lots of the kinds of Journals that the Altmann editorial aspires to. That’s precisely why Mehler and Bever started their journal. Altmann appears to think that psychology needs one more.

If this read is right, then it is not surprising that Cognition’s content profile has changed over the years. It is not merely that new topics get hot and old ones get stale. Rather, it is that what was once a journal interested in bridging disciplines, critically investigating big issues and provoking thought, “grew up” and happily assumed the role of purveyor of “normal” science. A nice well behaved journal, just like most of the others.

Last two points. Given the apparent dearth of interest in theory, it is not a surprise (to me) that work on language is less represented in the new Cognition. Anything that takes linguistic theory seriously in psycho study will be suspect to those with a great respect for psychological techniques (we don’t gather data the right way, there is a distance between competence and performance, we think that minds are not all purpose learners etc.). Thus taking results in linguistic theory as starting points will go against the intellectual grain where theory is less important than data points. This need not have been so. But that it is so is not surprising.

Second, there is a weird part of Altmann’s editorial concerning the “collaborative” nature of science and how this should be reflected in the “editorial structure” of the journal. Basically, it seems to be signaling a departure from past methods. I don’t really know how the Mehler era operated “editorially.” But it would not surprise me were he (and Bever) more activist editors than is commonplace. This would go far, IMO, in explaining why the old Cognition was such a great journal. It expressed the excellent taste of its leaders. This is typically true of great journals. At one time the leading figures edited journals and imposed their tastes on the field, to its benefit. Max Planck edited the journals that published Einstein’s groundbreaking (and very unconventional) papers.[5] Keynes edited the most important economics journal of his day. Mehler and Bever were intellectual leaders and Cognition reflected their excellent taste in questions and problems. It strikes me that the Altmann editorial is a non too subtle critique of this. It’s saying that going forward editorial decisions would be more balanced and shared. In other words, more watered down, more common denominatorish, less quirky, more fashionnable. There is room for this ideal, one where the aim is to reflect the scientific “consensus.” Today, in fact, this is what most journals do. Mehler and Bever’s Cognition did not.

To end: Cognition has changed. Why? Because it wanted to. It has managed to achieve exactly what the new regime was aiming for. The old Cognition stood apart, had a broad vision and had the courage of its new ideas. The new Cognition has re-joined the fold. A good journal (no doubt). But no longer a distinctive one. It’s not where people go to see the most important ideas in cognition vigorously debated. It’s become a professional’s journal, one among many. Does it publish good papers? Sure. Is it the indispensible journal in cogsci that it once was? Not so much. IMO, that’s really too bad. However, it is educational, for now you know how to make $1,000,000 dollars. Just be sure to start off with $10,000,000.




[1] This is all premised on the assumption that the topic model methodology used in the paper accurately reflects what has been going on. This may be incorrect. However, I confess that it accurately reflects what many people I know have noted anecdotally.
[2] Is this PoMo or what? With a tinge of the Wachowskis thrown in.
[3] And you know many of the others. To name a few: Chomsky, two Fodors, Gleitman, Gallistel, Katz, Pylyshyn, Gellman, Garrett, Block, Carey, Spelke, Berwick, Marr, Marcus, a.o.
[4] E.g. Eberhardt & Danks, Brown & Love, Bowers & Davis, Marcus have all appeared in other venues. See here, here and here form some discussion and references.
[5] A friend of mine in theoretical physics once told me that he doubted that papers like Einstein’s great 1905 quartet could be published today. Even by the standards in 1905 they looked strange. Moreover, they were from a nobody working in a patent office. It’s a good thing for Einstein that Planck, one of the leading physicists of his day, was the editor of Annalen der Physik.