Comments

Friday, March 20, 2015

Some weekend reads

1. This blog entry discusses the review processes in sociology and how it is being affected by Nate Silver’s online 538 sight. Lindner (the author of the scooped paper) makes a provocative point concerning the value added of the review process. Here’s a taste of the argument:

Silver writes for a new technocratic audience and produces posts with “outputs from multivariate regression analyses, resplendent with unstandardized coefficients, standard errors, and R2s.” It might not have quite the rigor of academic papers, but it yields many of the same results. Even more importantly, “Unlike academics, Silver is unburdened by the constraining forces of peer review, turgid and esoteric disciplinary jargon, and the unwieldy format of academic manuscripts. He need not kowtow to past literature, offer exacting descriptions of his methods, or explain in tedious detail how his findings contribute to existing theory.”

2. Johan Bolhuis sent me this fascinating pair of papers (here, here) on cortical computations in mammals and birds. It seems that birds and mammals share a common cortical circuitry strongly suggesting that what we have and what they have brains wise is pretty much the same thing. As the Harris paper puts it:

Perhaps intelligence isn’t such a hard trick after all: a basic circuit capable in principle of supporting advanced cognition might
have evolved hundreds of millions of years ago, but only adapted to this purpose when the benefits actually outweighed the costs
of increased head size, development time, and energy use. Tool-use wouldn’t do much for a sheep; those few times intelligence was favored by evolution, it may have appeared with remarkably little effort, by repurposing an ancient circuit most animals use for other things.

Johan sent me this short additional very suggestive comment. Very soon the question may not be why we have merge but why every thing else doesn’t? Or maybe they do, but we can’t see it yet?

There you have it! More evidence to suggest that you don’t need big neural changes to achieve ‘cognitive’ changes. As I like to put it in talks: the basic neurogenetic machinery is there in all vertebrates, it’s what you do with it that counts. In the case of humans and songbirds (and a few other taxa) - but not apes or mice - auditory-vocal imitation learning evolved with it, and in the case of humans (but not other species) ‘Merge’ evolved with it, possibly as the result of a mutation that led to a rewiring of the cortex.

3. Rainer Mausfield sent me a reference to a study that relates to the three-remark rule that I mentioned here. To repeat: it’s been my experience that if someone hears something three times at a conference and nobody objects, it becomes accepted wisdom.  It seems that there is some (weakish) data to back this up. Rainer sent me this link and this paper. Though the experiments are useful, the logic behind this seems sound to me. Say you are at a specialist conference and someone gives a talk that is not in your area of expertise, yet it seems not that persuasive. A reasonable strategy is to defer to the experts, who, you hope, can be counted upon to make the status of the (perhaps more controversial bits) clear in the question period. If this does not happen, then it is reasonable (on Bayes grounds, I believe) to conclude that the absence of criticism is a sign of the accepted truth of the claims.  So, if you hear something that you think objectionable it is incumbent upon you to speak up. Others will be taking their cure from you. This is part of what makes scientific inquiry a collective enterprise.

4. Caveat lector; this piece is from Aeon (the Vyvyan Evans venue)!! The author is Andreas Wagner (here) and he has a cross appointment at the Sante Fe Institute. At any rate, I found the discussion provocative and relevant for the EVOLANG discussion concerning the emergence of merge in the species. Here’s a useful quote:

How do random DNA changes lead to innovation? Darwin’s concept of natural selection, although crucial to understand evolution, doesn’t help much. The thing is, selection can only spread innovations that already exist. The botanist Hugo de Vries said it best in 1905: ‘Natural selection can explain the survival of the fittest, but it cannot explain the arrival of the fittest.’ (Half a century earlier, Darwin had already admitted that calling variations random is just another way of admitting that we don’t know their origins.)

This is obviously relevant to the question of how linguistic facility arose; where did it come from? If Chomsky is right in thinking that there is nothing quite like merge in our ancestors, then we need something other than a natural selection (NS) account for how it arose. Of course, once it arises, we can ask why it did not disappear (this is where NS would come in). But how it arrived? NS has nothing to say about this (though 2 above suggests that more may lay dormant than might meet the untutored eye).

Wagner believe that the “arrival” space is highly organized, thereby constraining possible evolutionary trajectories independently of the effects of NS (sound familiar? Think UG and “learning”). Wagner discusses how the space of genetic possibilities might be organized so as to be searchable by NS. He mentions that absent such a structure, the size of the possibility space would make evolution miraculous:

If you had to find a text on a specific subject in such a library – without a catalogue – you would get utterly lost. Worse than that, if missteps can be fatal, you would quickly die. Yet life not only survived, it found countless new meaningful texts in these libraries. Understanding how it did that requires us to build the catalogue that evolution lacks. It demands that we work out how these libraries are organised to comprehend how innovation through blind search is possible.


I have no idea whether this is right (though it sounds plausible to me). So if anyone has a grasp on these matters, please enlighten the rest of us. What seems clear is that if something like this is correct, it fits well with other work that constrains evolutionary trajectories coming from the Evo-Devo literature. There is more than a slight analogy between this kind of discussion and the one we had in cognition/linguistics 60 years ago.  When considering the mechanics of change two factors will always loom large: the set of possible trajectories and the set of factors that choose between these options. NS is a factor of the second type. Until recently (or so it seems to an outsider like me) the main idea has been that the range of possible trajectories was so humongous that the bulk of an evolutionary explanation would be carried by NS like factors. This no longer seems as clear. It looks like in evolution (as in “learning”) the range of options available are tightly constrained and so a (large?) part of any evolutionary account will advert to these circumscribed possibilities. It goes without saying (but I will say it) that any complete story will likely have factors of both kinds. However, it seems clear that the narrower the options, the less role there will be for NS/learning, the wider the options the greater the causal efficacy of NS/learning.  It’s interesting that biologists have started focusing on the restricted possibility space, as a major factor in evolutionary change.

Wednesday, March 18, 2015

A (shortish) Whig history of Generative Grammar (part 1)

 0. One of the striking characteristics of modern scientific practice, at least in the successful domains of inquiry, is that it is possible to stand on the shoulders of one’s predecessors (of various heights), thereby allowing one to (occasionally successfully) address questions that would otherwise be out of reach.  This cumulative aspect of scientific inquiry is so distinctive (and important) that it is not unreasonable to use it as a measure of how much scientific traction a domain of investigation has achieved.

The converse is also true: to the degree that a field of inquiry enjoys a “revolution” every decade to that degree it suggests that it has not yet made the jump from (possibly informed) speculation to scientific inquiry.  Revolutions tend to discard the past, not build on it and this implies that what is discarded is not worth retaining. Permanent scientific revolution within a field is a leading indicator of ignorance.

Given this epistemological yardstick it is germane to ask how well modern Generative Grammar (GG) has measured up.  Not well, is one standard view.  Here we argue that this misreads the history of GG. We aim to back this judgment up by providing a Whig history (WH) of GG that displays its cumulative character.  Before proceeding, a word of caution. WHs are not “real” histories. They present the past as “an inevitable progression towards ever greater…enlightenment.” Real history is filled with dead ends, lucky breaks, misunderstandings, confusions, petty rivalries, and more.  WHs are not. They focus on the “successful chain of theories and experiments that led to the present-day science, while ignoring failed theories and dead ends” (see here). A less pejorative term for WH is “rational reconstruction.” At its best, a WH can expose the logic behind a set of questions and the attendant inquiries into them, exposing how the theories we consider today have built on the results of earlier theoretical and empirical investigations.  A domain with a credible WH is a domain where the growth of knowledge has in fact been cumulative even if the path leading to this body of knowledge has been extremely crooked. GG has a very credible WH, as we intend to demonstrate.

1. Where to begin?  The modern Generative enterprise rises from the consideration of a handful of obvious facts:
            (1) A competent speaker of a given natural language (NL) has the capacity to deal with an effectively infinite number of linguistic objects.
            (2) The linguistic objects in question are pairings of meanings with “sounds.” Thus, among the things a native speaker of English knows is that Dogs chase cats does not mean the same thing as Cats chase dogs while Cats are chased by dogs does. There are an unbounded number of such facts that a competent native speaker of a given NL knows.
            (3) Any human child can acquire competence in any NL when placed in the appropriate speech community. Thus, any child of any parentage will grow up speaking English if it grows up in NYC, Chinese if it grows up in Shanghai, Hungarian in Budapest, etc. Moreover, all children regardless of the language, or the child, do this in essentially the same way.
            (4) This capacity to acquire anything like an NL is a species-specific capacity that humans have. In other words, no other animals do language like we humans do language.

The first fact suggests that linguistic competence consists (in part) in mastery of a system of rules that specify the NL mastered.  Why a rule system? Because that is the only way to specify an effectively infinite capacity.  We cannot just list the objects in the domain of a native speaker’s competence. The capacity can only be specified in terms of a finite procedure that describes (aka: generates) it. Thus, we conclude that linguistic mastery consists (in part) in acquiring a set of rules (aka: a Grammar (G)) that generate the kinds of linguistic objects that a native speaker has competence with.   

The second fact tells us something more about these Gs. They must specify pairings of meanings with sounds. Thus the rule systems that native speakers have mastered are rules that generate objects with two distinctive properties. Gs are generative procedures that tie a specific meaning profile together with a specific sound profile, and they do this over an effectively infinite domain. So Gs are functions whose range is meaning-sound pairs, viz. an infinite number of objects like this: <m,s>. What’s the domain? Some finite set of “atoms” that can combine again and again to yield more and more complex <m,s> pairs. Let’s call these atoms ‘morphemes.’

Putting this all together, we know from the basic facts and some very elementary reasoning that native speakers master Gs (recursive procedures) that map morphemes into an unbounded range of <m,s>s. This we know. What we don’t know is what the specific rules that Gs contain look like (or for that matter what the ‘m’s and ‘s’s look like). And that brings us to our first research question: describe some rules of NL Gs and specify their variety.

We know another thing about these Gs. They can be acquired by any human when placed in the appropriate linguistic environment. You might be a native speaker of English, but you could just as well have been a native speaker of Swahili but for an accident of birth location. There is nothing that intrinsically makes you a native speaker of English (nor of Swahili, Dutch, Japanese, etc.), though, given the third and fourth facts noted above, there is likely to be something biologically intrinsic to you that makes you capable of acquiring a NL (i.e. a G that generates the NL) at all. Let’s call the capacity that humans have to acquire a NL, the ‘Faculty of Language’ (FL).  Note, postulating that there is a FL does not tell us what’s in FL. It simply names a fact: that humans (qua humans) are able to acquire any NL G in more or less the same way. We can think of FL as a function whose range is Gs of NLs.  An obvious research question is what’s the fine structure of this function FL.

This second research question has been addressed in at least two different ways. The first has been to inspect many NL Gs to induce what all Gs have in common? Note, that doing this requires having a largish set of candidate Gs and a reasonable variety of them (e.g. romance Gs, Germanic Gs, Semitic Gs, Austonesian Gs, East Asian Gs, etc.).  Clearly, we cannot inspect their commonalities without having them to inspect.[1]

There is a second way of addressing this question. We can ask what FL must contain in order to produce even a single a G only guided by the kinds of data the child uses. Call the data the child actually uses to guide its G acquisition ‘Primary Linguistic Data’ (PLD).  We can investigate the structure of FL by asking what (if anything) we must assume about the structure of FL to allow it to use the PLD(NL) (i.e. the PLD of a given NL, e.g. uttered bits of English) to arrive at the G(NL) (i.e. The G of that NL, e.g. the grammar of English) that the child attains. 

What do we know about the PLD? Actually quite a bit. We know that it consists of relatively simple linguistic forms. There are not many sentences in the PLD that the child has access to (e.g. child directed language, a superset, most likely, of what it actually uses) that involve more than 2 levels of embedding. Indeed, most of the PLD seems to consist of simple phrases and sentences.  Moreover, there are virtually no grammatically ill-formed utterances addressed to the child and no correction of mistakes that the child spontaneously makes. Thus, to a good first approximation, we can take the PLD to be simple well formed phrases and sentences addressed to the child. From this the child builds a G that permits it to generate an effectively unbounded number of phrases and sentences, both simple (like the one’s it is exposed to) and complex (language bits with structures unattested in the PLD). The idea is that we can investigate how FL is structured using the following argument form, called ‘The Poverty of the Stimulus Argument’ (POS): assume that anything that the child can acquire on the basis of the PLD it does acquire in this way. However, anywhere that the PLD is insufficient to fix the attained property implicates some built-in (i.e. innate) feature of FL.  In contrast to the comparative method noted above, the POS licenses inferences about FL based on the properties of a single G(NL).  The method is effectively subtractive: what you can acquire from PLD assume is so acquired, what’s left after you subtract this out is due to fixed (viz. innate) features of FL.[2]

Two last points before proceeding further: note (i) that investigating FL in either way requires that we have some candidate Gs. The output of FL is a G, so until we know something about the operations that NL Gs contain, it is fruitless to pursue this question about FL. (ii) that the two forms of hunting for innate features of FL are complementary and both are useful. As we will see in our WH of GG below, what we learned in studying multiple particular Gs , and POS evaluations of single Gs have both contributed to GG’s understanding of FL’s structure.

Given (4) we know that FL is a biological novelty. That means that there was a time at which our ancestors did not have a FL. A reasonable question to ask is how FL arose in humans. Note that just as having candidate Gs is a precondition for studying the properties of FL, having plausible candidate FLs is a precondition for studying how FL arose in the species.  Let’s be a little clearer here.

Let’s divide FL into those features (i) that are domain specific to the use and mastery of FL (call these principles ‘Universal Grammar’ (UG)), (ii) that FL shares with other cognitive capacities (call these ‘domain general cognitive features’ (DGCF)), (iii) features that FL has by physical necessity (PN).  The more that FL is constructed from operations and principles drawn from (ii) and (iii) to that degree the story of how FL could have arisen in the species can be simplified. For example, if all of the ingredients but one necessary to build our FL are in DGCF or PN then we can trace the rise of FL in humans to the emergence of that one distinctive property of UG.  Conversely, the more that must be packed into UG the more involved will be an explanation for how FL arose.

The difficulty is further exacerbated if FL is a relatively recent cognitive innovation, as it leaves less time for (the oft believed) gradual methods of natural selection to work their magic.[3] This leads to the following conclusion: in the best of all possible worlds, FL uses operations and principles that cognitively pre-date the emergence of FL and that these suffice to construct an FL with the properties ours has when the (hopefully, very) small number (possibly zero, but more likely one or two) of language specific cognitive innovations (i.e. UG) are added to the prior mix.  This line of reasoning suggests another clear project: show UG is pretty sparse and that very modest UGs can derive the operations and principles of richer UGs when combined with identifiable DGCFs and PNs. 

The above formulation highlights an important tension between explaining how Gs arise in a single individual and explaining how FL arose in the species.  The more we pack into UG (operations and principles specific to linguistic capacity), the easier we make the child’s task of projecting a G(NL) from the PLD(NL) it exploits.  However, the richer the UG component of FL, the harder it is for our cognitive ancestors (who, by assumption, were sans FL) to evolve an FL like ours.  That’s the tension, and contemporary GG tries to address it. However, for now, let’s just observe the tension and note that the enterprise of discussing how FL arose in the species can only get productively started once we have some idea of what properties FL has, and for this it is very useful to have some understanding of what operations and principles of FL might be linguistic specific (i.e. part of UG).

To conclude our little conceptual tour of the problem: we have here outlined the logic of the GG research program based on some pretty elementary facts. This project addresses three questions:[4]
           
(5)       a. What properties do individual Gs have?
                        b. What properties must FL have to so as to enable it to acquire these Gs?
                        c. Which of these properties of FL are proprietary to language and which
are more general?

Our WH will illustrate how research in GG can be understood as progressively addressing each of these questions, thereby setting the stage for a fruitful investigation of the next questions. This is what we should expect given our observation that (5a) is a precondition for (fruitfully) addressing (5b) and (5b) is a precondition for fruitfully addressing (5c). Note that ‘precondition’ is here used in the conceptual sense.  It does not mean that answers to (5b) might not lead to rethinking claims about (5a) or (5c) to claims about (5b). Nor does it mean that these questions cannot be (and aren’t) pursued in tandem. As a matter of practice, all three questions are often addressed simultaneously. Rather, what we intend by ‘precondition’ is that it is nugatory to pursue the latter questions without some answers to the previous ones given the nature of the questions asked.

One more caveat before we get the show on the road: this is a very idealized account of the problem and its empirical boundary conditions. For example, most GGers do not think that humans have a single grammar of their NL (i.e. it is almost certain that humans develop multiple Gs). Nor do they think that NLs are natural kinds (e.g. looked at carefully, there is nothing like ‘English’ that all so-called speakers of English speak). Ontologically, Gs are more “real” than the NLs they are related to. However, this idealization is adopted because it is recognized that describing the properties of G and FL even under these idealized assumptions is already very hard and, GGers believe that relaxing the idealization will not significantly affect the shape of the answers provided.  Of course, this may be incorrect. But we doubt it and what follows does not much worry about the legitimacy of so idealizing.[5]

2. So given these three questions, we can divide the WH of GG into three epochs. Early work in GG (say from the mid 50s to the early 70s) concentrated on constructing sample Gs for fragments of a given NL.  The second epoch goes from the mid 70s to the early 90s. This concentrated on simplifying the rules/operations of FL. This involved, categorizing the various rule types a G could have, factoring out common features within these types and enriching UG to prevent massive over-generation (a natural consequence of simplifying the rules, as we shall see). The third epoch goes from the mid 90s to the present. This period focuses on simplifying FL, trying to figure out which aspects of FL’s properties are language specific and which follow from more general cognitive/ and/or computational principles. The aim here has been to factor out those features of FL that are computationally general, leaving, it is hoped, a very small domain specific residue. So, three epochs: (i) exploring the rules Gs contain and how they interact, (ii) simplifying the structure of Gs by articulating the structure of UG and (iii) simplifying FL by separating the computationally general wheat from the linguistically specific chaff.

In what follows we describe in slightly more detail the kinds of results each epoch delivered and illustrate the progressive nature of the GG enterprise. Let’s begin with the first epoch.


[1] Though this sounds like an obvious method to pursue in studying the structure of FL, it is actually quite a bit harder to do than one might think. The reason is that Gs do not tend to have the same rules. What they have in common is far more abstract: e.g. all the rules adhere to a common rule schema, or all the rules obey similar constraints in the sense of no rules within Gs showing evidence of disobeying them.  This makes any simple-minded process of looking for commonalities quite difficult.
[2] It is worth observing that the POS sets a very high standard for concluding that some property reflects innate structural features of FL. POS assumes that any feature of G that could be a data driven acquisition is one.  However, this conclusion does not follow. Nonetheless, POS allows one to isolate a whole slew of promising candidate generalizations useful in probing the built-in structure of FL.
[3] Even if natural selection can operate quickly, time pressures may matter.
[4] There are others: How is FL instantiated in the brain? How are Gs used in linguistic performance? How is FL used to acquire a G in real time? We return to these.
[5] This does not mean to say that GGers have not explored models that relax these idealizations. For example, a staple of work in historical linguistics is to assume that speakers have multiple Gs that compete. Change is then modeled as the changing dominance relations between these multiple Gs. 

Saturday, March 14, 2015

Defeating linguistic agnotology; or why very vigorous attack on bad ideas is important

In a recent piece I defended the position that intemperance in debate is defensible. More specifically, I believe that there are some positions that are so silly and/or misinformed that treating them with courtesy amounts to a form of disinformation. As any good agnotologist will tell you, simply being admitted to the debate as a reasonable option is 90% of the battle. So, for example, the smoking lobby became toast once it was widely recognized that the "science" it relied on was not wrong, but laughable and fraudulent. You can fight facts with other facts, but once a position is the subject of late night comedy, it's a goner.

Why should this be so? Omer Preminger sent me this piece in discussing the Dunning-Kruger Effect (DKE), something that I had never heard of before. It is described as follows:

 ….those with limited knowledge in a domain suffer from a dual burden: Not only do they reach mistaken conclusions and make regrettable errors, but their incompetence robs them of the ability to realize it.

If the DKE is real, then it is not particularly surprising that some debates are interminable. The only constraint on their longevity will be the energy that the parties bring to the "debate." In other words, holding a silly uninformed position is not something that will become evident to the holder no matter how clearly this is pointed out. Thus, the idea that gentle debate will generally result in consensus, though a nobel ideal, cannot be assumed to be a standard outcome.

The piece goes on to describe a second apparent fact. It seems that when there is "disagreement" people tend to split the difference.

people have an “equality bias” when it comes to competence or expertise, such that even when it’s very clear that one person in a group is more skilled, expert, or competent (and the other less), they are nonetheless inclined to seek out a middle ground in determining how correct different viewpoints are.

I've seen this in action, especially as regards non-experts. I call it the three remark rule. You go to a talk and someone says something that you believe might be controversial. Nobody in the audience objects. The third time there is no objection, you assume that the point made was true. After all, you surmise, were it false someone would have said so. Of course, there are many other reasons not to say anything. Politeness being the highest motive on the list. However, in such cases reticence, even when motivated by the best social motives, have intellectual consequences. Saying nothing, being "nice" can spread false belief. That's why, IMO, it is incumbent on linguists to be very vociferous.

Let me put this another way: reasonable debate among reasonable positions is, well, reasonable. However, there are many views, especially robust in the popular intellectual press, that we know to be garbage. The Everett and Evans stuff on universals is bunk. It is without any intellectual interest once the puns are clarified. Much of the stuff one finds from empiricist minded colleagues (read, psychologists) is similarly bereft of insight or relevance. These positions should not be treated graciously. So, nest time you are at a talk and hear these views expounded, make it clear that you think that they are wrong, or better still miss the point or beg the relevant questions. If you can do this with a little humor, so much the better. Whatever you do, do not rest the views with respect. Do this and we have every reason to believe that this will help the silliness spread.

Thursday, March 12, 2015

Getting a grant

Kleanthes Grohmann (recently featured here) sent me this paper (here) which surveys the costs and benefits of grant writing. The participants are from astronomy and psychology (social and personality). The paper is wroth taking a look at for it confirms what many might have thought reflecting on their own experience. So, for example, grant writing is VERY labor intensive. Conservatively, it seems to require about 10 times the amount of time of other scholarly pursuits. Moreover, large amounts of good work is never funded and the collateral benefits of writing the grant and not getting funded seem pretty mild. At any rate, whatever their collateral benefits, it does not seem that these are sufficient to inspire more grant writing effort.

There is more that is interesting:


  • Success does not seem to be gender biased
  • Funding is very competitive, more so if one has not landed a grant in the few years before applying
  • Amount of time spent on writing grant not apparently correlated with greater funding
  • There is a correlation between applying for more grants and getting more funding, though "the causal direction might go either way"
  • Many good things don't get funded
None of these results are surprising. But take a look. It's interesting. 

Sunday, March 8, 2015

A good Sunday read

Bob Berwick sent me this old review by Richard Lewontin. I'm a big fan of Lewontin's. His paper in the Invitation to CogSci on the evolution of cognition (here) is (or should be) standard reading for anyone interested in Evolang. At any rate, this review of a Carl Sagan book is a terrific (funny) discussion of the scientific method (there is none), the overhyping of science by scientists, the increasing cynicism of the "unwashed" towards this overhyping, science as show business, the cultural bases of (some of) the public's antipathy towards science in the USA, and more. If I had read this before, I didn't remember. I certainly enjoyed it this time around.

Saturday, March 7, 2015

How to make $1,000,000

My mother once told me about an easy way to become a millionaire: start with $10 million. This seems to be advice that second generations are particularly good at following. And not only as regards inter-generational wealth transfer. Like families, journals also enjoy life cycles, with founders giving way to a next generation. And as in families, regression to the mean (i.e. the headlong rush to average) seems to be an inexorable force. However, in contrast to the rise and decline of wealth in families, the move from provocative to staid in journals is rarely catalogued. Rarely, but not never. Here is a paper (by Priva and Austerweil (P&A)) that charts the change in intellectual focus of Cognition, a (once?) very important cogsci journal. What does P&A show? Two things: (i) It shows that the mix of papers in the journal has substantially changed. Whereas in the beginning, there was a fair mix of theory and experimental papers (theory papers predominating), since the mid 2000s the mix has dramatically changed, with experimental papers forming the bulk of academic product. Theory papers have not entirely disappeared, but they have been substantially overtaken by their experimental kin. (ii) That papers on language and development have gone from central topics of interest to a somewhat marginal presence.[1]

How surprising is this?  Let me start by discussing (i), the decline of “theory,” a favorite obsession of mine (see here). Well, first off, from one common perspective, some decline might be expected. We all know the Kuhnian trope; “revolutionary” periods of scientific inquiry where paradigms are contested, big ideas are born and old ones die out (one old fogey at a time in Plank time) give way to periods of “normal science” where the solid majestic wall of scientific accomplishment is carefully and industriously built brick by careful empirical brick. The picture offered is one in which the basic framework ideas get hashed out and then their implications are empirically fleshed out. I never really liked this way of conceptualizing matters (there is a lot of hashing and fleshing going on all the time in serious work), but I think that this picture has some appeal descriptively. Sadly, it also seems to have normative attractions, especially to next generation editors. Here’s what I mean.

Editing a journal is a lot of work. Much of it thankless. So before I go off the deep end here in a minute, let me personally thank those that take this task on, for their work and commitment is invaluable and what we think of as science could not succeed without such effort. That said, precisely because of how hard it is to do, you need to be driven (nuts?) to start a journal. What drives you? The feeling that there is something new to say but that there is no good place to say it. Moreover, not only is that something new, it must be important and new. And it is not possible to say these new important things in the current journals because the new ideas cut things up in new ways or approach problems from premises that don’t fit into the existing journalistic matrix.[2] So, at the very least, the extant venues are not congenial places to publish, and in some cases are outright hostile.

The emergence of cogsci (something that happened when I was growing up intellectually) had this feel to it. There was a self-conscious cognitive revolution, with very self-conscious revolutionaries. Furthermore, this revolution was fought on several fronts: linguistics, psychology, computer science and philosophy being the four main ones. Indeed, for a while, it was not clear where one left off and the other began. Linguists read philosophy and psychology papers, psychologists knew about Transformational Grammar and could Locke from Descartes, philosophers debated what to make of innate ideas, representations and rule following based on work in linguistic and psychology and computer scientists (especially in AI) worried about the computational properties of mental representations (think Marr and Marcus for example).  Cogsci lived at the intersection of these many disciplines, was nurtured by their cross disciplinary discussions and, for someone like me, cogsci became identified as the investigation of the structures of minds (and one day brains) using the techniques and methods of thought that each discipline brought to the feast. Boy was this exciting. Not surprisingly, the premiere journal for the advancement of this vision was Cognition. Why not surprisingly? Because the founding editors, Jacques Mehler and Tom Bever, were two people that thoroughly embodied this new combined intellectual vision (and were and are two of its leading lights) and they built Cognition to reflect it.[3]

A nice way of seeing this is to read Mehler’s “farewell remarks” here. It is very explicit about what gap the journal was intended to fill:

Our aim was to change the publishing landscape in psychology and related disciplines that became part of “Cognitive Science.” …[P]sychology had turned almost exclusively into an experimental discipline with an overt disdain for theory…Linguistics had become a descriptive discipline often favoring normative or purely descriptive over theoretical approaches. Professional journals in line with this outlook generally obliged contributors to write their papers in standard format that privileged the shortest possible introductions and conclusions, methods and procedures used in experiments. Papers by non-experimental scientists say, philosophers of mind or theoretical linguists, were rarely even accepted…. (p. 7)

In service of this, the journal was the venue of lots of BIG debates concerning connectionism, the representational theory of mind, compositionality, AI models of mind, prototypes, domain specificity, computational complexity, core knowledge and much much more. In fact, Cognition did something almost miraculous: It became a truly inter-disciplinary journal, something that administrators and science bureaucrats (including publishers) love to talk about (but, it seems, often fail appreciate when it happens).

P&A records that this Cognition now seems to be largely gone. It is no longer the journal its editors founded. There is little philosophy and little linguistics or linguistically based psychology. Nor does it seem to any longer be the venue where big ideas are thrashed out. Three illustrations: (i) the critical discussions concerning Bayesian methods in psychology have not occurred in the pages of Cognition,[4] (ii) nor have the Gallistel-like critiques of connectionist neuro-science gotten much of an airing, (iii) nor have extensive critiques of resurgent “language” empiricism (e.g. Tomasello) made an appearance. These have gotten play elsewhere, and that is a good thing, but these dogs have not barked in Cognition, and their absence is a good indicator of how much Cognition has changed. Moreover, this change is no accident. It was policy.

How so? Well, in the same issue that Mehler penned his farewell the new incoming editor Gerry Altmann gave his inaugural editorial (here). It’s really worth reading the Mehler and Altmann pieces side by side, if nothing else as an exercise in the sociology of science. I’ve rarely read anything that so embodies (and embraces) the Kuhnian distinction between revolutionary vs normal science. Altmann’s editorial is six pages long. After some standard boilerplate thanking Mehler & Co. for its path-breaking efforts, Altmann sets out his vision of the future. It comes in two parts.

First the ideal paper:

To be published in Cognition, articles must be robust in respect of the fit between the theory, the data and the literature in which the work is grounded. They should have a breadth to them that enables the specific research they describe to make contact with more general issues in cognition; the more explicit this contact, the greater the impact of the research beyond the confines of the specialized research community. (2)

It’s worth contrasting this ideal with the more expansive one provided by Mehler above. In Altmann’s, there is already an emphasis on “data” that was missing from Mehler’s discussion. In other words, Altmann’s ideal has an up front experimental tilt. Data’s the lede. The vision thing is filler. To see this, read the two sentences in reverse order. The sense of what is important changes. In the actual quoted order what matters is data fit then idea quality. Reverse the sentences and we get first idea quality and then data fit. Moreover, unlike Mehler’s pitch, what’s clear here is that Altmann does not envision papers that might be good and worthwhile even were they bereft of data to fit. It more or less assumes that the conceptual issues that were at the foundation of the cogsci revolution have all been thoroughly investigated and understood (or were largely irrelevant to begin with (dare I say, maybe even BS?)). More charitably, it assumes that if something new does rise under the cognitive sun, it will arise from the carefully fitted data. In short, the main job of the cogscientist is see how the theories fit the facts (or vice versa). Theory alone is so your grandparent’s cognition.

The second part of the editorial reinforces this reading. The last 3 pages (i.e. half the editorial), section 3, concerns “the appropriate analyses of data” (4). It’s a long discussion of what stats to use and how to use them. There is no equally long section discussing thos hot topics/problems, what issues are worth addressing and why. This reinforces the conclusion that what Cognition will henceforth worry most about is data fit and experimental procedure. Sounds like the kind of journal that Mehler and Bever had hoped that Cognition would displace. Indeed, prior to Cognition’s founding, psychology had lots of the kinds of Journals that the Altmann editorial aspires to. That’s precisely why Mehler and Bever started their journal. Altmann appears to think that psychology needs one more.

If this read is right, then it is not surprising that Cognition’s content profile has changed over the years. It is not merely that new topics get hot and old ones get stale. Rather, it is that what was once a journal interested in bridging disciplines, critically investigating big issues and provoking thought, “grew up” and happily assumed the role of purveyor of “normal” science. A nice well behaved journal, just like most of the others.

Last two points. Given the apparent dearth of interest in theory, it is not a surprise (to me) that work on language is less represented in the new Cognition. Anything that takes linguistic theory seriously in psycho study will be suspect to those with a great respect for psychological techniques (we don’t gather data the right way, there is a distance between competence and performance, we think that minds are not all purpose learners etc.). Thus taking results in linguistic theory as starting points will go against the intellectual grain where theory is less important than data points. This need not have been so. But that it is so is not surprising.

Second, there is a weird part of Altmann’s editorial concerning the “collaborative” nature of science and how this should be reflected in the “editorial structure” of the journal. Basically, it seems to be signaling a departure from past methods. I don’t really know how the Mehler era operated “editorially.” But it would not surprise me were he (and Bever) more activist editors than is commonplace. This would go far, IMO, in explaining why the old Cognition was such a great journal. It expressed the excellent taste of its leaders. This is typically true of great journals. At one time the leading figures edited journals and imposed their tastes on the field, to its benefit. Max Planck edited the journals that published Einstein’s groundbreaking (and very unconventional) papers.[5] Keynes edited the most important economics journal of his day. Mehler and Bever were intellectual leaders and Cognition reflected their excellent taste in questions and problems. It strikes me that the Altmann editorial is a non too subtle critique of this. It’s saying that going forward editorial decisions would be more balanced and shared. In other words, more watered down, more common denominatorish, less quirky, more fashionnable. There is room for this ideal, one where the aim is to reflect the scientific “consensus.” Today, in fact, this is what most journals do. Mehler and Bever’s Cognition did not.

To end: Cognition has changed. Why? Because it wanted to. It has managed to achieve exactly what the new regime was aiming for. The old Cognition stood apart, had a broad vision and had the courage of its new ideas. The new Cognition has re-joined the fold. A good journal (no doubt). But no longer a distinctive one. It’s not where people go to see the most important ideas in cognition vigorously debated. It’s become a professional’s journal, one among many. Does it publish good papers? Sure. Is it the indispensible journal in cogsci that it once was? Not so much. IMO, that’s really too bad. However, it is educational, for now you know how to make $1,000,000 dollars. Just be sure to start off with $10,000,000.




[1] This is all premised on the assumption that the topic model methodology used in the paper accurately reflects what has been going on. This may be incorrect. However, I confess that it accurately reflects what many people I know have noted anecdotally.
[2] Is this PoMo or what? With a tinge of the Wachowskis thrown in.
[3] And you know many of the others. To name a few: Chomsky, two Fodors, Gleitman, Gallistel, Katz, Pylyshyn, Gellman, Garrett, Block, Carey, Spelke, Berwick, Marr, Marcus, a.o.
[4] E.g. Eberhardt & Danks, Brown & Love, Bowers & Davis, Marcus have all appeared in other venues. See here, here and here form some discussion and references.
[5] A friend of mine in theoretical physics once told me that he doubted that papers like Einstein’s great 1905 quartet could be published today. Even by the standards in 1905 they looked strange. Moreover, they were from a nobody working in a patent office. It’s a good thing for Einstein that Planck, one of the leading physicists of his day, was the editor of Annalen der Physik.