Comments

Thursday, July 13, 2017

Some recent thoughts on AI

Kleanthes sent me this link to a recent lecture by Gary Marcus (GM) on the status of current AI research. It is a somewhat jaundiced review concluding that, once again, the results have been strongly oversold. This should not be surprising. The rewards to those that deliver strong AI (“the kind of AI that would be as smart as, say a Star Trek computer” (3)) will be without limit, both tangibly (lots and lots of money) and spiritually (lots and lots of fame, immortal kinda fame). And given hyperbole never cripples its purveyors (“AI boys will be AI boys” (and yes, they are all boys)), it is no surprise that, as GM notes, we have been 20 years out from solving strong AI for the last 65 years or so. This is a bit like the many economists who predicted 15 of the last 6 recessions but worse. Why worse? Because there have been 6 recessions but there has been pitifully small progress on strong AI, at least if GM is to be believed (and I think he is). 

Why despite the hype (necessary to drain dollars from “smart” VC money) has this problem been so tough to crack? GM mentions a few reasons.

First, we really have no idea how open ended competence works. Let me put this backwards. As GM notes, AI has been successful precisely in “predefined domains” (6). In other words, where we can limit the set of objects being considered for identification or the topics up for discussion or the hypotheses to be tested we can get things to run relatively smoothly. This has been true since Winograd and his block worlds. Constrain the domain and all goes okishly. Open the domain up so that intelligence can wander across topics freely and all hell breaks loose. The problem of AI has always been scaling up, and it is still a problem. Why? Because we have no idea how intelligence manages to (i) identify relevant information for any given domain and (ii) use that information in relevant ways for that domain. In other words, how we in general figure out what counts and how we figure out how much it counts once we have figured it out is a complete and utter mystery. And I mean ‘mystery’ in the sense that Chomsky has identified (i.e. as opposed to ‘problem’).

Nor is this a problem limited to AI.  As FoL has discussed before, linguistic creativity has two sides. The part that has to do with specifying the kind of unbounded hierarchical recursion we find in human Gs has been shown to be tractable. Linguists have been able to say interesting things about the kinds of Gs we find in human natural languages and the kinds of UG principles that FL plausibly contains. One of the glories (IMO, the glory) of modern GG lies in its having turned once mysterious questions into scientific problems. We may not have solved all the problems of linguistic structure but we have managed to render them scientifically tractable.

This is in stark contrast to the other side linguistic creativity: the fact that humans are able to use their linguistic competence in so many different ways for thought and self-expression. This is what the Cartesians found so remarkable (see here for some discussion) and that we have not made an iota of progress understanding. As Chomsky put it in Language & Mind (and is still a fair summary of where we stand today):

Honesty forces us to admit that we are as far today as Descartes was three centuries ago from understanding just what enables a human to speak in a way that is innovative, free from stimulus control, and also appropriate and coherent. (12-13)[1]

All-things-considered judgments, those that we deploy effortlessly in every day conversation, elude insight. That we do this is apparent. But how we do this remains mysterious. This is the nut that strong AI needs to crack given its ambitions. To date, the record of failure speaks for itself and there is no reason to think that more modern methods will help out much.

It is precisely this roadblock that limiting the domain of interest removes. Bound the domain and the problem of open-endedness disappears.

This should sound familiar. It is the message in Fodor’s Modularity of Mind. Fodor observes that modularity makes for tractability. When we move away from modular systems, we flat on our faces precisely because we have no idea how minds identify what is relevant in any given situation and how it weights what is relevant in a given situation and how it then deploys this information appropriately. We do it all right. We just don’t know how.

The modern hype supposes that we can get around this problem with big data. GM has a few choice remarks about this. Here’s how he sees things (my emphasis):

I opened this talk with a prediction from Andrew Ng: “If a typical person can do a mental task with less than one second of thought, we can probably automate it using AI either now or in the near future.” So, here’s my version of it, which I think is more honest and definitely less pithy: If a typical person can do a mental task with less than one second of thought and we can gather an enormous amount of directly relevant data, we have a fighting chance, so long as the test data aren’t too terribly different from the training data and the domain doesn’t change too much over time. Unfortunately, for real-world problems, that’s rarely the case. (8)

So, if we massage the data so that we get that which is “directly relevant” and we test our inductive learner on data that is not “too terribly different” and we make sure that the “domain doesn’t change much” then big data will deliver “statistical approximations” (5). However, “statistics is not the same thing as knowledge” (9). Big data can give us better and better “correlations” if fed with “large amounts of [relevant!, NH] statistical data”. However, even when these correlational models work, “we don’t necessarily understand what’s underlying them” (9).[2]

And one more thing: when things work it’s because the domain is well behaved. Here’s GM on AlphaGo (my emphasis):

Lately, AlphaGo is probably the most impressive demonstration of AI. It’s the AI program that plays the board game Go, and extremely well, but it works because the rules never change, you can gather an infinite amount of data, and you just play it over and over again. It’s not open-ended. You don’t have to worry about the world changing. But when you move things into the real world, say driving a vehicle where there’s always a new situation, these techniques just don’t work as well. (7)
 
So, if the rules don’t change, you have unbounded data and time to massage it and the relevant world doesn’t change, then we can get something that approximately fits what we observe. But fitting is not explaining and the world required for even this much “success” is not the world we live in, the world in which our cognitive powers are exercised. So what does AI’s being able to do this in artificial worlds tell us about what we do in ours? Absolutely nothing.

Moreover, as GM notes, the problems of interest to human cognition have exactly the opposite profile. In Big Data scenarios we have boundless data, endless trials with huge numbers of failures (corrections). The problems we are interested in are characterized by having a small amount of data and a very small amount of error. What will Big Data techniques tell us about problems with the latter profile? The obvious answer is “not very much” and the obvious answer, to date, has proven to be quite adequate.

Again, this should sound familiar. We do not know how to model the everyday creativity that goes into common judgments that humans routinely make and that directly affects how we navigate our open-ended world. Where we cannot successfully idealize to a modular system (one that is relatively informationally encapsulated) we are at sea. And no amount of big data or stats will help.

What GM says has been said repeatedly over the last 65 years.[3] AI hype will always be with us. The problem is that it must crack a long lived mystery to get anywhere. It must crack the problem of judgment and try to “mechanize” it. Descartes doubted that we would be able to do this (indeed this was his main argument for a second substance). The problem with so much work in AI is not that it has failed to crack this problem, but that it fails to see that it is a problem at all. What GM observes is that, in this regard, nothing has really changed and I predict that we will be in more or less the same place in 20 years.

Postscript:

Since penning(?) the above I ran across a review of a book on machine intelligence by Gary Kasparov (here). The review is interesting (I have not read the book) and is a nice companion to the Marcus remarks. I particularly liked the history on Shannon’s early thoughts on chess playing computers and his distinction on how the problem could be solved:

At the dawn of the computer age, in 1950, the influential Bell Labs engineer Claude Shannon published a paper in Philosophical Magazine called “Programming a Computer for Playing Chess.” The creation of a “tolerably good” computerized chess player, he argued, was not only possible but would also have metaphysical consequences. It would force the human race “either to admit the possibility of a mechanized thinking or to further restrict [its] concept of ‘thinking.’” He went on to offer an insight that would prove essential both to the development of chess software and to the pursuit of artificial intelligence in general. A chess program, he wrote, would need to incorporate a search function able to identify possible moves and rank them according to how they influenced the course of the game. He laid out two very different approaches to programming the function. “Type A” would rely on brute force, calculating the relative value of all possible moves as far ahead in the game as the speed of the computer allowed. “Type B” would use intelligence rather than raw power, imbuing the computer with an understanding of the game that would allow it to focus on a small number of attractive moves while ignoring the rest. In essence, a Type B computer would demonstrate the intuition of an experienced human player.

As the review goes on to note, Shannon’s mistake was to think that Type A computers were not going to materialize. They did, with the result that the promise of AI (that it would tell us something about intelligence) fizzled as the “artificial” way that machines became “intelligent” simply abstracted away from intelligence. Or, to put it as Kasparov is quoted as putting it:  “Deep Blue [the machine that beat Kasparov, NH] was intelligent the way your programmable alarm clock is intelligent.”

So, the hope that AI would illuminate human cognition rested on the belief that technology and brute calculation would not be able to substitute for “intelligence.” This proved wrong, with machine learning being the latest twist in the same saga, per the review and Kasparov. 

All this fits with GM’s remarks above. What both do not emphasize enough, IMO, is something that many did not anticipate; namely that we would revamp our views of intelligence rather than question whether our programs had it.  Part of the resurgence of Empiricism is tied to the rise of the technologically successful machine. The hope was that trying to get limited machines to act like we do might tell us something about how we do things. The limitations of the machine would require intelligent design to get it to work thereby possibly illuminating our kind of intelligence. What happened is that getting computationally miraculous machines to do things in ways that we had earlier recognized as dumb and brute force (and so telling us nothing at all) has transformed into the hypothesis that there is no such things as real intelligence at all and everything is “really” just brute force. Thus, the brain is just a data cruncher, just like Deep Blue is. And this shift in attitude is supported by an Empiricist conception of mind and explanation. There is no structure to the mind beyond the capacity to mine the inputs for surfacy generalizations. There is no structure to the world beyond statistical regularities. On this Eish viw, AI has not failed, rather the right conclusion is that there is less to thinking than we thought. This invigorated Empiricism is quite wrong. But it will have staying power. Nobody should underestimate the power that a successful (money making) tech device can have on the intellectual spirit of the age.


[1] Chomsky makes the same point recently, and he is still right. See here for discussion and links to article.
[2] This should again sound familiar. It is the moral that Chang drew on her work on faces as discussed here.
[3] I myself once made similar points in a paper with Elan Dresher. Few papers have been more fun to write. See here for the appraisal (if interested).

Thursday, July 6, 2017

The logic of adaptation

I recently ran across a nice paper on the logic of adaptive stories (here), along with a nice short discussion of its main points (here) (by Massimo Pigliucci (P)). The Olson and Arroyo-Santos paper (OAS) argues that circularity (or “loopiness”) is characteristic of all adaptive explanations (indeed, of all non-deductive accounts) but that some forms of loopiness are virtuous while others are vicious. The goal, then, is to identify the good circular arguments from the bad ones, and this amounts to distinguishing small uninteresting circles from big fat wide ones. Good adaptive explanations distinguish themselves from just-so stories in having independent data afforded from the three principle kinds of arguments evolutionary biologists deploy. OAS adumbrates the forms of these arguments and uses this inventory to contrast lousy adaptive accounts from compelling ones. Of particular interest to me (and I hope FoLers) is the OAS claim that looking at things in terms of how fat a circular/loopy account is will make it easy to see why some kinds of adaptive stories are particularly susceptible to just-soism. What kinds? Well ones like those applied to the evolutions of language, as it turns out. Put another way, OAS leads to Lewontin like conclusions (see here) from a slightly different starting point.

An example of a just-so story helps to illustrate the logic of adaptation that OAS highlights.  Why do giraffes have long necks? So as to be able to eat leaves from tall trees. Note, that giraffes eat from tall trees confirms that having long necks is handy for this activity, and the utility of being able to eat from tall trees would make having a long neck advantageous. This is the loopiness/circularity that OAS insists is part of any adaptational account. OAS further insists that this circularity is not in itself a problem. The problem is that in the just-so case the circle is very small, so small as to almost shrink to a point. Why? Because the evidence for the adaptation and the fact that the adaptation explains is the same: tall necks are what we want to explain and also constitute the evidence for the explanation. As OAS puts it:

…the presence of a given trait in current organisms is used as the sole evidence to infer heritable variation in the trait in an ancestral population and a selective regime that favored some variants over others. This unobserved selective scenario explains the presence of the observed trait, and the only evidence for the selective scenario is trait presence (168).

In other words, though ‘p implies p’ is unimpeachably true, it is not interestingly so. To get some explanation out of an account that uses these observations we need a broader circle. We need a big fat circle/loop, not an anorexic one.

OAS’s main take home message is that fattening circles/loops is both eminently doable (in some cases at least) and is regularly done. OAS lists three main kinds of arguments that biologists use to fatten up an adaptation account: comparative arguments, population arguments, and optimality arguments. Each brings something useful to the table. Each has some shortcomings. Here’s how OAS describes the comparative method (169):

The comparative method detects adaptation through convergence (Losos 2011). A basic version of comparative studies, perhaps the one underpinning most state- ments about adaptation, is the qualitative observation of similar organismal features in similar selective contexts.

The example OAS discusses is the streamlined body shapes and fins in animals that live in water. The observation that aquatic animals tend to be sleek and well built for moving around in water strongly suggests that there is something about the watery environment that is driving the observed sleekness.  As this example illustrates, a hallmark of the comparative method is “the use of cross-species variation” (170). The downside of this method is that it “does not examine fitness or heritability directly” and it “often relies on ancestral character state reconstructions or assumptions of tempo and mode that are impossible to test” (171, table 1).

A second kind of argument focuses on variations in a single population and sees how this affects “heritability and fitness between potentially competing individuals” (171). These kinds of studies involve looking at extant populations and seeing how their variations tie up with heritability. Again OAS provides an extensive example involving “the curvature of floral nectar spurs” in some flowers (171) and shows how variation and fitness can be precisely measured in such circumstances (i.e. where it is possible to do studies of  “very geographically and restricted sets of organisms under often unusual circumstances” (172)).

This method, too, has a problem.  The biggest drawback is that the population method “examines relatively minor characters that have not gone to fixation” and “extrapolation of results to multiple species and large time scales” is debatable (171, table 1). In other words, it is not that clear whether the situation in which population arguments can be fully deployed reveal the mechanisms that are at play “in generating the patterns of trait distribution observed over geological time and clades” because it is unclear whether the “very local population phenomena are…isomporphic with the factors shaping life on earth at large” (172).

The third type of argument involves optimality thinking. This aims to provide an outline of the causal mechanisms “behind a given variant being favored” and rests on a specification of the relevant laws driving the observed effect (e.g. principles of hydronamics for body contour/sleekness in aquatic animals). The downside to this mode of reasoning is that it is not always clear what variables are relevant for optimization.

OAS notes that adaptive explanations are best when one can provide all three kinds of reasons (as one can in the case, for example, of aquatic contour and sleekness (see figure 4 and the discussion in P). Accounts achieve just-so status when none of the three methods can apply and none have been used to generate relevant data. The OAS discussion of these points is very accessible and valuable and I urge you take a look.

The OAS framing also carries an important moral, one that both OAS and P note: if going from just-so to serious requires fattening with comparative, population and optimization arguments then some fashionable domains of evolutionary speculation relying on adaptive consideration are likely to be very just-soish. Under what circumstances will getting beyond hand waving prove challenging? Here’s OAS (184, my emphasis):

Maximally supported adaptationist explanations require evidence from comparative, populational, and optimality approaches. This requirement highlights from the outset which adaptationist studies are likely to have fewer layers of direct evidence available. Studies of single species or unique structures are important examples. Such traits cannot be studied using comparative approaches, because the putatively adaptive states are unique (cf. Maddison and FitzJohn 2015). When the traits are fixed within populations, the typical tools of populational studies are unavailable. In humans, experimental methods such as surgical intervention or selective breeding are unethical (Ruse 1979). As a result, many aspects of humans continue to be debated, such as the female orgasm, human language, or rape (Travis 2003; Lloyd 2005; Nielsen 2009; MacColl 2011). To the extent that less information is available, in many cases it will continue to be hard to distinguish between different alternative explanations to decide which is the likeliest (Forber 2009).

Let’s apply these OAS observations to a favorite of FoLers, the capacity for human language. First, human language capacity is, so far as we can tell, unique to humans. And it involves at least one feature (e.g. hierarchical recursion) that, so far as we can tell, emerges nowhere else in biological cognition. Hence, this capacity cannot be studied using comparative methods. Second, it cannot be studied using population methods, as, modulo pathology, the trait appears (at least at the gross level) fixed and uniform in the human species (any kid can learn any language in more or less the same way). Experimental methods, which could in principle be used (for there probably is some variation across individuals in phenomena that might bear on the structure of the fixed capacity (e.g. differences in language proficiency and acquisition across individuals) will, if pursued, rightly land you in jail or at the World Court in the Hague. Last, optimization methods also appear useless for it is not clear what function language is optimized for and so the dimensions along which it might be optimized are very obscure.  The obvious ones relating to efficient information transmission are too fluffy to be serious.[1]

P makes effectively the same point, but for evo-psych in general, not just evo-lang. In this he reiterates Lewontin’s earlier conclusions. Here is P:

If you ponder the above for a minute you will realize why this shift from vicious circularity to virtuous loopiness is particularly hard to come by in the case of our species, and therefore why evolutionary psychology is, in my book, a quasi-science. Most human behaviors of interest to evolutionary psychologists do not leave fossil records (i); we can estimate their heritability (ii) in only what is called the “broad” sense, but the “narrow” one would be better (see here); while it is possible to link human behaviors with fitness in a modern environment (iii), the point is often made that our ancestral environment, both physical and especially social, was radically different from the current one (which is not the case for giraffes and lots of other organisms); therefore to make inferences about adaptation (iv) is to, say the least, problematic. Evopsych has a tendency to get stuck near the vicious circularity end of Olson and Arroyo-Santos’ continuum.

There is more, much more, in the OAS paper and P's remarks are also very helpful. So those interested in evolang should take a look. The conclusion both pieces draw regarding the likely triviality/just-soness of such speculations is a timely re-re-re-reminder of Lewontin and the French academy’s earlier prescient warnings. Some questions, no matter how interesting, are likely to be beyond our power to interestingly investigate given the tools at hand.

One last point, added to annoy many of you. Chomsky’s speculations, IMO, have been suitably modest in this regard. He is not giving an evolang account so much as noting that if there is to be one then some features will not be adaptively explicable. The one that Chomsky points to is hierarchical recursion. Given the OAS discussion it should be clear that Chomsky is right in thinking that this will not be a feature liable to an adaptive explanation. What would “variation” wrt Merge be? Somewhat recursive/hierarchical? What would this be and how would the existence of 1-merge and 2-merge systems get you to unbounded Merge? It won’t, which is Chomsky’s (and Dawkins’) point (see here for discussion and references). So, there will be no variation and no other animals have it and it doesn’t optimize anything. So there will be no available adaptive account. And that is Chomsky’s point! The emergence of FL whenever it occurred was not selected for. Its emergence must be traced to other non adaptive factors. This conclusion, so far as I can tell, fits perfectly with OAS’s excellent discussion. What Chomsky delivers is all the non-trivial evolang we are likely to get our hands on given current methods, and this is just what OAS, P and Lewontin should lead us to expect.



[1] Note that Chomsky’s conception of optimal and the one discussed by OAS are unrelated. For Chomsky, FL is not optimized for any phenotypic function. There is nothing that FL is for such that we can say that it does whatever better than something else might. For example structure dependence has no function so that Gs that didn’t have it would be worse in some way than ones (like ours) that do.

Friday, June 30, 2017

Statistical obscurantism; math destruction take 2

I've mentioned before that statistical knowledge can be a dangerous thing (see here). It's a little like Kabbala, something that is dangerous in the hands of inexperienced, the ambitious and lazy.  This does not mean that in its place stats are not valuable tools. Of course they are. But there is a reason for the slogan "lies, damn lies and statistics." A few numbers can cover up the most awful thinking, sort of like pretty pix of brains in the NYT can sell almost any new cockamamie idea in cog-neuro. So, in my view, stats is a little like nitroglycerine; useful but dangerous on unsteady ground.

Now, even I don't really respect my views on these matters. What the hell do I know, really? Well, very little. So I will buck this view up by pointing you to an acknowledged expert on the subject who has come to a very similar conclusion. Here is Andrew Gelman despairing of the view that done right stats is the magic empirical elixir, able to get something out of any data set, able to spin scientific gold from any experimental foray:

In some sense, the biggest problem with statistics in science is not that scientists don’t know statistics, but that they’re relying on statistics in the first place.
How is stats the problem? Because it covers up dreadful thinking:
Just imagine if papers such as himmicanes, air rage, ages-ending-in-9, and other clickbait cargo-cult science had to stand on their own two feet, without relying on p-values—that is, statistics—to back up their claims. Then we wouldn’t be in this mess in the first place.
So, one problem with stats is that they can make drek look serious. Is this a problem with the good use of stats? No, but given the current culture, it is a problem. And as these pair of quotes suggests, if something absent the stats sounds dumb, then one should be very very very wary of the stats. In fact, one might go further: if the idea sans stats looks dumb then the best reaction on hearing that idea with stats is to reach for your wallet (ore your credulity).

So what does Gelman suggest we do? Well, he is a reasonable man so he says reasonable things:
I’m not saying statistics are a bad idea. I do applied statistics for a living. But I think that if researchers want to solve the reproducibility crisis, they should be doing experiments that can successfully be reproduced—and that involves getting better measurements and better theories, not rearranging the data on the deck of the Titanic.
Yup, it looks like he is recommending thinking. Not a bad idea.  The problem is that stats has the unfortunate tendency of replacing thought. It gives the illusion of being able to substitute technique for insight. Stats are often treated as the Empiricist's perfect tool: it is the method that allows the data speak for itself. And this is the illusion that Gelman is trying to puncture.

German has (given his posts of late) come to believe that this illusion is deeply desired. Here he is again replying to the suggestion that misuse of stats is largely an educational problem:
Not understanding statistics is part of it, but another part is that people—applied researchers and also many professional statisticians—want statistics to do things it just can’t do. “Statistical significance” satisfies a real demand for certainty in the face of noise. It’s hard to teach people to accept uncertainty. I agree that we should try, but it’s tough, as so many of the incentives of publication and publicity go in the other direction.
I would add, you will not be surprised to hear, that there is also the Eish dream I mentioned above wherein the aim is to minimize the human factor mediating data and theory. Rationalists believe that the world must be vigorously interrogated (sometimes put under extreme duress) to reveal its deep secrets. Es don't think that it has deep secrets as they don't really believe that the world has that much hidden structure. Rather the problem is with us: we fail to see what is before our eyes if we gather the data carefully and inspect it with an open heart. The data will speak for itself (which is what stats correctly applied will allow it to do). This Eish vision has its charms. I never underestimate it. I think that it partially lies behind the failure to appreciate Gelman's points.









Wednesday, June 28, 2017

Facing the nativist facts

One common argument for innateness rests on finding some capacity very early on. So imagine that children left the womb speaking Yiddish (the langauge FL/UG makes available with all unmarked values for parameters). The claim that Yiddish was innate would (most likely) not be a hard sell. Actually, I take this back: there will always be unreconstructed Empiricists that will insist that the capacity is environmentally driven, no doubt by some angel that co-habits the womb with the kid all the while sedulously imparting Yiddish competence.

Nonetheless, early manifestation of competence is a pretty good reason for thinking that the manifest capacity rests on biologically given foundations, rather than being the reflex of environmental shaping.  This logic applies quite generally and it is interesting to collect examples of it beyond the language case. The more Eism stumbles the easier it is to ignore it in my own little domain of language.

Here is the report of paper in Current Biology that makes the argument that face recognition is pre-wired in. The evidence? Kids in utero distinguish face like images from others. Given the previous post (here) this conclusion should not be very surprising. There is good evidence that face competence relies on abstract features used to generate a face space. Moreover, these features are not extracted from exemplars and so would appear to be a pre-condition (rather than consequence) for face experience. At any rate, the present article reports on a paper that provides more evidence for this conclusion. Here’s the abstract:

It's well known that young babies are more interested in faces than other objects. Now, researchers have the first evidence that this preference for faces develops in the womb. By projecting light through the uterine wall of pregnant mothers, they found that fetuses at 34 weeks gestation will turn their heads to look at face-like images over other shapes.
Pulling this experiment off required some technical and conceptual breakthroughs: a fancy 4D ultrasound and the appreciation that light could penetrate into the uterus. This realized, the kid in utero responded to faces as infants outside the uterus respond to them. “The findings suggest that babies' preference for faces begins in the womb. There is no learning or experience after birth required.” This does not mean that face recognition is based on innate features. After all, the kid might have acquired the knowledge underlying its discriminative abilities by looking at degraded faces projected through the womb, sort of a fetus’s version Plato’s Cave. This is conceivable, but I doubt that it is believable. Here’s one reason why. It apparently takes some wattage to get the relevant facial images to the in utero kid. Decent reception requires bright lights, hence the author’s following warning:

Reid says that he discourages pregnant mothers from shining bright lights into their bellies.

So, it’s possible that the low passed filter images that the kid sees bouncing around the belly screen is what drives the face recognition capacity. But then the Easter Bunny and Santa Clause are also logically possible.

This work looks ready to push back the data at which kids capacities are cognitively set. First faces, then numbers and quantities. Reid and colleagues are rightly ambitious to push back the time line on the alter two now that faces have been studied. In my view, this kind of evidence is unnecessary as the case for substantial innate machinery was already in place absent this cool stuff (very good for parties and small talk). However, far be it from me to stop others from finding this compelling. What matters is that we dump the blank slate view so that we can examine what the biological givens are. It would be weird were there not substantial innate capacity, not that there is. The question is not whether this is true, but which possible version is.

Last point: for all you skeptics out there: note this is standard infant cognition run in a biology journal. I fail to see any difference in the logic behind this kind of work and analogous work on language. The question is what’s innate. It seems that finding out what is so is a question of biological interest, at least if the publishing venue is a clue. So, to the degree that linguists’ claims bear on the innate mental structures underlying human linguistic facility, to that degree they are doing biology. Unless of course you think that research in biology gets is bona fides via its tools; no 4D ultrasounds and bright lights no biology. But who would ever confuse a discipline with its tools?
  

Wednesday, June 21, 2017

Two things to read

Here are a couple of readables that have entertained me recently.

The first is a NYT report (here) on what is taken to be an iconoclastic view of the role of animal aesthetics in evolution. According to the article, a female’s aesthetic preferences can drive evolutionary change. This, apparently, was Darwin’s view but it appears to be largely out of favor today. More utilitarian/mundane conceptions are favored. Here’s the mainstream view as per the NYT:

All biologists recognize that birds choose mates, but the mainstream view now is that the mate chosen is the fittest in terms of health and good genes. Any ornaments or patterns simply reflect signs of fitness.

The old/new view wants to allow for forces based on fluffier considerations:

The idea is that when they are choosing mates — and in birds it’s mostly the females who choose — animals make choices that can only be called aesthetic. They perceive a kind of beauty. Dr. Prum defines it as “co-evolved attraction.” They desire that beauty, often in the form of fancy feathers, and their desires change the course of evolution.

The bio world contrasts these two approaches, favoring the more “objective” utility-based one over the more “subjective” aesthetic one.  Why? I suspect because the former seems so much more hard-headed and, thus, “scientific.” After all, why would any animal prefer something on aesthetic grounds! If there is no cash value to be had, clearly there is not value to be had at all! (Though this reminds one of the saying about knowing the price of everything and the value of nothing).

An aside: I suspect that this preference hangs on importing the common sense understanding of ‘fitness’ into the technical term. The technical term dubs fit any animal that sends more of its genes into the next generation whatever the reason for this. So, if being a weak effete pretty boy allows greater reproductive success than being a tough successful but ugly looking tough guy than pretty boyhood is fitter than ugly tough guy even if the latter appears to eat more, control more territory and fight harder. Pretty boys may be less fit on the colloquial sense, but they are not less fit technically if they can get more of their genes into the next generation. So, strictly speaking appealing to a female’s aesthetics (if there is such a thing) in such a way as to make you more alluring to her and making it more likely that your genes will mix with hers makes you more fit even if you are slower, weaker and more pusillanimous (i.e. less fit in common parlance).  

Putting the aside aside, focusing on the less fluffy virtues may seem compelling when it comes to animals, though even here the story gets a bit involved and slightly incredulous.  So for example here’s one story: peahens prefer peacocks with big tails because if a peacock can make it in peacock world despite schlepping around a whopping big tail that makes doing anything at all a real (ahem) challenge, then that peacock must be really really really fit (i.e. stronger, tougher, etc.) and so any rational peahen would want its genes for its own offspring. The evaluation is purely utilitarian and the preference for the badly engineered results (big clumsly tail) are actually the hidden manifestations of a truer utilitarian calculus (really really fit because even with handicap it succeeds).

And what would the alternative be? Well, here’s a simple possibility: peahens find big tails hot and are attracted to showy males because they prefer hotties with big tails. There is nothing underneath the aesthetic judgment. It is not beautiful because of some implied utility. It’s simple lust for beauty driving the train. Beauty despite the engineering-wise grotesque baggage. Of course, believing that there is something like beauty that is not reducible to having a certain (biological) price is a belief that can land you in the poorly paid Arts faculty and exiled from the hard headed Science side of campus. Thus, they are unlikely to be happily entertained. However, it is worth noting how little there often is behind the hard headed view beside the (supposed) “self evident” fact that it is hard headed.  Nonetheless, that appears to be the debate being played out in the bio world as reported by the NYT, and, as regards animals, maybe ascribing aesthetics to them is indulgent anthropomorphism.

Why do I mention this? Because what is going on here is similar to what goes on in evo accounts concerning humans as well. The point is discussed in a terrific Jerry Fodor review of Pinker’s How the Mind Works in the LRB about 20 years ago (here). If you’ve never read it, go there now and delight yourself. It is Fodor at his acerbic (and analytical) best.

At any rate, he makes the following point in discussing Pinker’s attempt to “explain” human preferences for fiction, friends, games, etc. in more prudential (adaptationist/ selectionist) terms.

I suppose it could turn out that one’s interest in having friends, or in reading fictions, or in Wagner’s operas, is really at heart prudential. But the claim affronts a robust, and I should think salubrious, intuition that there are lots and lots of things that we care about simply for themselves. Reductionism about this plurality of goals, when not Philistine or cheaply cynical, often sounds simply funny. Thus the joke about the lawyer who is offered sex by a beautiful girl. ‘Well, I guess so,’ he replies, ‘but what’s in it for me?’ Does wanting to have a beautiful woman – or, for that matter, a good read – really require a further motive to explain it? Pinker duly supplies the explanation that you wouldn’t have thought that you needed. ‘Both sexes want a spouse who has developed normally and is free of infection … We haven’t evolved stethoscopes or tongue-depressors, but an eye for beauty does some of the same things … Luxuriant hair is always pleasing, possibly because … long hair implies a long history of good health.’

Read the piece: the discussion of why we love literature and want friends is quite funny. But the serious point is that aside from being delightfully obtuse, the more hard headed “Darwinian” account ends up sounding unbelievably silly. Just so stories indeed! But that’s what you get when in the end you demand that all values reduce to their cash equivalent.

So, the debate rages.

The second piece is on birdsongs in species that don’t bring up their own kids (here). Cowbirds are brood parasites. They are also songbirds. And they are songbirds that learn the cowbird song and not that of their “adoptive” hosts. The question is how do they manage to learn their own song and not that of their hosts (i.e. ignore the song of their hosts and zero in on that of their conspecifics? The answer seems to be the following:

…a young parasite recognizes conspecifics when it encounters a particular species-specific signal or "password" -- a vocalization, behavior, or some other characteristic -- that triggers detailed learning of the password-giver’s unique traits.

So, there is a certain vocal signal (a “password” (PW)) that the young cowbird waits for and this allows it to identify its conspecific and this triggers the song learning that allows the non cowbird raised bird to learn the cowbird song. In other words, it looks a very specific call (the “chatter call”) triggers the song learning part of the brain when it is heard. As the article puts it:

Professor Hauber's "password hypothesis" proposes that young brood parasites first recognize a particular signal, which acts as a password that identifies conspecifics, and the parasites learn other species­-specific characters only after encountering that password. One of the important features of the password hypothesis is that the password must be innate and familiar to the animal from a very early age. This suggests that encountering the password triggers specific neural responses early in development -- neural responses can actually be seen and measured.

It seems that some of the biochemistry behind this PW triggering process has been identified.

…cowbirds' brains change after the birds hear the chatter call by rapidly increasing production of a protein known as ZENK. This protein is ephemeral; it is produced in neurons after exposure to a new stimuli, but disappears only a few hours later, and it is not produced again if the same stimuli is encountered. The production of ZENK occurs in the neurons in the auditory forebrain, which are regions in the songbird brain that respond to learned vocalizations, such as songs, and also to specific unlearned calls.

So, hear PW, get ZENKed, get song. It gets you to think: why don’t humans have the same thing wrt language? Why aren’t there PWs for English, Chinese etc? Or more exactly, why isn’t it the case that humans come biologically differentiated so that they are triggered to learn different languages? Or, why is it that any child can acquire any language in the same way as any other child?  You might think that if language evolved with a P&P architecture and that different Gs were simply different settings of the same parameters with different values (i.e. each G was a different vector of such values) that evolution might have found it useful to give those most likely to grow up speaking Piraha or Hungarian a leg up by prepopulating their parameter space with Piraha or Hungarian values. Or at least endowing these offspring with PWs that when encountered triggered the relevant G values. Why don’t we see this?

Here’s one non-starter of an answer: there’s not enough time for this to have happened. Wrong! If we can have gone from lactose intolerant to lactose tolerant in 5,000 years then why couldn’t evo give some kids PWs in that time? Maybe too much intermixing of populations? But we know that there have been long stretches of time during which populations were quite isolated (right?). So this could have happened, and indeed it did with cowbirds. So why not with us? [1]

At any rate, it is not hard to imagine what the cowbird linguistic equivalent would be. Hear a sentence like “Mr Phelps, should you agree to take this mission then, as you know, should you or any of your team be captured, the government will disavow any knowledge of your activities” and poof, out pops English. Just think of how much easier second language acquisition would be. Just a matter of finding the right PWs.  But this is not, it seems, how it works with us. We are not cowbirds. Why not?

So, enjoy the pieces, they amused me. Hopefully they will amuse you too.


[1] This would particularly apposite given Mark Baker’s speculations (here; p. 23):

..it could be that linguistic diversity has the desirable function of making it hard for a greedy or dangerous outsider to join your group and get access to your resources and skills. You are less vulnerable to manipulation or deception by a would-be exploiter who cannot communicate with you easily.

In this context, a PW for offspring might be just what Dr Darwing might have ordered. But it appears not to exist.

Friday, June 16, 2017

Vapid and vacuous

It’s hard to be both vapid and vacuous (V&V), but some papers succeed. Here is an example. It is, of course, a paper on the evolution of language (evolang) and it is, of course, critical of the Chomsky-Berwick (and many others) approach to the problem. But the latter is not what makes it V&V. No, the combination of banality and emptiness starts from the main failing of many (most? all?) these evolang papers. It fails to specify the capacity the evolution of which it aims to explain. And this necessarily leads to a bad end. Fail to specify the question and nothing you say can be an answer. Or, if you have no idea what properties of what capacity you aim to explain, it should be no surprise that you fail to add anything of cognitive (vs phatic) content to the ongoing conversation.

This point is not a new one, even for me (see, for example, here). Nor should it be a controversial one. Nor, to repeat, does it require that you endorse Chomsky’s claims. It simply observes the bare minimum required to offer an evo account of anything. If you want to explain how X evolved then you need to specify X. And if X is “complex” then you need to specify each property whose evolution you are interested in. For example, if you are interested in the evolution of language, and by this I mean the capacity for language in humans, then you need to specify some properties of the capacity. And a good place to start is  to look at what linguists have been doing for about 60 years.

Why? Because we know a non trivial thing or two about human natural language. We know many things about the Gs (rules) that humans can acquire and something about the properties required to acquire such Gs (UG). We have discovered a large number of non-trivial “laws” of grammar. And given this, we can ask how a system with these laws, generating these Gs (might have) evolved. So, we can ask, as Chomsky does, how a capacity to acquire recursive Gs of the kind characteristic of natural language Gs (might have) evolved. Or we can ask how a G with these properties hooked up to articulation systems (which we can also describe in some detail) might have evolved. Or we can ask how the categorization system we find in natural language Gs (might have) evolved. We can ask these question in a non trivial, non vacuous non vapid way because we can specify (some of) the properties whose evolution we are interested in. We might not give satisfactory answers mind you. By and large the answers are less interesting than the questions right now. But we can at least frame a question. Absent a specification of the capacity of interest there is no question, only the appearance of one.

Given this, the first thing one does in reading an evolang paper is to looks for a specification of the capacity of interest. Note saying that one is interested in explaining the evolution of “language” without further specification of what “language” is and what capacities are implicated is not to give a specification. Unfortunately this is what generally happens in the evolang world. As evidence, witness the recent paper by Michael Corballis linked to above. 

It fails to specify a single property of language (more exactly the capacity for language for it is this, not language, whose evolution everyone is interested in) yet spends four pages talking about how it must have evolved gradually. What’s the it that has so evolved? Who knows! The paper is mum. We are told that whatever it is is communicatively efficacious (without saying what this means or might mean). We are told that language structure is a reflection of thought and not something with its own distinctive properties but we are not given a single example of what this might mean in concrete terms. We are told that “language derives” from mental properties like the “generative capacities to travel mentally in space and time and into the minds of others” without having a specification of the either the relevant generative procedures of these two purported cognitive faculties nor a discussion of how linguistic structures, whose properties we know a fair bit about, are simple reflections of these more general capacities. In other words, we are given nothing at all but windy assertions with nary a dollop of content. 

Let me fess up: I for one would love to see how theory of mind generates the structure of polar questions or island effects or structure dependency or c-command or anything at all of linguistic specificity. Ditto for the capacity for mental time travel. Actually, I’d love to see a specification of what these two capacities consists in. We know that people can think counterfactually (which is what this seems to amount to more or less) but we have no idea how this is done. It is a mystery how it is that people entertain counterfactual thoughts (i.e. what cognitive powers undergird this capacity) though it cannot be doubted that humans (and maybe other animals) do this. Of course unless we can specify what this capacity consists in (at least in part) we cannot ask if linguistic properties are simple reflections of these. So, virtually all of the claims to the effect that theory of mind (not much of a theory by the way as we have no idea how people travel into other minds either!) and time travel suffice to get us linguistic structures is empty verbiage. Let me repeat this: the claims are not false, they are EMPTY, VACUOUS, CONTENTLESS.   

And sadly, this is quite characteristic of the genre. Say what you will about Chomsky’s proposal it does have the virtue of specifying the capacity of interest. What he is interested in is how the generative capacity that give rise to certain kinds of structured arose and argues that given its formal properties it could not have arisen gradually. Recursion is an all or nothing property. You either got it or you don’t. So whenever it arose it did not do so in small steps, first 2-item structures, then 3, then 4, then unboundedly many. That’s not sensible, as I’ve mentioned more than a few times before (see, e.g. here and here). So Chomsky may be wrong about many things, but at least he can be wrong for he has a hypothesis which starts with a specified capacity. This is a very rare thing in the evolang world, it appears.

Actually, it’s worse than this. So rare is it that journals do not realize that absent such specifications papers purportedly dealing with the topic are empty. The Corballis paper appears in TiCS. Do the editors know that it is contentless? I doubt it. They think there is a raging “debate” and they want to be the venue where those interested in the “debate” go to be titillated (and maybe informed). But these is no debate because at least the majority of the discussants don’t say anything. The most that one can say of many contributions (the Corballis paper being one) is that they strongly express the opinion that Chomsky is wrong. That there is nothing behind this opinion, that it is merely phatic expression, is not something the editors have likely noticed.

The Corballis paper is worth looking at as an object lesson. For those that want more handholding through the vices, there is also a joint reply (here) by a gang of seven (overkill IMO) showing how there is no there there, and pointing out that, in addition, the paper seems unaware of much of modern evolutionary biology.  I cannot comment on the last point competently.[1] I can say that the reply is right in noting the Corballis paper “leave[s] the problem [regarding evolang, NH] exactly where it was, adding nothing” precisely because it fails to specify “the mechanisms of recursvie thought” in time travel or theory of mind and “how might lead to the feat that has to be explained” [i.e. how language with its distinctive properties might have arisen NH].

So can a paper be both vapid and vacuous? It appears that it can. For those interested in writing one, the Corballis paper provides a perfect model. If only it were an outlier!



[1] Though I can believe it. The paper cites Evans and Levinson, Tomasello and Everett as providing solid critiques of modern GG. This is sufficient evidence that the Corballis paper is not serious. As I’ve beaten all of these horses upside the head repeatedly, I will refrain from doing so again here. Suffice it to say, that approving citations of this work suffice by themselves to cast doubt on the seriousness of the paper citing it.