Comments

Showing posts with label evolution. Show all posts
Showing posts with label evolution. Show all posts

Thursday, February 23, 2017

Optimal Design

In a recent book (here), Chomsky wants to run an argument to explain why the Merge, the Basic Operation, is so simple. Note the ‘explain’ here. And note how ambitious the aim. It goes beyond explaining the “Basic Property” of language (i.e. that natural language Gs (NLG) generate an unbounded number of hierarchically structured objects that are both articulable and meaningful) by postulating the existence of an operation like Merge. It goes beyond explaining why NLGs contain both structure building and displacement operations and why displacement is necessarily to c-commanding positions and why reconstruction is an option and why rules are structure dependent. These latter properties are explained by postulating that NLGs must contain a Merge operation and arguing that the simplest possible Merge operation will necessarily have these properties. Thus, the best Merge operation will have a bunch of very nice properties.

This latter argument is interesting enough. But in the book Chomsky goes further and aims to explain “[w]hy language should be optimally designed…” (25). Or to put this in Merge terms, why should the simplest possible Merge operation be the one that we find in NLGs? And the answer Chomsky is looking for is metaphysical, not epistemological.

What’s the difference? It’s roughly this: even granted that Chomsky’s version of Merge is the simplest and granted that on methodological grounds simple explanations trump more complex ones, the question remains, given all of this why should the conceptually simplest operation be the one that we in fact have.  Why should methodological superiority imply truth in this case?  That’s the question Chomsky is asking and, IMO, it is a real doozy and so worth considering in some detail.

Before starting, a word about the epistemological argument. We all agree that simpler accounts trump more complex ones. Thus if some account A is involves fewer assumptions than some alternative account A’ then if both are equal in their empirical coverage (btw, none of these ‘if’s ever hold in practice, but were they to hold then…) then we all agree that A is to be preferred to A’. Why? Well because in an obvious sense there is more independent evidence in favor of A then there is for A’ and we all prefer theories whose premises have the best empirical support. To get a feel for why this is so let’s analogize hypotheses to stools. Say A is a three legged and A’  a four legged stool. Say that evidence is weight that these stools support. Given a constant weight each leg on the A stool supports more weight than each of the A’ stool, about 8% more.  So each of A’s assumption are better empirically supported than each of those made by A’. Given that we prefer theories whose assumptions are better supported to those that are less well supported A wins out.[1]

None of this is suspect. However, none of this implies that the simpler theory is the true one. The epistemological privilege carries metaphysical consequences only if buttressed by the assumption that empirically better supported accounts are more likely to be true and, so far as I know, there is actually no obvious story as to why this should be the case short of asking Descarte’s God to guarantee that our clear and distinct ideas carry ontological and metaphysical weight. A good and just God would not deceive us, would she?

Chomsky knows all of this and indeed often argues in the conventional scientific way from epistemological superiority to truth. So, he often argues that Merge is the simplest operation that yields unbounded hierarchy with many other nice properties and so Merge is the true Basic Operation. But this is not what Chomsky is attempting here. He wants more! Hence the argument is interesting.[2]

Ok, Chomsky’s argument. It is brief and not well fleshed out, but again it is interesting. Here it is, my emphasis throughout (25).

Why should language be optimally designed, insofar as the SMT [Strong Minimalist Thesis, NH] holds? This question leads us to consider the origins of language. The SMT hypothesis fits well with the very limited evidence we have about the emergence of language, apparently quite recently and suddenly in the evolutionary time scale…A fair guess today…is that some slight rewiring of the brain yielded Merge, naturally in its simplest form, providing the basis for unbounded and creative thought, the “great leap forward” revealed in the archeological record, and the remarkable difference separating modern humans from their predecessors and the rest of the animal kingdom. Insofar as the surmise is sustainable, we would have an answer to questions about apparent optimal design of language: that is what would be expected under the postulated circumstances, with no selectional or other pressures operating, so the emerging system should just follow laws of nature, in this case the principles of Minimal Computation – rather the way a snowflake forms.

So, the argument is that the evolutionary scenario for the emergence of FL (in particular its recent vintage and sudden emergence) implies that whatever emerged had to be “simple” and to the degree we have the evo scenario right then we have an account for why Merge has the properties it has (i.e. recency and suddenness implicate a simple change).[3] Note again, that this goes beyond any methodological arguments for Merge. It aims to derive Merge’s simple features from the nature of selection and the particulars of the evolution of language. Here Darwin’s Problem plays a very big role.

So how good is the argument? Let me unpack it a bit more (and here I will be putting words into Chomsky’s mouth, always a fraught endeavor (think lions and tamers)). The argument appears to make a four way identification: conceptual simplicity = computational simplicity = physical simplicity = biological simplicity. Let me elaborate.

The argument is that Merge in its “simplest form” is an operation that combines expressions into sets of those expressions. Thus, for any A, B: Merge (A, B) yields {A, B}. Why sets? Well the argument is that sets are the simplest kinds of complex objects there are. They are simpler than ordered pairs in that the things combined are not ordered, just combined. Also, the operation of combining things into sets does not change the expressions so combined (no tampering). So the operation is arguably as simple a combination operation that one can imagine. The assumption is that the rewiring that occurred triggered the emergence of the conceptually simplest operation. Why?

Step two: say that conceptually simple operations are also computationally simple. In particular assume that it is computationally less costly to combine expressions into simple sets than to combine them as ordered elements (e.g. ordered pairs). If so, the conceptually simpler an operation then the less computational effort required to execute it. So, simple concepts imply minimal computations and physics favors the computationally minimal. Why?

Step three: identify computational with physical simplicity. This puts some physical oomph into “least effort,” it’s what makes minimal computation minimal. Now, as it happens, there are physical theories that tie issues in information theory with physical operations (e.g. erasure of information plays a central role in explaining why Maxwell’s demon cannot compute its way to entropy reversal (see here on the Landauer Limit)).[4] The argument above seems to be assuming something similar here, something tying computational simplicity with minimizing some physical magnitude. In other words, say computationally efficient systems are also physically efficient so that minimizing computation affords physical advantage (minimizes some physical variable). The snowflake analogy plays a role here, I suspect, the idea being that just as snowflakes arrange themselves in a physically “efficient” manner, simple computations are also more physically efficient in some sense to be determined.[5] And physical simplicity has biological implications. Why?

The last step: biological complexity is a function of natural selection, thus if no selection, no complexity. So, one expects biological simplicity in the absence of selection, the simplicity being the direct reflection of simply “follow[ing] the laws of nature,” which just are the laws of minimal computation, which just reflect conceptual simplicity.

So, why is Merge simple? Because it had to be! It’s what physics delivers in biological systems in the absence of selection, informational simplicity tied to conceptual simplicity and physical efficiency. And there could be no significant selection pressure because the whole damn thing happened so recently and suddenly.

How good is this argument? Well, let’s just say that it is somewhat incomplete, even given the motivating starting points (i.e. the great leap forward).

Before some caveats, let me make a point about something I liked. The argument relies on a widely held assumption, namely that complexity is a product of selection and that this requires long stretches of time.  This suggests that if a given property is relatively simple then it was not selected for but reflects some evolutionary forces other than selection. One aim of the Minimalist Program (MP), one that I think has been reasonably well established, is that many of the fundamental features of FL and the Gs it generates are in fact products of rather simple operations and principles. If this impression is correct (and given the slippery nature of the notion “simple” it is hard to make this impression precise) then we should not be looking to selection as the evolutionary source for these operations and principles.

Furthermore, this conclusion makes independent sense. Recursion is not a multi-step process, as Dawkins among others has rightly insisted (see here for discussion) and so it is the kind of thing that plausibly arose (or could have arisen) from a single mutation. This means that properties of FL that follow from the Basic Operation will not themselves be explained as products of selection. This is an important point for, if correct, it argues that much of what passes for contemporary work on the evolution of language is misdirected. To the degree that the property is “simple” Darwinian selection mechanisms are beside the point. Of course, what features are simple is an empirical issue, one that lots of ink has been dedicated to addressing. But the more mid-level features of FL a “simple” FL explains the less reason there is for thinking that the fine structure of FL evolved via natural selection. And this goes completely against current research in the evo of language. So hooray.

Now for some caveats: First, it is not clear to me what links conceptual simplicity with computational simplicity. A question: versions of the propositional calculus based on negation and disjunction or negation and disjunction are expressively equivalent. Indeed, one can get away with just one primitive Boolean operation, the Sheffer Stroke (see here). Is this last system more computationally efficient than one with two primitive operations, negation and/or conjunction/disjunction? Is one with three (negation, disjunction and conjunction) worse?  I have no idea. The more primitives we have the shorter proofs can be. Does this save computational power? How about sets versus ordered pairs? Is having both computationally profligate? Is there reason to think that a “small rewiring” can bring forth a nand gate but not a neg gate and a conjunction gate? Is there reason to think that a small rewiring naturally begets a merge operation that forms sets but not one that would form, say, ordered pairs? I have no idea, but the step from conceptually simple to computationally more efficient does not seem to me to be straightforward.

Second, why think that the simplest biological change did not build on pre-existing wiring? So, it is not hard to imagine that non-linguistic animals have something akin to a concatenation operation. Say they do. Then one might imagine that it is just as “simple” to modify this operation to deliver unbounded hierarchy as it is to add an entirely different operation which does so. So even if a set forming operation were simpler than concatenation tout court (which I am not sure is so), it is not clear that it is biologically simpler to derive hierarchical recursion from a modified conception of concatenation given that it already obtains in the organism then it is to ignore this available operation and introduce an entirely new one (Merge). If it isn’t (and how to tell really?) then the emergence of Merge is surprising given that there might be a simpler evolutionary route to the same functional end (unbounded hierarchical objects via descent with modification (in this case modification of concatenation)).[6]

Third, the relation between complexity of computation and physical simplicity is not crystal clear for the case at hand. What physical magnitude is being minimized when computations are more efficient? There is a branch of complexity theory where real physical magnitudes (time, space) are considered, but this is not the kind of consideration that Chomsky has generally thought relevant. Thus, there is a gap that needs more than rhetorical filling: what links the computational intuitions with physical magnitudes?

Fourth, how good are the motivating assumptions provided by the great leap forward? The argument is built by assuming that Merge is what gets the great leap forward leaping. In other words, the cultural artifacts that are proxy for the time when the “slight rewiring” that afforded Merge that allowed for FL and NLGs. Thus the recent sudden dating of the great leap forward are the main evidence for dating the slight change. But why assume that the proximate cause of the leap is a rewiring relevant to Merge, rather than say, the rewiring that licenses externalization of the Mergish thoughts so that they can be communicated. 

Let me put this another way. I have no problem believing that the small rewiring can stand independent of externalization and be of biological benefit. But even if one believes this, it may be that large scale cultural artifacts are the product of not just the rewiring but the capacity to culturally “evolve” and models of cultural evolution generally have communicative language as the necessary medium for cultural evolution. So, the great leap forward might be less a proxy for Merge than it is of whatever allowed for the externalization of FL formed thoughts. If this is so, then it is not clear that the sudden emergence of cultural artifacts shows that Merge is relatively recent. It shows, rather, that whatever drove rapid cultural change is relatively recent, and this might not be Merge per se but the processes that allowed for the externalization of merge generated structures.

So how good is the whole argument? Well let’s say that I am not that convinced. However, I admire it for it tries to do something really interesting. It tries to explain why Merge is simple in a perfectly natural sense of the word.  So let me end with this.

Chomsky has made a decent case that Merge is simple in that it involves no-tampering, a very simple “conjoining” operation resulting in hierarchical sets of unbounded size and that has other nice properties (e.g. displacement, structure dependence). I think that Chomsky’s case for such a Merge operation is pretty nice (not perfect, but not at all bad). What I am far less sure of is that it is possible to take the next step fruitfully: explain why Merge has these properties and not others.  This is the aim of Chomsky’s very ambitious argument here. Does it work? I don’t see it (yet). Is it interesting? Yup! Vintage Chomsky.



[1] All of this can be given a Bayesian justification as well (which is what lies behind derivations of the subset principle in Bayes accounts) but I like my little analogy so I leave it to the sophisticates to court the stately Reverend.
[2] Before proceeding it is worth noting that Chomsky’s argument is not just a matter of axiom counting as in the simple analogy above. It involves more recondite conceptions of the “simplicity” of one’s assumptions. Thus even if the number of assumptions is the same it can still be that some assumptions are simpler than others (e.g. the assumption that a relation is linear is “simpler” than that a relation is quadratic). Making these arguments precise is not trivial. I will return to them below.
[3] So does the fact that FL has been basically stable in the species ever since it emerged (or at least since humans separated). Note, the fact that FL did not continue to evolve after the trek out of Africa also suggests that the “simple” change delivered more or less all of what we think of as FL today. So, it’s not like FLs differ wrt Binding Principles or Control theory but are similar as regards displacement and movement locality. FL comes as a bundle and this bundle is available to any kid learning any language.
[4] Let me fess up: this is WAY beyond my understanding.
[5] What do snowflakes optimize? The following see here, my emphasis [NH]):

The growth of snowflakes (or of any substance changing from a liquid to a solid state) is known as crystallization. During this process, the molecules (in this case, water molecules) align themselves to maximize attractive forces and minimize repulsive ones. As a result, the water molecules arrange themselves in predetermined spaces and in a specific arrangement. This process is much like tiling a floor in accordance with a specific pattern: once the pattern is chosen and the first tiles are placed, then all the other tiles must go in predetermined spaces in order to maintain the pattern of symmetry. Water molecules simply arrange themselves to fit the spaces and maintain symmetry; in this way, the different arms of the snowflake are formed.

[6] Shameless plug: this is what I try to do here, though strictly speaking concatenation here is not among objects in a 2-space but a 3-space (hence results in “concatenated” objects with no linear implications.

Monday, December 7, 2015

How deep are typological differences?

Linguists like to put languages into groups. Some of these, as in biology, are groupings based on historical descent (Germanic vs Romance), some of long standing (Indo-European vs Ural Altaic vs Micronesian). Some categorizations show more sensitive to morpho-syntactic form (analytic vs agglutinative) and some are tied to whether they got to where they are spoken by tough guys who rode little horses over long distances (Finno-Ugaric (and Basque?)).  There is a tacit agreement that these groupings are significant typologically and hence linguistically significant as well. In what follows, I want to query the ‘hence.’ I would like to offer a line of argument that concludes that typological differences tell us nothing about FL. Or, to put this another way, the structure of FL in no way reflects the typological differences that linguists have uncovered. Or, to put this in reverse, typological distinctions among languages have no FL import. If this is correct, then typology is not a good probe into the structure of FL. And so if your interest is the structure of FL (i.e. if liming the fine structure of FL is how you measure linguistic significance), you might be well advised to study something other than typology.

Before proceeding let me confess that I am not all that confident about the argument that follows. There are several reasons for this. First, I am unsure that the premises are as solid as I would like them to be. As you will see, it relies on some semi-evolutionary speculation (and we all know how great that is, not!). Second, even given the premises, I am unsure that the logic is airtight. However, I think that the argument form is interesting and it rests on widely held minimalist premises (based on a relatively new and, IMO, very important observation regarding the evolutionary stability of FL), so even if the argument fails it might tell us something about these premises.  So, with these caveats, cavils, hedges and CYAs out of the way, here is the argument.

Big fact 1: the stability of FL. Chomsky has emphasized this point recently. It is the observation that whatever change (genetic, epi-genetic, angelic) led to the re-wiring of the human brain thus supporting the distinctive species specific nature of human linguistic facility, whatever change that was, it has remained intact and unchanged in the species since its biological entrance.  How do we know?

We know because of Big Fact 2: any kid can learn any language and any kid learning any language does so in essentially the same way. Thus, for example, a kid from NYC raised in Papua New Guinea (PNG) will acquire the local argot just like a native (and in the same way, with the same stages, making the same kinds of mistakes etc.). And vice versa for a PNGer in NYC, despite the relative biological isolation of PNGers for a pretty long period of time. If you don’t like this pair, plug in any you would like, say Piraha speakers and German speakers or Hebrew Speakers and Japanese. A child’s biological background seems irrelevant to which Gs it can acquire and how it acquires them. Thus, since humans separated about 100kya (trek out of Africa and all that), FL has remained biologically stable in the species. It has not changed. That’s the big fact of interest.

Now, observe that 100k years is more than enough time for evolution to work its magic. Think of Darwin’s finches. As soon as a niche opened up, these little critters evolved to exploit it. And quickly filling niches is not reserved just for finches. Humans do the same thing. Think of lactase persistence (here). The capacity to usefully digest milk products arose with the spread of cattle domestication (i.e. roughly 5-10kya).[1] So, humans also evolutionarily track novel “environmental” options and change to exploit them at a relatively rapid rate. If 5-10k years is enough for the evolution of the digestive system, then 100k years should be enough for FL to “evolve” should there be something there to evolve. But, as we saw above, this seems to be false. Or, more accurately, Big Fact 2 implies Big Fact 1 and Big Fact 1 denies that FL has evolved in the last 100k years. In sum, it seems that once the change allowing FL to emerge occurred nothing else happened evolution wise to differentially affect this capacity across humans. So far as we can tell, all human FLs are the same.

We can add to this a third “observation,” or, more accurately, something I believe that linguists think is likely to be true though we probably only have anecdotal evidence for it. Let’s call this Big Fact 3 (understanding the slight tendentiousness of the “fact” part): kids can learn multiple first languages simultaneously and do so in the same way despite the languages involved. [2] So, LADs can acquire English and Hebrew (a Germanic and Semitic language) as easily as German and Swedish (two Germanic languages), or Navajo and French or Basque and Spanish as easily as French and Spanish or… In fact, kids will acquire any two languages no matter how typologically distinct in effectively the same way. In short, typological difference has no discernable impact on the course of acquisition of two first languages. So, not only is there no ethnically-biologically based genetic pre-disposition among FLs for some Gs over others, there is not even a cognitive preference for acquiring Gs of the same type over Gs that are typologically radically different.

If thess “facts” are indeed facts, the conclusion seems obvious: to the degree that we understand FL as that cognitive-neural feature of humans that underlies our capacity to acquire Gs then it is the same across all humans (in the same sense that hearts or kidneys are, (i.e. abstracting from normal variation)) and this implies that it has not evolved despite apparently sufficient time for doing so.[3]

This raises an obvious question: why not? Why does the process of language acquisition not care about typological differences? Or, if typological differences run deep then why have they had no impact on the FLs of people who have lived in distinct linguistic eco-niches?

Here’s one obvious answer: typological differences are irrelevant to FL. However big these differences may seem to linguists, FL sees these typologically “different” languages as all of a piece. In other words, from the point of view of FL, typological variation is just surface fluff.

Same thing, said differently: there is a difference between variation and typology. Variation is a fact, languages appear on the surface to have different properties. Typology is a mid-level theoretical construct. It is the supposition that variation comes in family types, or, that variation is (at least in part) grammatically principled. The argument above does not question the fact of variation. It calls into question whether this variation is in any FL sense principled, whether the mid level construct is FL significant. It argues it isn’t.

Let me put this last point more positively. Variation establishes a puzzle for linguists in that kids acquire Gs that result in different surface features. So, FL plus a learning theory must be able to accommodate variation. However, if the above is on the right track, then this is not because typological cleavages reflect structural fault lines (or G-attractors) in FL or the learning theory. How exactly FL and learning theories yield distinctive Gs is currently unknown. We have good cases studies of how experience can fix different Gs with different surface properties but I think it is fair to say that there is still lots more fundamental work to be done.[4] Nonetheless, even without knowing how this happens, the argument above suggests that it does not happen in virtue of a typologically differentiated FL.

Let me end with one last observation. Say that the above is correct, it seems to me that a likely corollary is that FL has no internal parameters. What I mean is that FL does not determine a finite space of possible Gs, as GB envisioned. Why not?

Well say that acquisition consisted in fixing the values of a finite series of FL internal open parameters. Then why wouldn’t evolution have fixed the FL of speakers of typologically isolated languages so that the relevant typological parameters were no longer open. On the assumption that “closing” such a parameter would yield an acquisition advantage (fixing parameters would reduce the size of the parameter space, so the more fixed parameters the better as this would simplify the acquisition problem), why wouldn’t evolution take advantage of the eco-niche to speed up G acquisition? Thus, why wouldn’t humans be like finches with FLs quickly specializing to their typological eco-niches? Doesn’t this suggest that parameters are not internal properties of FL?

I am pretty sure that readers will find much to disagree with here. That’s great. I think that the line of reasoning above is reasonable and hangs together. Moreover, if correct, I believe that it is important for pretty obvious reasons. But, there are sure to be counter-arguments and other ways of understanding the “facts.” Can’t wait.


[1] The example provided by Bill Idsardi. Thanks.
[2] By two “first” languages I intend to signal the fact that this is a different process from second language acquisition. Form the little I know about this process, there is no strcit upper bound on how many first languages one can simultaneously acquire, though I am willing to bet that past 3 or 4 the process gets pretty hairy.
[3] It also strongly casts doubt on the idea that FL itself is the product of an evolutionary process. If it is, the question becomes why did it stop when it did and not continue after humans separated? Why no apparent changes in the last 100k years?
[4] Charles Yang has a forthcoming book on this topic (which I heartily recommend) and Jeff Lidz has done some exemplary work showing how to think of FL and learning theory together to deliver integrated accounts of real time language acquisition. I am sure that there is other work of this kind. Feel free to mention them in comments.

Wednesday, February 25, 2015

More reading for the curious

Deep learning: Here are two papers that investigate the “psychological reality” of  some popular deep learning models. They are particularly important for those wishing to borrow insights from this literature for cognitive ends. What the papers show is that it is possible to construct stimuli that the systems systematically classify (i.e. classify with very high confidence) as objects that no human would mistake them for.  Thus, deep neural networks are easily fooled: see http://arxiv.org/abs/1412.1897

These papers do something that linguists commonly do. The papers are about negative data. Negative data describe what humans do not do (e.g. native English speakers do not accept sentences like “*who did you meat a man who saw”). If deep learning models are to be understood as psychological theories, they need to agree both on the good and the bad data (i.e. on what we accept and reject). So far, much of the discussion has been on the positive capacities of such systems. They can be trained to spot a dog in a picture. However, these papers observe that current systems spot dogs that are not there, or, more precisely, categorize some picture as a dog photo that no human would so categorize. Or as the first paper puts it:

A recent study revealed that changing an image (e.g. a lion) in a way imperceptible to humans can cause a DNN [deep neural network NH] to label; the image as something else entirely (e.g. mislabeling a lion a library). Here we show a related result: it is easy to produce images that are completely unrecognizable to humans, but the state of the art DNNs believe to be recognizable objects with 99.99% confidence (e.g. labeling with certainty that white noise state is a lion).

Why do DNNs do this? Right now, it seems that nobody knows. Need I say that these are important results for the psychological “reality” of DNNs? As every linguist knows, explaining negative data is critical in evaluating any proposal aimed at describing our mental powers.

A note on evolution:
http://www.newyorker.com/news/daily-comment/evolution-catechism. This is an interesting discussion of the obvious political import that theories of evolution have in the US. The points are mainly obvious and congenial (to me). However, there is one distinction I would have made that Gopnik does not; the difference between the fact of evolution versus the centrality of the mechanism of natural selection (NS) as the prime causal force behind evolution. The fact is completely uncontroversial. Indeed, it was considered commonplace before Darwin, though Darwin did a lot to cement the truth of this fact. What is somewhat controversial today is how large a part NS plays in explaining this fact. All agree that it plays some role, the question is how big.

In many ways the recent Evo-Devo discoveries replay discussions similar to those in the early cog revolution. The Evo-Devo stuff suggests that the range of options that NS has to pick from is quite a bit narrower than earlier believed (viz. there are very few ways of doing anything (e.g. building an eye) and these tend to be strongly conserved over evolutionary time scales. Of course, the fewer the options available, the less one looks to NS to explain the outcome. Why? Because NS relies on the idea that were you are is heavily dependent on the path you took to getting there. But if the number of paths to get anywhere is very small in number then why you got to where you are is less dependent on a long series of linked choices than on the one or two you made at the very start. So, it is not that NS plays no role. Rather the importance of NS’s role depends on how wide the range of possibilities. The question is then not either/or but how much. And these are theoretical/empirical questions.

So, Gopnik is quite right that denying the fact of evolution is a nutty thing to do, sort of like denying that the earth is round or the earth orbits the sun. However, questioning the size of NS effects to evolutionary trajectories is not. NS is a theory. Evolution is the fact. As Gopnik notes, theories evolve and change. One of the changes being currently contemplated is that NS is a less potent factor than heretofore believed. Even a Republican can believe this in scientific good faith.

A neat paper on sources of evolution making a splash:
Some scene setting from Bob Berwick:
This is an interesting evolutionary finding because it uncovered a new mechanism by which evolution can work very quickly in a very complex setting.  We don’t know much about how complex organ systems evolve – the “major transitions in evolution” – like our brains. There are so many genes involved – how is it all put together without it getting all tangled up?  But now somebody’s got their foot in the door about one of the biggest transitions of all – how placental mammals evolved pregnancy, and went from laying eggs externally to growing them inside their bodies. Turns out that again (surprise!) this involved a set of regulatory genes, plus – the real surprise –what are called ‘transposons’ – bits of genes that can leap whole genes at a single bound and insert themselves even across chromosomes.[1] (Lynch et al., Ancient Transposable Elements Transformed the Uterine Regulatory Landscape and Transcriptome during the Evolution of Mammalian Pregnancy, Cell Reports (2015.[2] Apparently, the transposons donated regulatory elements to the genes that were recruited to alter the immune system so that the mother wouldn’t reject fetal embryos as foreign (remember a fetus has all those unknown genes from daddy). Lynch et al. demonstrated that this involved thousands of genes in a carefully coordinated orchestration led in part by the transposons, enabling exceptionally rapid evolution.  As Lynch notes, nobody expected to find that evolution could work this way to evolve large complex organ systems. Seems we still have a lot to learn about the basic evolutionary machinery, more than 150 years after Darwin.
More on Minds and Bodies: Those that enjoyed Chomsky’s discussion of the mind-body problem (here) (or should I say the non-existence of the problem given Newton’s excision of body from the equation) and were (rightly) dissatisfied with my discussion (here) might enjoy a real philosophical exposition of the state of the art by John Collins (here). It engages with lots of the philo literature on these matters and is eminently readable, even for linguists. He discusses and defends a position that he calls it “methodological naturalism,” which, if understood and adopted as the standard in the cog-neuro sciences (including linguistics) will remove most of the metaphysical and epistemological underbrush that hinders fruitful collaboration between linguistics and neuro-types. So, take a look and pass it onto your friends (and enemies) in the neurosciences who ignore most of what you have to say.




[1] First discovered by Barbara McClintock in corn, in the 1940s.  Nobody believed her at first, because it violated everything people thought they new about Mendelism and genes, but she hammered away at it.  Forty years later she got a Nobel prize.

Tuesday, September 16, 2014

Some fodder for lunch conversation

The inimitable Bill Idsardi sent me two links to a recent paper on Foxp2 (here and here). The paper, a collaborative effort between Ann Graybiel's lab at MIT and researchers at the Max Planck in Leipzig studied how mice equipped with a "humanized" form of the Foxp2 gene learned to run mazes. It seems that it helps, well at least some times in some ways. The big advantage the humanized gene provides os to facilitate the transition between declarative (deliberative) and procedural (automatic) forms of storing new info. At any rate, the mice did better at some tasks than those without the humanized form of the gene. The reports go onto speculate (and I do mean speculate) about how all of this might have something to do with language. Here's the AAAS version for the non-expert: "The results suggest the human version of the FOXP2 gene may enable quick switch to repetitive learning- an ability that could have helped infants 200,000 years ago better communicate with their parents." The emphasis in the previous quote is mine. I don't know if it is possible to make a more hedged "suggestion" but I sincerely doubt it.  Even so, the report from Science does quote a skeptic who is "not sure how relevant the findings are to speech" given that the test relies on visual cues while speech relies on auditory ones.  I think that were they to ask me I would have been more skeptical still as I am not sure I see what the bridging assumptions are that take one from facilitated routinization of maze running to even word learning (the capacity that Grabile cites in the MIT piece as possibly enhanced by this version of FOXP2 (is there a difference between FOXP2 and foxp2?  I suspect that the former is the human one and the latter the non-human analogue. At any rate, …). Maybe, but it would have been nice to hear a little of how these two capacities might be related.

This might be interesting and important work. I am told that Graybiel is a big deal. Still, it is odd how little attempt there is to link this language gene to any language like effect.  I suspect that the reasons for this (aside form the fact that it's probably hard at MIT to find anyone (e.g. a linguist) who knows anything about language (and yes that was sarcastic)) is that biologists are really flummoxed by language. The Science article notes in passing as if it were obvious the following: "As a uniquely human trait, language has long baffled evolutionary biologists" (2)." Funny, when I say things like this (e.g. that language is a species specific special capacity and that evolution has little to say about it) furor immediately erupts. However, it seems to be conventional wisdom, at least for Science writers (and both they and I are right about this). At any rate, take a look. It won't take long.

Here's one more thing that you might find interesting. Aaron White sent me this link to Michael Jordan where he discusses deep learning. His discussion of supervised vs unsupervised learning is useful coming from him.  It's also short and he is also a big shot in this area so it's worth a quick look.

Thanks again to Bill and Aaron for this. Let me make it official: if you find something that you think would be of general interest, please send it along to me. One hope is that the blog can exploit the wisdom of crowds to make us all more aware with what is happening elsewhere that might be of general interest to us.