Comments

Friday, December 4, 2015

Two readables

Here are two things to look at to stir the blood.

The first is a report of a recent breakthrough for storing information on DNA. This is not a new discovery (see here), but it appears to be scaling up rather dramatically. The NYT reports the following:
The new research demonstrates that specific digital files can be retrieved from a potentially vast pool of data. The new storage technology would also be capable of keeping immense amounts of information safely for a millennium or longer, researchers said.
This is a big deal, if you keep in mind how often science follows technology. If DNA becomes the basis of long term memory storage in industry, how long before the idea hits and hits and hits again that humans do the same with their DNA. This would solve "read" part of the required read-write memory issue that the Gallistel conjecture requires. The scientists quoted in the piece acknowledge that at present the "write" part is the "bottleneck."

The scientists acknowledge that their current bottleneck is in the ability to write the information in DNA, but they say they expect that technology to begin to improve rapidly.
And when it does, we will have a realized proof of concept for the the Gallistel conjecture. And then bye bye neural nets. Can't wait.

The second piece is a blog post by my distinguished colleague Colin Phillips. He takes on the vexed issues of grad education in linguistics suggesting that it might be time for a re-think. Many of the heavyweights chime in with their views and the discussion is very informative, if also a tad defensive in parts (I won't name names, you can see for yourself). One thing that Colin does not mention, but I think is germane (and probably impolitic to mention) is that there is an understory to the debate. Like it or not, I believe that linguistics is now at an intellectual cross-roads. How so?

Part of the field wants it to stay pretty much as it has always been, and by 'always' I mean for the last several hundred years. For these the subject matter of linguistics is language and the methods basically philological. Sure GG is an important advance, but mainly because the methods developed are philology on steroids. At any rate, for this group, the mantra is "linguists study language."

The other group, of which Colin is an important leader, takes the subject matter of linguistics to be a mental faculty, either FL or G in both the narrow and wider sense (to borrow two Fitch Hauser and Chomsky terms). One studies language to study these objects that are contributing causes to linguistic "behavior." On this view, philological methods are useful but not exclusively so. This group believes that other methods both can and are contributing to understanding Gs and FL as cognitive objects.

These divergent views will take differing views on a proper linguistics education. Both will respect the standard ling methods as both agree that those methods have and continue to tell us a lot about language and G/FL. The latter however believe that other methods are no less central to understanding the cognitive (and ultimately, biological) object and so these methods should have a place in a good grad curriculum.

Now the problem, there are only so many hours in a grad education 5 years and so what to cover, given that not everything can be covered. And this is a problem where departments will differ influenced by what their view of the subject matter is. The fact, however, is that something will have to give and that allowing people to develop skills beyond the philological will require that the depths of their philological education will have to be sacrificed.  This, IMO, is a sign of progress. Grad school is there to allow neophytes to become professionals. Professionalism requires specialization. However, one cannot specialize in everything. So, choices must be made. What's the best way to make them? Well here people can legitimately disagree. However, if you are a cog-bio-linguist then you will believe that training in non philological techniques will be as legit as the other methods. This does not mean ignoring the standard methods, but it does mean making room for the others and understanding their importance for certain kinds of investigations.

Linguistics, I believe, will soon move the way of other developed sciences. Papers will be written by gangs of researchers as the expertise required will transcend the expertise of any one person. This will mean learning to work with people who know stuff you don't and that you can talk with. This does not entail becoming the other, but it does mean learning to understand the other's techniques and modes of thinking.

At any rate, take a look at the discussion. It is nothing if not amusing watching the heavyweights tussle.













Thursday, December 3, 2015

"Fuck nuance"

Tim Hunter sent me this interesting paper by sociologist Kieran Healy that argues against nuance as a theoretical virtue. The title, “Fuck Nuance” tersely provides the paper’s main conclusion. It is a paper that linguists might find interesting for, IMO, many of the vices that nuance has wrt sociological theory carry over pretty directly to theory within linguistics. Indeed, as Healy argues, nuance is often the scourge of theory. Healy describes it as follows:

It is the act of making—or the call to make—some bit of theory “richer” or “more
sophisticated” by adding complexity to it, usually by way of some additional dimension, level, or aspect, but in the absence of any strong means of disciplining or specifying the relationship between the new elements and the existing ones. Theorists do this to themselves and demand it of others. It is typically a holding maneuver. It is what you do when faced with a question that you do not yet have a compelling or interesting answer to. Thinking up compelling or interesting ideas is quite difficult, and so often it is easier to embrace complexity than cut through it. (2)

In other words, nuance is the name for the impulse to do anything and everything to cover that last data point. It is “the free-floating request that something more be added” and this sort of request, if indulged, becomes “a pernicious and invasive weed” that strangles the hope of explanation. As Healy writes:

…the kudzu of nuance … makes us shy away from the riskier aspects of
abstraction and theory-building generally, especially if it is the first and most frequent response we hear. Instead of pushing some abstraction or argument along for a while to see where it goes, there is a tendency to start hedging theory with particulars. People complain that you’re leaving some level or dimension out, and tell you to bring it back in. Crucially, “accounting for”, “addressing”, or “dealing” with the missing item is an unconstrained process. That is, the question is not how a theory can handle this or that issue internally, but rather the suggestion to expand it with this new term or terms. (5)

I am very sympathetic to Healy’s observations. Indeed, I think that I am on record lamenting this tendency within current linguistic theory (see here). Rare is the paper or interaction where theoretical profligacy is resisted and a data point or two left stranded in the deductive wilderness. Indeed, tolerance for “counter-examples”[1] is taken to be sure evidence of scientific felony, and to forestall such a charge we lard our papers and stories with ad hocry that serves to both obscure the interesting explanatory points being made and to confuse us into mistaking description for explanation. One of the virtues of the early Minimalist papers was to warn against this tendency, alas to little apparent effect. What makes this truly unfortunate is that these demands, as Healy notes, have baleful results.

The result is a lot of unproductive blocking. Both specific explanations and more
abstract concepts and theories suffer. By calling for a theory to be more comprehensive, or for an explanation to include additional dimensions, or a concept to become more flexible and multi-faceted, we paradoxically end up with less clarity. We lose information by adding detail. A further odd consequence is that the apparent scope of theories increases even as the range of their actually-accomplished application in explanations narrows. (6)

The paper makes other interesting points and I encourage you to read it. And please don’t think that this is just dumb sociologists and that his observations do not apply to linguistics. Here’s one more observation that I believe hits the mark concerning the motivations behind looking for nuance:

There is a strong tendency to embrace the fine-grain, both as a means of
defense against criticism and as a moral guarantor of the value of everyone’s empirical research project. (4)

So, I am with Healy here. Fuck theoretical nuance!



[1] These are scare quotes. I am currently hard riding a hobby horse by the name of “most people don’t really understand what a serious counter-example is.” Maybe I will write on this sometime soon, but right now I am loving the verbal ride.

Tuesday, December 1, 2015

More (and more polite) critical comments on Evans book

Asya Pereltsvaig dedicates (as of today) three posts examining Vyvyan Evan's claims in his recent book. I have already made my comments on them (many times e.g. here). For those still interested in the topic, Asaya's posts are thorough and helpful. Plus, she is not nearly as polemical as I was, so if that offended you, her posts will be more congenial to your sensibilities (and reach effectively the same conclusions that mind did).  For the interested here are three links (here, here, here)

Monday, November 30, 2015

In case you missed it; a remark on Gallistel's conjecture.

Patrick Trettenbrein posted a comment on the recent little post on Gallistel's conjecture (here). He included a little paper of his that reviews some interesting new work on aplysia that provides further evidence for the Gallsitel conjecture. Here are the concluding paragraphs.

All in all, it seems that there indeed are two different processes at work in learning and memory, as Chen et al. (2014) also point out. While the exact details about both remain obscure, there appears to be a dissociation between the way in which learning occurs and how memory works. We do not know how the brain implements a read/write memory, but there is good evidence that it does. Similarly, there is ample and convincing evidence, also in Chen et al. (2014), that synaptic conductivity and connectivity play a role in regulating behavior. Consequently, it appears that synaptic plasticity might not so much be a precondition for learning as it is a consequence of it, so that the observed rewiring of synaptic connections might constitute the brain's way of ensuring an “efficient,” or possibly even close to “optimal” (Cherniak et al., 2004; Sporns, 2012), connectivity and therefrom resulting activity pattern that is appropriate to environmental (and presumably also “internal”) conditions. Synaptic plasticity thus might be reinterpreted as a way of regulating behavior (i.e., activity and connectivity patterns) only after learning has already occurred (i.e., after relevant information has been extracted from the environment and stored in memory).

Extrapolating Chen et al.'s (2014) findings stemming from work on Aplysia to claims about much more complex nervous systems is, of course, speculative in nature, to say the least. However, it seems to be no more speculative than the almost universally accepted idea of the synapse being the locus of memory. Similarly to Johansson et al. (2014), the work of Chen et al. (2014) shows that (1) there is plenty of “room” for the implementation of symbols other than synapses, and (2) substantiates the understanding that the network approach of connectionism might indeed best be seen as an implementational theory (Fodor and Pylyshyn, 1988) that still requires representation, computation, and a Turing architecture (i.e., a read/write memory). Gallistel and Balsam (2014) proclaimed that is was about time to rethink the neural mechanisms of learning and memory, Chen et al.'s experimental results add to the urgency of this claim.

 I particularly like the speculation that the wiring is there not to code the relevant information but for efficient use and that it is a consequence of learning rather than a pre-condition for it. I also like the observation, that Gallistel and Matzel also emphasize (see here), that there is, at best, paltry evidence for the standard assumption that the "synapse [is] the locus of memory." The Gallistel conjecture is generally assumed to be some daring edge of thought kind of speculation for which there is little evidence in contrast to the well-established "fact" that memory lives in inter-neural connections. Vast academic enterprises are based on this assumption. This may be more truthy than true, however.

At any rate, take a look at Patrick's note and the Chen paper he links to. It seems that the Gallsitel conjecture is daily becoming less "exciting." About time.

A typological trap

Are the critics right? Is there scant evidence for universals? The right answer is that it all depends what you mean by ‘universal.’ If by this you intend a Greenberg Universal (GU) then it might be right (in fact, as you will see below, I think it should be right). If by this you mean a Chomsky Universal (CU) then you are not likely right. There is a big difference between these two and the empirical success of GG rests on keeping them firmly distinguished. Why? Because a priori there is little reason to think that there are many GUs out there. There may be a few, but the standard GG universals if understood as candidate GUs are not likely among them. When critics of GG argue that it’s hard to find universals expressed in the world’s many languages, they understand ‘universals’ as GUs. And here they may well be right!  However, as GG commits itself to CUs and not GUs it requires quite some ancillary argument to conclude from the absence of GUs to the non-existence of CUs.

Indeed, from a perfectly normal “scientific” point of view, we should not expect to find many GUs. What do I mean? Well, think of the analogues of GUs in a real science, like Newtonian mechanics. We observe that bodies fall. We ask what makes bodies fall. We propose that bodies “fall” because of gravitational attraction. In particular, there is a force G that causes bodies (i.e. masses) to attract. Larger masses exert a stronger pull than smaller ones. Thus, bodies much smaller than the earth, say a ball, will “fall” in the sense that the mass of the earth will more strongly attract the mass of the ball than vice versa. This will make it appear that the ball “falls” to earth when it is released (rather than the earth “falls” to the ball).  That’s the story, and a good one it is, though we now know that it needs amending, especially if the ball is travelling near the speed of light. At any rate, what’s this have to do with GUs?

Well, is it indeed phenomenologically accurate that when we observe falling bodies in the wild we observe them acting in accord with Newton’s law of gravitation? Nope. Note even close. A leaf drops from a tree.  Does it appear to fall in accordance with the law of falling bodies. Not on your life.  Drop a ball into lake and see how long it takes to hit bottom (or if it hits bottom at all). Does it appear to drop in accordance with the law of falling bodies? Nahh! Or take a body that is electrically charged and drop it in an electrical field and see if Newton’s law suffices to describe its trajectory. It doesn’t. What’s the right conclusion: that gravity is not a cause of falling bodies, that things don’t universally attract (i.e. fall)? Not on your life. Why?

Here’s the conventional wisdom. We understand that the law of falling bodies is not intended as a description of what we see outside our window. It describes one relevant force in causing what we see. And this force in complex interaction with many other factors, causes observed physical behavior. Thus, we know that shape matters (not just mass) if the object is not dropped in a vacuum (and vacuums are pretty rare out there in the real world). We know that the consistency of the space into which an object drops also matters (less frictional resistance in air than in water  and less in air than in mercury). We know that electrical charges exert forces on electrically charged objects and so this, as well as mass, can effect a falling object’s trajectory. If the object is small enough, then other factors may intercede as well. If shaped a certain way, drop an object into water and it will float rather than fall. So many factors stand between dropping and falling and nonetheless gravity does explain why bodies fall.

What’s this mean? It means that whatever the law of falling bodies is, it is not a description of what we see when we look at the world outside our window. In other words, it is not GUish.[1] It is not a description of the immediately phenomenal world, but a proposal about one of the fundamental forces that act on bodies and this force is (often) an important causal factor in determining how bodies that we see fall actually do fall. Pure expressions of the law take careful experimental set up. Indeed, we must control for many other factors, before we can “see” gravity’s effects.[2]

There is an excellent discussion of just how complex this is in Cartwright (here, chapter 4). I have discussed her main points in previous posts (see here). For current purposes, Cartwright makes two very important observations. First, that it takes a lot of work to hook a law up to observation. This is what a good experiment does. It establishes a way of observing the effects of abstract non-observable features to visible effects. It creates a “nomological machine,” a way of hooking up the underlying capacities to surface regularities. And, this is the important part:

There is no fact of the matter what a system can do just in virtue of having a given capacity. What it does , depends on its setting, and the kinds of settingsnecessary for it to produce systematic and predictable results are very exceptional (73).

So, gravity can be seen in action, but only if we arrange things very carefully! And if this is true of gravity, why should it be less true of a principle of UG?  Of course it might be different in the mental sciences, but it might not be, and assuming that linguistic “laws” (aka principles of UG) must be apparent to inspection in the wild is little more than methodological dualism (a real no-no).

Returning to the main topic, GUs are typological generalizations. They describe (and are intended to describe) generalizations thought to be observable across languages, surface generalizations. Why are we surprised that not many can be found? Why are we surprised that the UG principles proposed are not “surface true”? Why should we expect the visible surface properties of language to express the underlying grammatical forces at work any more than we expect the phenomenological observables of real world events to distinctly manifest their underlying causes (e.g. the law of gravity in bodies observed falling around us). We don’t in the latter, and shouldn’t in the former. Which brings us to CUs.

Chomsky understood universals to be properties of FL, FL being the specifically linguistic contribution that minds exploit to build language particular Gs. From the get-go, these were understood to be quite abstract, and to not be inducible from the simple inspection of the surface properties of sentences. Thus, CUs were not intended to be surface true, anymore than gravity is. Thus, the absence of GUs does not imply the non-existence of CUs any more than the phenomenological inadequacy of the laws of gravity to describe what happens when any object falls any time anywhere invalidates Newton’s theory of gravity and its explanation for the law of falling bodies.

IMO, none of this is or should be controversial. I mention it because it seems easily forgotten.  Linguists (or many of them) are currently quite skeptical that we have discovered any universals. But this is because many forget the distinction between GUs and CUs. Doing so leads to skepticism precisely because there is every reason to believe that universals understood as Greenbergian objects are not (and should not be) thick on the linguistic ground. Thus, when critics point out that such GUs are not pervasive we should agree and say that nobody thought (or should have thought) they would be. And then loudly repeat that GUs are not CUs and CUs is what we are looking for.

Why the warning? Because, it seems to me that typological work invites the inference that linguists are on the hunt for GUs and that GGers agree with critics of the Chomsky program that universals ought to be understood as GUs. But this is a mistake, one that misunderstands what GG is about. To repeat a venerable theme: GG takes the object of study to be the structure of FL/UG, not the properties of languages. These latter are interesting to the degree that they illuminate the former. And there is no reason to think that linguistic principles, any more than any other scientific principles, will be visible in the data used to investigate them.

Let me make this point another way. IMO, there is no way that something like FL/UG does not exist (see here and here for a defense). That FL/UG exists is a virtual truism. What’s in FL/UG is not. Thus, what’s up for grabs is the fine structure of FL/UG, not whether it exists. Here’s another triviality: language exhibits the properties of FL/UG only in interaction with many other adventitious linguistic factors, many non-linguistic cognitive factors and probably much else (like the weather, time of day, and who knows what else). This means that we expect the fine structure of FL/UG to be hard to discern and we do not expect it to sit out there waiting to be spotted by (even careful) observation. 

In fact, I would go further (as you knew I would). I suspect that the only really good way to argue for a CU is via something like a POS argument.  Looking at lots of languages and Gs might be helpful (see here), but if you want to zero in on potential candidate universals, there is nothing like a POS argument. Why? Because, POSs limn the borders of the grammatically possible. That’s what’s so nice about them. Inductive surveys of many Gs cannot do this. POSs are the linguistic analogue of Cartwright’s nomological machines. They afford the most direct access to CUs, and for those interested in FL/UG, CUs are the principle objects of interest.

So be careful out there. Languages and their fabulous intricacies can be confusing. It’s not that hard to mistake Greenberg Universals for Chomsky Universals, and it’s a slippery slope from there the dreaded vice of Empiricism (and its concomitant horrors (e.g. connectionism). So watch your step when you go into the field.


[1] As I’ve noted before, there is a tendency to understand universals as patterns in the data waiting to be revealed. Finding universals is then roughly a problem in signal processing in which the judicious use of statistical techniques will find the signal in the often very noisy noise. This conception understands universals as GUs. It is not the right model of a CU. For discussion see here. Incidentally, mistaking GUs for CUs will eventually lay low Deep Learning/Big Data approaches to language. The latter count on the fact that all universals will be GUs. If this is false, and it is, then such approaches cannot succeed, and so they won’t. Of course it will take time for this to become evident and by then another fad will sweep the Empiricist world.
[2] There is an excellent discussion of just how complex this is in Cartwright (here, chapter 4). I have discussed her main points in previous posts (see here).

Wednesday, November 25, 2015

Ok, tell me that this shouldn't be part of every Ling PhD defense?

Talk about outreach! Here is a screening of this year's social science winner of the "Dance your PhD competition." I believe that our talented Grads could do a whole lot better.

Tuesday, November 24, 2015

Minsky on Gallistel

I once heard of a class tight in the great days of literary theory entitled something like "The influence of Philip Roth on Charles Dickens."  My memory tingles the suggestion that I have the names wrong here, but I am pretty sure that I got the gist right. A linguistic version of this might be "The influence of Chomsky on von Humboldt." The idea is that we see the past more clearly, when we see the present concepts more clearly. The inimitable intellectual archivist Bob Berwick sent me this great quote from Marvin Minsky:

“Unfortunately, there is still very little definite knowledge about, and not even any generally accepted theory of, how information is stored in nervous systems, i.e., how they learn. … One form of theory would propose that short-term memory is ‘dynamic’—stored in the form of pulses reverberating around closed chains of neurons. … Recently, there have been a number of publications proposing that memory is stored, like genetic information, in the form of nucleic-acid chains, but I have not seen any of these theories worked out to include plausible read-in and read-out mechanisms. (Minsky 1967, 66). Minsky, Finite and Infinite Machines.
So, it seems that Randy's conjecture has a distinguished pedigree and we cog-neuro has investigated the theory of genetic information storage largely by ignoring it. Let's hope that this time around this alternative hypothesis, one which really would challenge long held views in cog-neuro, is carefully vetted. Conceptually, the Gallistel view seems to me very strong. This does not mean that it is right, but it does mean that a perfectly reasonable alternative view has not even been pursued.