Comments

Showing posts with label Gary Marcus. Show all posts
Showing posts with label Gary Marcus. Show all posts

Monday, November 26, 2018

What's innate?

Johan Bolhuis sent me a copy of a recent comment in TiCS(Priors in animal and artificial intelligence (henceforth Priors))on the utility of rich innate priors in cognition, both in actual animals and artificially in machines. Following Pinker, Priorsframes the issue in terms of the blank slate hypothesis (BSH) (tabula rasafor you Latin lovers). It puts the issue as follows (963):

Empiricists and nativists have clashed for centuries in understanding the architecture of the mind: the former as a tabula rasa, and the latter as a system designed prior to experience…The question, summarized in the debate between the nativist Gary Marcus and the pioneer of machine learning, Yann LeCun, is the following: shall we search for a unitary general learning principle able to flexibly adapt to all conditions, including novel ones, or structure artificial minds with driving assumptions, or priors, that orient learning and improve acquisition speed by imposing limiting biases?

Marcus’ paper (here) (whose philosophical framework Priorsuses as backdrop for its more particular discussion) relates BSH to the old innateness question, which it contends revolves around “trying to reduce the amount of innate machinery in a given system” (1). I want to discuss this way of putting things, and I will be a bit critical. But before diving in, I want to say that I really enjoyed both papers and I believe that they are very useful additions to the current discussion. They both make excellent points and I agree with almost all their content. However, I think that the way they framed the relevant issue, in terms of innateness and blank slates, is misleading and concedes too much to the Empiricist (E) side of the debate. 

My point will be a simple one: the relevant question is not how much innate machinery, but what kindof innate machinery. As Chomsky andQuine observed a long time ago, everyonewho discusses learning and cognition is waste deep in a lot of innate machinery. The reason is that learning without a learning mechanismis impossible. And if one has a learning mechanism in terms of which learning occurs, then that learning mechanism is not itself learned. And if it is note learned then it is innate. Or, to put this more simply, the mechanism that allows for learning is a precondition for learning and preconditions are fixed prior to that which they precondition. Hence all features of the learning mechanism are innate in the simple sense of not themselves being learned. This is a simple logical point, and all who discuss these issues are aware of this point. So the question is not, never has been, and never could not have been is there innate structure?Rather the question is, always has been and always will be what structure is innate?

Why is putting things in this way important? Because arguing about the amountof innate structure gives Eists the argumentative edge. Ockham like considerations will always favor using less machinery rather than more all things being equal. So putting things as the Marcus paper and Priorsdoes is to say that the Eist position is methodologically preferable to the Rationalist (R) one. Putting things in terms of what kinds of innate machinery is required (to solve a given learning problem), rather than how much considerably levels the methodological playing field. If both E and R conceptions require boatloads of innate machinery to get anywhere, then the question moves from whether innate structure is needed (as the BSH slyly implicates) to what sort is needed (which is the serious empirical question).

This said, let’s zero in on some specifics. What makes an approach Eist? There are two basic ingredients. The first important ingredient is associationism (Aism). This is the glue that holds “ideas” together. However, this is not all. There is a second important ingredient: perceptualism (Pism). Pism is the idea that all mental contents are effectively reducible to perceptual contents, which are themselves effectively reducible to sensory concepts (sensationalism (Sism)). 

This pair of claims lies at the center of Eist theories of mind. And the notion of the blank slate emphasizes the second. We find this reflected in a famous Eish slogan: “there is nothing in the mind that is not first in the senses.” The Eish conception unites P/Sism with Aism to get to the conclusion that all mental contents are either primitive sensory/perceptual “ideas” or constructed out of sensory/perceptual input via association. The problems with Eism arise from both sources and revolve around two claims: the denial that mental concepts interrelate other than by association (they have no further interesting logical structure) and that all ideas are congeries of sensory perceptions. These two assumptions combine to provide a strong environmentalist approach to cognition wherein the structure of the environment largely shapes the contents of the mind through the probabilistic distributions of sensory/perceptual inputs. Rism denies bothclaims. It argues that association is not the fundamental conceptual glue that relates mental contents anddenies that all complex mental contents are combinations of sensory/perceptual inputs. To wax metaphorical, for Eists, only sensation can write on our mental blank slates and the greater the sensations the more vivid the images that appear. Rists think this is empirical bunk.

Note that this combination of cognitive assumptions has a third property. Given Eist assumptions, cognition is general purpose. If cognition is nothing but tracking the frequencies of sensory inputs then all cognition is of a piece, the only difference being the sensations/perceptions being tracked. There is no modularity or domain specificity beyond that afforded by the different sensory mechanisms, nor rules of “combination” beyond those tracking the differential exposure to some sensations over others. Thus for Eists, the domain generality of cognition is not an additional assumption. It is the consequence of Eisms two foundational premises.

Now, we actually know today that Eism will not work (actually, we knew this way back way back when). In particular, Pism/Sism was very thoroughly explored at the turn of the 20thcentury and shown to be hopeless. There were vigorous attempts to reduce our conceptual contents to sense data. And these efforts completely failed! Pism/Sism, in other words, is a hopeless position. So hopeless, in fact, that the only place it still survives is in AI and certain parts of psychology. Deep Learning (DL), it seems, is the latest incarnation of P/Sism+Aism right now. BothPriorsand Marcus elegantly debunk DLs inflated pretentions by showing both that the assumptions are biologically untenable and that they are adhered to more in the PR discussions than in the practice of the parade cases meant to illustrate successful AI learners.[1] I refer you to their useful discussions. See especially their excellent points concerning how much actual learning in humans and animals is based on very little input (i.e. from a very limited number of examples). DL requires Big Data (BD) to be even remotely plausible. And this data must be quite carefully curated (i.e. supervised) to be of use. Both papers make the obvious point that much biological learning is done from very few example cases (sparse data) and is unsupervised (hence notcurated). This makes most of what DLers have “discovered” largely irrelevant as models for biologically plausible theories of cognition. Sadly, the two papers do notcome right out and say this, though they hint at it furiously. It seems that the political power of DL is such that frankly saying that this emperor is hardly clothed will not be well rewarded.[2]Hence, though the papers make this point, it is largely done in a way that bends over backwards to emphasize the virtues of DL and not appear to be critically shrill. IMO, there is a cost to this politeness.

One last point and I stop. Priorsmakes a cute observation, at least one that I never considered. Eists of the DL and connectionist variety loveplasticity. They want flexible minds/brains because these are what the combination of Aism and P/Sism entails.Priorsmakes the nice observation that if flexibility is understood as plasticity then plasticity is something that biology only values in smalldoses. Brains cease being plastic after a shortish critical period. This Priorsnotes implies that there is a biological cost of being relentlessly open minded. You can see why I might positively reverberate to this observation.

Ok, nuff said. The two papers are very good and are shortish as well. Priorsis perfect for anyone wanting to have a non human case to illustrate Rish themes in a class on language and mind. The Marcus piece is part of a series of excellent papers he has been putting out reviewing the hype behind DL and taking it down several pegs (though, again, I wish he were less charitable). From these papers and the references they cite, it strikes me that the hype that has surrounded DL is starting to wear thin. Call me a hopeless romantic, but maybe when the overheated PR dies down and it becomes clear that the problems the latest round of Eish accounts solved were not the central problems in cognition, we can return to some serious science.  


[1]An aside: there is more than a passing similarity between the old attempts to reduce mental contents to sense data and the current fad in DL of trying to understand everything in terms of pixel distributional properties. History seems to constantly repeat; the first time as insight, the second time as a long con. Not surprisingly, the attempt to extract the notion “object” or “cat” from pixel distributions is no more successful today than were prior attempts to squeeze such notions from sense data. Ditto with algebraic structure from associations. It is really useful to appreciate how long we have known that Eism cannot be a serious basis for cognition. The failures Priorsand Marcus observe are not new ones, just the same old failures gussied up in technically spiffier garb.

[2]Some influential voices are becoming far more critical. Shalizi (here) notes that much of DL is simply a repackaging of perceptrons (“extracting features from the environment which work in that environment to make a behaviorally-relevant classificationor prediction or immediate action”) and will have roughly the same limitations that perceptrons had (viz. “This sort of perception is fast, automatic, and tuned to very, very particular features of the environment… They generalize to more data from their training environment, but not to new environments…”).  Shalizi, like Marcus andPriors, locates the problems with these systems in their lack of “abstract, compositional, combinatorial understanding we (and other animals) show in manipulating our environment, in planning, in social interaction, and in the structure of language.” 
            In other words, DL is basically the same old stuff repackaged for the credulous “smart” technopilic shopper. You cannot keep selling perceptrons, so repackage and sell it as DeepLearning (the ‘deep’ here is, no doubt, the contribution of the marketing department). The fact is that the same stuff that was problematic before is problematic still. There is no way to “abstract” out compositional and combinatorial principles and structures from devices aimed to track “particular features of the environment.” 

Wednesday, January 24, 2018

Gary Marcus on deep learning

An “unknown” commentator left links to two very interesting Gary Marcus (GM) pieces (here1 and here2) on the current state of Deep Learning (DL) research. His two pieces make the points that I tried to make in a previous post (here), but do so much more efficiently and insightfully than I did. They are MUCH better. I strongly recommend that you take a look if you are interested in the topics.

Here are, FWIW, a couple of reactions to the excellent discussion these papers provide.

Consider first here1.

1. GM observes that the main critiques of DL contend not that DL is useless or uninteresting, but (i) that it leaves out a lot if one’s research interests lie with biological cognition, and (ii) that the part that DL leaves out is precisely what theories promoting symbolic computation have always focused on. In other words, the idea that DL suffices as a framework for serious cognition is what is up for grabs not whether it is necessary. Recall, Rs are comfortable with the kinds of mechanisms DLers favor. The E mistake is to think that this is all there is. It isn’t. As GM puts it (here1:4): DL is “not a universal…solvent, but simply…one tool among many…”

I am tempted to go a bit farther (something that Lake et. al. (see here) moot as well). I suspect that if one’s goal is to understand cognitive processes then DL will play a decidedly secondary explanatory role. The hard problem is figuring out the right representational format (the kinds of generalizations it licenses and categorizations it encourages). These fixed, DL can work its magic. Without these, DL will be relatively idle. These facts can be obscured by DLers that do not seem to appreciate the kinds of Rish debts their own programs actually incur (a point that GM makes eloquently in here 2). However, as we all know a truly blank slate generalizes not at all. We all need built-ins to do anything. The only relevant question is which ones and how much, not whether. DLers (almost always of an Eish persuasion) seem to have a hard time understanding this or drawing the appropriate conclusions from this uncontentious fact.

2. GM makes clear (here1:5) in what sense DL is bad at hierarchy. The piece contrasts “feature-wise hierarchy” from systems that “can make explicit reference to the parts of larger wholes.” GM describes the former as a species of “hierarchical feature detection; you build lines out of pixels, letters out of lines, words out of letters and so forth.” DL is very good at this (GM: “the best ever”). But it cannot do the second at all well, which is the kind of hierarchy we need to describe, say, linguistic objects with constituents that are computationally active. Note, that what GM calls “hierarchical feature detection” corresponds quite well with the kind of discovery procedures earlier structuralism advocated and whose limitations Chomsky exposed over 60 years ago. As GM notes, pure DL does not handle at all well the kinds of structures GGers regularly make use of to explain the simplest linguistic facts. Moreover, DL fails for roughly the reasons that Chomsky originally laid out; it does not appreciate the particular computational challenges that constituency highlights.

3. GM has a very nice discussion of where/how exactly DLs fail. It relates to “extrapolation” (see discussion of question 9, 10ff). And why? Because DL networks “don’t have a way of incorporating prior knowledge” that involve “operations over variables.” For these kinds of “extrapolations” we need standard symbolic representations, and this is something that DL eschews (for typically anti-nativist/rationalist motives). So they fail to do what humans find trivially easy (viz. to “learn from examples the function you want and extrapolate it”). Can one build into DL systems that employ operations over variables? GM notes that they can. But in doing so they will not be pure DL devices and will have to allow for symbolic computations and the innate (i.e. given) principles and operations that DLers regularly deny is needed.

4. GM’s second paper also has makes for very useful reading. It specifically discusses the AlphaGO programs recently in the news for doing for Go what other programs did for chess (beat the human champions). GM asks whether the success of these programs support the anti R conclusions that its makers have bruited about? The short answer is ‘NO!”. The reason, as GM shows, is that there is lots of specialized pre-packaged machinery that allows these programs to succeed. In other words, they are elbow deep into very specific “innate” architectural assumptions without which the programs would not function.

Nor should this be surprising for this is precisely what one should expect. The discussion is very good and anyone interested in a good short discussion of innateness and why it is important should take a look.

5. One point struck me as particularly useful. If what GM says is right then it appears that the non nativist Es don’t really understand what their own machines are doing. If GM is right, then they don’t seem to see how to approach the E/R debate because they have no idea what the debate is about. The issue is not whether machines can cognize. The issue is what needs to be in a machine that cognizes. I have a glimmer of a suspicion that DLers (and maybe other Eish AIers) confuse two different questions: (a) Is cognition mechanizable (i.e does cognition require a kind of mentalistic vitalism )? versus (b) What goes into a cognitively capable mind: how rasa can a cognitively competent tabula be?
These are two very different questions. The first takes mentalism to be opposed to physicalism, the suggestion being that mental life requires something above and beyond the standard computational apparatus to explain how we cognize as we do. The second is a question within physicalism and asks how much “innate” (i.e. given) knowledge is required to get a computational system to cognize as we do. The E answer to the second question is that not much given structure is needed. The Rs beg to differ. However, Rs are not committed to operations and mechanisms that transcend the standard variety computational mechanisms we are all familiar with. No ghosts or special mental stuff required. If indeed DLers confuse these two questions then it explains why they consider whatever program they produce (no matter how jam packed with specialized “given” structures (of the kind that GM notes to be the case with AlphaGO)) as justifying Eism. But as this is not what the debate is about, the conclusion is a non-sequitur. AlphaGo is very Rish precisely because it is very non rasa tabularly.[1]

To end: These two pieces are very good and important. DL has been massively oversold. We need papers that keep yelling about how little cloth surrounds the emperor. If your interests are in human (or even animal) cognition then DL cannot be the whole answer. Indeed, it may not even be much or the most important part of the answer. But for now if we can get it agreed that DL requires serious supplementation to get off the ground, that will be a good result. GM’s papers are a very good at getting us to this conclusion.



[1] I should add, that there are serious mental mysteries that we don’t know how to account for conutationally. Chomsky describes these as the free use of our capacities and what Fodor discusses under the heading central systems. We have no decent handle on how we freely exercise our capacities or how the complex judgments work. These are mysteries, but these mysteries are not what the E/R debate is mostly about.

Thursday, July 13, 2017

Some recent thoughts on AI

Kleanthes sent me this link to a recent lecture by Gary Marcus (GM) on the status of current AI research. It is a somewhat jaundiced review concluding that, once again, the results have been strongly oversold. This should not be surprising. The rewards to those that deliver strong AI (“the kind of AI that would be as smart as, say a Star Trek computer” (3)) will be without limit, both tangibly (lots and lots of money) and spiritually (lots and lots of fame, immortal kinda fame). And given hyperbole never cripples its purveyors (“AI boys will be AI boys” (and yes, they are all boys)), it is no surprise that, as GM notes, we have been 20 years out from solving strong AI for the last 65 years or so. This is a bit like the many economists who predicted 15 of the last 6 recessions but worse. Why worse? Because there have been 6 recessions but there has been pitifully small progress on strong AI, at least if GM is to be believed (and I think he is). 

Why despite the hype (necessary to drain dollars from “smart” VC money) has this problem been so tough to crack? GM mentions a few reasons.

First, we really have no idea how open ended competence works. Let me put this backwards. As GM notes, AI has been successful precisely in “predefined domains” (6). In other words, where we can limit the set of objects being considered for identification or the topics up for discussion or the hypotheses to be tested we can get things to run relatively smoothly. This has been true since Winograd and his block worlds. Constrain the domain and all goes okishly. Open the domain up so that intelligence can wander across topics freely and all hell breaks loose. The problem of AI has always been scaling up, and it is still a problem. Why? Because we have no idea how intelligence manages to (i) identify relevant information for any given domain and (ii) use that information in relevant ways for that domain. In other words, how we in general figure out what counts and how we figure out how much it counts once we have figured it out is a complete and utter mystery. And I mean ‘mystery’ in the sense that Chomsky has identified (i.e. as opposed to ‘problem’).

Nor is this a problem limited to AI.  As FoL has discussed before, linguistic creativity has two sides. The part that has to do with specifying the kind of unbounded hierarchical recursion we find in human Gs has been shown to be tractable. Linguists have been able to say interesting things about the kinds of Gs we find in human natural languages and the kinds of UG principles that FL plausibly contains. One of the glories (IMO, the glory) of modern GG lies in its having turned once mysterious questions into scientific problems. We may not have solved all the problems of linguistic structure but we have managed to render them scientifically tractable.

This is in stark contrast to the other side linguistic creativity: the fact that humans are able to use their linguistic competence in so many different ways for thought and self-expression. This is what the Cartesians found so remarkable (see here for some discussion) and that we have not made an iota of progress understanding. As Chomsky put it in Language & Mind (and is still a fair summary of where we stand today):

Honesty forces us to admit that we are as far today as Descartes was three centuries ago from understanding just what enables a human to speak in a way that is innovative, free from stimulus control, and also appropriate and coherent. (12-13)[1]

All-things-considered judgments, those that we deploy effortlessly in every day conversation, elude insight. That we do this is apparent. But how we do this remains mysterious. This is the nut that strong AI needs to crack given its ambitions. To date, the record of failure speaks for itself and there is no reason to think that more modern methods will help out much.

It is precisely this roadblock that limiting the domain of interest removes. Bound the domain and the problem of open-endedness disappears.

This should sound familiar. It is the message in Fodor’s Modularity of Mind. Fodor observes that modularity makes for tractability. When we move away from modular systems, we flat on our faces precisely because we have no idea how minds identify what is relevant in any given situation and how it weights what is relevant in a given situation and how it then deploys this information appropriately. We do it all right. We just don’t know how.

The modern hype supposes that we can get around this problem with big data. GM has a few choice remarks about this. Here’s how he sees things (my emphasis):

I opened this talk with a prediction from Andrew Ng: “If a typical person can do a mental task with less than one second of thought, we can probably automate it using AI either now or in the near future.” So, here’s my version of it, which I think is more honest and definitely less pithy: If a typical person can do a mental task with less than one second of thought and we can gather an enormous amount of directly relevant data, we have a fighting chance, so long as the test data aren’t too terribly different from the training data and the domain doesn’t change too much over time. Unfortunately, for real-world problems, that’s rarely the case. (8)

So, if we massage the data so that we get that which is “directly relevant” and we test our inductive learner on data that is not “too terribly different” and we make sure that the “domain doesn’t change much” then big data will deliver “statistical approximations” (5). However, “statistics is not the same thing as knowledge” (9). Big data can give us better and better “correlations” if fed with “large amounts of [relevant!, NH] statistical data”. However, even when these correlational models work, “we don’t necessarily understand what’s underlying them” (9).[2]

And one more thing: when things work it’s because the domain is well behaved. Here’s GM on AlphaGo (my emphasis):

Lately, AlphaGo is probably the most impressive demonstration of AI. It’s the AI program that plays the board game Go, and extremely well, but it works because the rules never change, you can gather an infinite amount of data, and you just play it over and over again. It’s not open-ended. You don’t have to worry about the world changing. But when you move things into the real world, say driving a vehicle where there’s always a new situation, these techniques just don’t work as well. (7)
 
So, if the rules don’t change, you have unbounded data and time to massage it and the relevant world doesn’t change, then we can get something that approximately fits what we observe. But fitting is not explaining and the world required for even this much “success” is not the world we live in, the world in which our cognitive powers are exercised. So what does AI’s being able to do this in artificial worlds tell us about what we do in ours? Absolutely nothing.

Moreover, as GM notes, the problems of interest to human cognition have exactly the opposite profile. In Big Data scenarios we have boundless data, endless trials with huge numbers of failures (corrections). The problems we are interested in are characterized by having a small amount of data and a very small amount of error. What will Big Data techniques tell us about problems with the latter profile? The obvious answer is “not very much” and the obvious answer, to date, has proven to be quite adequate.

Again, this should sound familiar. We do not know how to model the everyday creativity that goes into common judgments that humans routinely make and that directly affects how we navigate our open-ended world. Where we cannot successfully idealize to a modular system (one that is relatively informationally encapsulated) we are at sea. And no amount of big data or stats will help.

What GM says has been said repeatedly over the last 65 years.[3] AI hype will always be with us. The problem is that it must crack a long lived mystery to get anywhere. It must crack the problem of judgment and try to “mechanize” it. Descartes doubted that we would be able to do this (indeed this was his main argument for a second substance). The problem with so much work in AI is not that it has failed to crack this problem, but that it fails to see that it is a problem at all. What GM observes is that, in this regard, nothing has really changed and I predict that we will be in more or less the same place in 20 years.

Postscript:

Since penning(?) the above I ran across a review of a book on machine intelligence by Gary Kasparov (here). The review is interesting (I have not read the book) and is a nice companion to the Marcus remarks. I particularly liked the history on Shannon’s early thoughts on chess playing computers and his distinction on how the problem could be solved:

At the dawn of the computer age, in 1950, the influential Bell Labs engineer Claude Shannon published a paper in Philosophical Magazine called “Programming a Computer for Playing Chess.” The creation of a “tolerably good” computerized chess player, he argued, was not only possible but would also have metaphysical consequences. It would force the human race “either to admit the possibility of a mechanized thinking or to further restrict [its] concept of ‘thinking.’” He went on to offer an insight that would prove essential both to the development of chess software and to the pursuit of artificial intelligence in general. A chess program, he wrote, would need to incorporate a search function able to identify possible moves and rank them according to how they influenced the course of the game. He laid out two very different approaches to programming the function. “Type A” would rely on brute force, calculating the relative value of all possible moves as far ahead in the game as the speed of the computer allowed. “Type B” would use intelligence rather than raw power, imbuing the computer with an understanding of the game that would allow it to focus on a small number of attractive moves while ignoring the rest. In essence, a Type B computer would demonstrate the intuition of an experienced human player.

As the review goes on to note, Shannon’s mistake was to think that Type A computers were not going to materialize. They did, with the result that the promise of AI (that it would tell us something about intelligence) fizzled as the “artificial” way that machines became “intelligent” simply abstracted away from intelligence. Or, to put it as Kasparov is quoted as putting it:  “Deep Blue [the machine that beat Kasparov, NH] was intelligent the way your programmable alarm clock is intelligent.”

So, the hope that AI would illuminate human cognition rested on the belief that technology and brute calculation would not be able to substitute for “intelligence.” This proved wrong, with machine learning being the latest twist in the same saga, per the review and Kasparov. 

All this fits with GM’s remarks above. What both do not emphasize enough, IMO, is something that many did not anticipate; namely that we would revamp our views of intelligence rather than question whether our programs had it.  Part of the resurgence of Empiricism is tied to the rise of the technologically successful machine. The hope was that trying to get limited machines to act like we do might tell us something about how we do things. The limitations of the machine would require intelligent design to get it to work thereby possibly illuminating our kind of intelligence. What happened is that getting computationally miraculous machines to do things in ways that we had earlier recognized as dumb and brute force (and so telling us nothing at all) has transformed into the hypothesis that there is no such things as real intelligence at all and everything is “really” just brute force. Thus, the brain is just a data cruncher, just like Deep Blue is. And this shift in attitude is supported by an Empiricist conception of mind and explanation. There is no structure to the mind beyond the capacity to mine the inputs for surfacy generalizations. There is no structure to the world beyond statistical regularities. On this Eish viw, AI has not failed, rather the right conclusion is that there is less to thinking than we thought. This invigorated Empiricism is quite wrong. But it will have staying power. Nobody should underestimate the power that a successful (money making) tech device can have on the intellectual spirit of the age.


[1] Chomsky makes the same point recently, and he is still right. See here for discussion and links to article.
[2] This should again sound familiar. It is the moral that Chang drew on her work on faces as discussed here.
[3] I myself once made similar points in a paper with Elan Dresher. Few papers have been more fun to write. See here for the appraisal (if interested).

Wednesday, May 25, 2016

AI now

Here are some papers on the current hype over AI. They are moderately skeptical about the current state of the art. Not so much about whether there are tech breakthroughs to be had. All agree that these are forthcoming. The skepticism concerns the implications of this. Let me a say a word or two about this.

There is a lot of PR concerning how Big Data will revolutionize our conceptions of how the mind works and how science should be conducted. Big Data is Empiricism on steroids. It is made possible because of hardware breakthroughs in memory and speed of computation. We can do more of what we have always done faster and this can make a difference. I doubt that this tells us much about human cognition. Or more accurately, what it does tell us is likely wrong. Big Data is often coupled with Deep Learning. And linguists have every reason to believe that Deep Learning is an incorrect model of human cognition. Why? Because it is a modern version of the old discovery procedure. Level1 generalizations are generalized again at level 2 and level two generalizations are generalized agains t level 3 and so on. As a model of cognition, this tells us that higher levels are just generalizations over lower ones (e.g. from phonemes we get morphemes and from morphemes we get we get phrase structure and from phrase structure we get...). GG started form the demonstration that this is an incorrect understanding of linguistic organization. Levels exist, but they are in no sense reducible to the ones lower down. Indeed, whether it makes sense to speak of 'higher' and 'lower' is quite dubious. The levels interact but don't reduce. And any theory of learning that supposes otherwise is wrong if this is right (and there is no reason to think that it is not right, or at least no argument has been presented arguing against level independence). So, Deep Learning and Big Data are. IMO, dead theories walking. We will see this very soon.

The interview with Gary Marcus (here) discusses these issues and notes that historically what we have here is more of the same that has in the past proven to be wildly oversold. He thinks (and I agree) that we are getting another snow job this time around too. The interview rambles somewhat (bad editing) but there is lots in here to provoke thought.

A second paper on a similar theme is here in Aeon. The fact that it is in Aeon should not be immediately held against it. True, there is reason to be suspicious given the track record, but the paper was not bad, IMO. It argues that there is no "coming of the machines."

Here is a third piece on programming and advice about how not to do it (the Masaya). It interestingly argues for a Marrian conception of programming. Understand the computational problem before you write code. Seems like reasonable advice.

Last point: I mentioned above that Big Data is not only influencing how we conceive of Minds and Brains but also on how we should do science. The idea seems to be that that with enough data, the search for basic causal architecture becomes quaint and unnecessary. We can just vacuum up the data and the generalizations of scientific utility will pop out. On this view, theories (and not only minds) are just compendia of data generalizations and given that we can now construct these compendia more efficiently and accurately by analyzing more and more data, theory construction becomes a quaint pastime. The only real problem with Empiricism is that we did not gather enough data fast enough. And now, Big Data can fix this. The dream of understanding without thinking is finally here. Oy vey!

Monday, July 20, 2015

Some things to read

I’ve recently run across three little articles that GGers should find interesting (Thx Colin).

The first features David Poeppel. He discusses some recent work by Edward Chang at UCSF who has discovered neurons that respond to “basic “phonemic features,” rather than to phonemes themselves, which are larger chunks of sound.” As David points out in his remarks, this is not news to linguists. But it is a nice example of how what we have discovered, using pretty conventional analytic methods within phono, have mapped pretty directly onto basic neuro units. So when someone asks how linguistics is relevant to the brain sciences, this is a nice case to pull out and wave about. One more thing caught my eye: the relation between brain and cognition is here executed at the level of primitives. As Chang put it, features are the “building blocks of speech and language.” It makes sense to think that the easiest bridges to the brain will be from our primitives to theirs. The reason is that once one gets by this very simple stage complexity grows rapidly and the mapping likely becomes more and more obscure. This recalls the Embick and Poeppel observations concerning the Granularity Mismatch Problem (see here for some discussion). If GG is to make contact with neuroscience (which, IMO, it had better do or risk irrelevance, or worse), it will need to find ways if breaking the complexities of grammar into simpler and simpler sub-units and primitive operations. Happily, GG has been doing this since the mid 80s when constructions were broken down into more primitive operations like ‘moe alpha.’ As all of you know, this is also the main animating conceit of the Minimalist Program. And if what we see above is indicative, this is a very good thing for it’s at the level of simple units and basic operations that we might hope to find links to the neurosciences.

Marcus continues this theme in two pieces. The first is a piece in the NYTs (here) in which Gary argues that we need to return to the idea that the brain is a kind of computer. Apparently, this idea is no longer in fashion. Gary explains why the reasons for dumping this analogy are very weak and actually counterproductive. As he points out the right question is not whether the brain is (like) a computer but what kind of computer the brain is like. This seems absolutely right. The brain is an information processing device. Thus, it computes.

Gary rebuts the standard reasons for ignoring this analogy between brains and computers. But more interesting still he provides an interesting proposal for what kind of computer the brain might be. In a paper with two colleagues (here), he suggests that the brain is very like a “FIELD programmable gate array.” What are these? Well, they

…consist of a large number of “logic block” programs that can be configured, and reconfigured, individually, to do a wide range of tasks. One logic block might do arithmetic, another signal processing, and yet another look things up in a table. The computation of the whole is a function of how the individual parts are configured. Much of the logic can be executed in parallel, much like what happens in a brain.

In the Science piece, Gary & Co propose (p. 552) that neuroscience ought to be looking not for “a single canonical circuit” but for

…a broad array of reusable computational primitives- elementary units- of processing akin to basic sets of instructions in a microprocessor- perhaps wired together in parallel…

This fits with the “basic units/operations” theme I alluded to above. Again, it is at this very basic level that we can rationally hope to make contact with the brain sciences.

One more observation: the Science piece reads a little like a minimalist manifesto for the neurosciences. Most of cognition is made up of several simple operations that link together in different ways for different ends. Thus we should expect to find the same primitives again and again across cognition. This is music to a minimalist’s ears: most knowledge of language resides in recycled basic cognitive circuits. If there is something linguistically special (and I believe that the evidence to date is that there is) then it is a pretty small addition to this basic inventory of primitives. Our job is two fold: (i) to identify that special linguistic operation and (ii) to show that all the rest of G competence can be reduced/analyzed in terms of the remaining cognitively general ones. Conceptually, this program in linguistics fits snugly with the one that Marcus & friends outline for cogneuro, and that is just great!

One really last observation: Gary & Co end their Science piece with the following inspirational peroration (my emphasis NH):

Neuroscience must develop precisely the sorts of experimental tools, detailed brian maps, and computational infrastructures that today’s brain initiatives aim to support, but also a new set of intellectual tools for understanding how, even in principle, systems might bridge from neuronal networks to symbolic cognition…


In other words, it ends in praise of research on how things could work as well as how they do work.  I could not agree more, not only in neuroscience, but in linguistics as well. The obsession with the actual is often a barrier to intellectual progress. For the big questions, it is indispensible. Sadly, IMO, it is often devalued. But you already know that I think this. I find comfort in finding that I am not alone.

Thursday, July 17, 2014

Big money, big science and brains

Gary Marcus here discusses a recent brouhaha taking place in the European neuro-science community. The kerfuffle, not surprisingly, is about how to study the brain. In other words, it's about money. The Europeans have decided to spend a lot of Euros (real money!) to try to find out how brains function. Rather than throw lots of it at many different projects haphazardly and see which gain traction, the science bureaucrats in the EU have decided to pick winners (an unlikely strategy for success given how little we know, but bureaucratic hubris really knows no bounds). And, here’s a surprise, many of those left behind are complaining. 

Now, truth be told, in this case my sympathies lie with (at least some) of those cut out.  One of these is Stan Dehaene, who, IMO, is really one of the best cog-neuro people working today.  What makes him good is his understanding that good neuroscience requires good cognitive science (i.e. that trying to figure out how brains do things requires having some specification of what it is that they are doing). It seems that this, unfortunately, is a minority opinion. And this is not good. Marcus explains why.

His op-ed makes several important points concerning the current state of the neuro art in addition to providing links to aforementioned funding battle (I admit it: I can’t help enjoy watching others fighting important “intellectual battles” that revolve around very large amounts of cash). His most important point is that, at this point in time, we really have no bridge between cognitive theories and neuro theories. Or as Marcus puts it:

What we are really looking for is a bridge, some way of connecting two separate scientific languages — those of neuroscience and psychology.

In fact, this is a nice and polite way of putting it. What we are really looking for is some recognition from the hard-core neuro community that their default psychological theories are deeply inadequate. You see, much of the neuro community consists of crude (as if there were another kind) associationists, and the neuro models they pursue reflect this. I have pointed to several critical discussions of this shortcoming in the past by Randy Gallistel and friends (here).  Marcus himself has usefully trashed the standard connectionist psycho models (here). However, they just refuse to die and this has had the effect of diverting attention from the important problem that Marcus points to above; finding that bridge.

Actually, it’s worse than that. I doubt that Marcus’s point of view is widely shared in the neuro community. Why? They think that they already have the required bridge. Gallistel & King (here) review the current state of play: connectionist neural models combine with associationist psychology to provide a unified picture of how brains and minds interact.  The problem is not that neuroscience has no bridge, it’s that it has one and it’s a bridge to nowhere. That’s the real problem. You can’t find what you are not looking for and you won’t look for something if you think you already have it.

And this brings us back to the aforementioned battle in Europe.  Markham and colleagues have a project.  It is described here as attempting to “reverse engineer the mammalian brain by recreating the behavior of billions of neurons in a computer.” The game plan seems to be to mimic the behavior of real brains by building a fully connected brain within the computer. The idea seems to be that once we have this fully connected neural net of billions of “neurons” it will become evident how brains think and perceive. In other words, Markham and colleagues “know” how brains think, it’s just a big neural net.[1] What’s missing is not the basic concepts, but the details. From their point of view the problems is roughly to detail the fine structure of the net (i.e. what’s connected to what). This is a very complex problem for brains are very complicated nets. However, nets they are. And once you buy this, then the problem of understanding the brain becomes, as Science put it (in the July 11/2014 issue), “an information technology” issue.[2]

And that’s where Marcus and Dehaene and Gallistel and a few notable others disagree: they think that we still don’t know the most basic features of how the brain processes information. We don’t know how it stores info in memory, how it retrieves it from memory, how it calls functions, how it binds variables, how, in a word, it computes. And this is a very big thing not to know. It means that we don’t know how brains incarnate even the most basic computational operations.

In the op-ed, Marcus develops an analogy that Gallistel is also fond of pointing to between the state of current neuroscience and biology before Watson and Crick.[3]  Here’s Marcus on the cognition-neuro bridge again:

Such bridges don’t come easily or often, maybe once in a generation, but when they do arrive, they can change everything. An example is the discovery of DNA, which allowed us to understand how genetic information could be represented and replicated in a physical structure. In one stroke, this bridge transformed biology from a mystery — in which the physical basis of life was almost entirely unknown — into a tractable if challenging set of problems, such as sequencing genes, working out the proteins that they encode and discerning the circumstances that govern their distribution in the body.
Neuroscience awaits a similar breakthrough. We know that there must be some lawful relation between assemblies of neurons and the elements of thought, but we are currently at a loss to describe those laws. We don’t know, for example, whether our memories for individual words inhere in individual neurons or in sets of neurons, or in what way sets of neurons might underwrite our memories for words, if in fact they do.

The presence of money (indeed, even the whiff of lucre) has a way of sharpening intellectual disputes. This one is no different. The problem from my point of view is that the wrong ideas appear to be cashing in. Those controlling the resources do not seem (as Marcus puts it) “devoted to spanning the chasm.” I am pretty sure I know why too: they don’t see one. If your psychology is associationist (even if only tacitly so), then the problem is one of detail not principle. The problem is getting the wiring diagram right (it is very complex you know), the problem is getting the right probes to reveal the detailed connections to reveal the full networks. The problem is not fundamental but practical; problems that we can be confident will advance if we throw lots of money at them.

And, as always, things are worse than this. Big money calls forth busy bureaucrats  whose job it is to measure progress, write reports, convene panels to manage the money and the science.  The basic  problem is that fundamental science is impossible to manage due to its inherent unpredictability (as Popper noted long ago). So in place of basic fundamental research, big money begets big science which begets the strategic pursuit of the manageable. This is not always a bad thing.  When questions are crisp and we understand roughly what's going on big science can find us the Higgs field or W bosons. However, when we are awaiting our "breakthrough" the virtues of this kind of research are far more debatable. Why? Because in this process, sadly, the hard fundamental questions can easily get lost for they are too hard (quirky, offbeat, novel) for the system to digest. Even more sadly, this kind of big money science follows a Gresham’s Law sort of logic with Big (heavily monied) Science driving out small bore fundamental research. That’s what Marcus is pointing to, and he is right to be disappointed.




[1] I don’t understand why the failure of the full wiring diagram of the nematode (which we have) to explain nematode behavior has not impressed so many of the leading figures in the field (Cristof Koch is an exception here).  If the problem were just the details of the wiring diagram, then the nematode “cognition” should be an open book, which it is most definitely not. 
[2] And these large scale technology/Big Data projects are a bureaucrats dream. Here there is lots of room to manage the project, set up indices of progress and success and do all the pointless things that bureaucrats love to do. Sadly, this has nothing to do with real science.  Popper noted long ago that the problem with scientific progress is that it is inherently unpredictable. You cannot schedule the arrival of breakthrough ideas.  But this very unpredictability is what makes such research unpalatable to science managers and why it is that they prefer big all encompassing sciency projects to the real thing. 
[3] Gallistel has made an interesting observation about this earlier period in molecular biology. Most of the biochemistry predating Watson and Crick has been thrown away.  The genetics that predates Watson and Crick has largely survived although elaborated.  The analogy in the cognitive neurosciences is that much of what we think of as cutting edge neuroscience might possibly disappear once Marcus’s bridge is built. Cognitive theory, however, will largely remain intact.  So, curiously, if the prior developments in molecular biology are any guide, the cognitive results in areas like linguistics, vision, face recognition etc. will prove to be far more robust when insight finally arrives than the stuff that most neuroscientists are currently invested in.  For a nice discussion of this earlier period in molecular biology read this. It’s a terrific book.