Comments

Showing posts with label neural nets. Show all posts
Showing posts with label neural nets. Show all posts

Sunday, November 20, 2016

Revisiting Gallistel's conjecture

I recently received two papers that explore Gallistel’s conjecture (see here for one discussion) concerning the locus of neuronal computation. The first (here) is a short paper that summarizes Randy’s arguments and suggests a novel view of synaptic plasticity. The second (here: http://www.nature.com/nature/journal/v538/n7626/full/nature20101.html)[1] accept Randy’s primary criticism of neural nets and couples a neural net architecture with a pretty standard external memory system. Let me say a word about each.

The first paper is by Patrick Trettenbrein (PT) and it appears in Frontiers in Systems Neuroscience. It does three things.

First, it reviews the evidence against the idea that brains store information in their “connectivity profiles” (2). This is the classical assumption that inter-neural connection strengths are the locus of information storage. The neurophysiological mechanisms for this are long term potentiation (LTP) and long term depression (LTD). LTP/D are the technical terms for whatever strengthens or weakens interneuron connections/linkages. I’ve discussed Gallistel and Matzel’s (G&M) critique of the LTP/D mechanisms before (see here). PT reviews these again and emphasizes G&M’s point that there is an intimate connection between this Hebbian “fire together wire together” LTP/D based conception of memory and associationist psychology. As PT puts it: “Crucially, it is only against this background of association learning that LTP and LTD seem to provide a neurobiologically as well as psychologically plausible mechanism for learning and memory” (88). This is why if you reject associationsim and endorse “classical cognitive science” and its “information processing approach to the study of the mind/brain” you will be inclined to find contemporary connectionist conceptions of the brain wanting (3).

Second, there is recent evidence that connection strength cannot be the whole story. PT reviews the main evidence. It revolves around retaining memory traces despite very significant alterations in connectivity profiles. So, for example, “memories appear to persist in cell bodies and can be restored after synapses have been eliminated” (3), which would be odd if memories lived in the synaptic connections. Similarly it has recently been shown that “changes in synaptic strength are not directly related to storage of new information in memory” (3). Finally, and I like this one the best (PT describes it as “the most challenging to the idea that the synapse is the locus of memory in the brain”), PT quotes a 2015 paper by Bizzi and Ajemian which makes the following point:

If we believe that memories are made of patterns of synaptic connections sculpted by experience, and if we know, behaviorally, that motor memories last a lifetime, then how can we explain the fact that individual synaptic spines are constantly turning over and that aggregate synaptic strengths are constantly fluctuating?

Third, PT offers a reconceptualization of the role these neural connections. Here’s an extended quote (5):

…it occurs to me that we should seriously consider the possibility that the observable changes in synaptic weights and connectivity might not so much constitute the very basis of learning as they are the result of learning.

This is to say that once we accept the conjecture of Gallistel and collaborators that the study of learning can and should be separated from the study of memory to a certain extent, we can reinterpret synaptic plasticity as the brain's way of ensuring a connectivity and activity pattern that is efficient and appropriate to environmental and internal requirements within physical and developmental constraints. Consequently, synaptic plasticity might be understood as a means of regulating behavior (i.e., activity and connectivity patterns) only after learning has already occurred. In other words, synaptic weights and connections are altered after relevant information has already been extracted from the environment and stored in memory.

This leaves a place for connectivity, but not as the mechanism of memory but as what allows memories to be efficiently exploited.[2] Memories live within the cell but putting these to good use requires connections to other parts of the brain where other cells store other memories. That’s the basic idea. Or as PT puts it (6):

The role of synaptic plasticity thus changes from providing the fundamental memory mechanism to providing the brain’s way of ensuring that its wiring diagram enables it to operate efficiently…

As PT notes, the Gallistel conjecture and his tentative proposal are speculative as theories of the relevant cell internal mechanisms don’t currently exist. That said, neuroiphsyiological (and computational, see below) evidence against the classical Hebbian view are mounting and the serious problems for storing memories in usable form in connections strengths (the bases of Gallistel’s critique) are becoming more and more well recognized.

This brings us to the second Nature paper noted above. It endorses the Gallistel critique of neural nets and recognizes that neural net architectures are poor ways of encoding memories. It adds a conventional RAM to a neural net and this combination allows the machine to “represent and manipulate complex data structures.”

Artificial neural networks are remarkably adept at sensory processing, sequence learning and reinforcement learning, but are limited in their ability to represent variables and data structures and to store data over long timescales, owing to the lack of an external memory. Here we introduce a machine learning model called a differentiable neural computer (DNC), which consists of a neural network that can read from and write to an external memory matrix, analogous to the random-access memory in a conventional computer. Like a conventional computer, it can use its memory to represent and manipulate complex data structures, but, like a neural network, it can learn to do so from data.

Note that the system is still “associationist” in that learning is largely data driven (and as such will necessarily run into PoS problems when applied to any interesting cognitive domain like language) but it at least recognizes that neural nets are not good for storing information. This latter is Randy’s point. The paper is significant for it comes from Google’s Deep Mind Project and this means that Randy’s general observations are making intellectual inroads with important groups. Good.

However, this said, these models are not cognitively realistic for they still don’t make room for the domain specific knowledge that we know characterizes (and structures) different domains. The main problem remains the associationism that the Google model puts at the center of the system. As we know that associationism is wrong and that real brains characterize knowledge independently of the “input,” we can be sure that this hybrid model will need serious revision if intended as a good cog-neuro model.

Let me put this another way. Classical cog sci rests on the assumption that representations are central to understanding cognition. Fodor and Pylyshyn and Marcus long ago agued convincingly that connectionism did not successfully accommodate representations (and, recall, that connectionist agreed that their theories dumped representations) and that this was a serious problem for connectionist/neural net architectures. Gallistel further argued that neural nets were poor models of the brain (i.e. and not only of the mind) because they embody a wrong concpetion of memory; one that that makes it hard to read/write/retrieve complex information (data structures) in usable form. This, Gallistel noted, starkly contrasts with more classical architectures. The combined Fodor-Pylyshyn-Marcus-Gallistel critique then is that connectionist/neural net theories were a wrong turn because they effectively eschewed representations and that this is a problem both from the cognitive and the neuro perspective. The Google Nature paper effectively concedes this point, recognizes that representations (i.e. “complex data structures) are critical  and resolves the problem by adding a classical RAM to a connectionist front end.

However, there is a second feature of most connectionist approaches that is also wrong. Most such architectures are associationist. They embody the idea that brains are entirely structured by the properties of the inputs to the system. As PT puts it (2):

Associationism has come in different flavors since the days of Skinner, but they all share the fundamental aversion toward internally adding structure to contingencies in the world (Gallistel and Matzel 2013).

Yes! Connectionists are weirdly attracted to associationism as well as rejecting representations. This is probably not that surprising. Once on thinks of representations then it quickly becomes clear that many of their properties are not reducible to statistical properties of the inputs. Representations have formal properties above and beyond what one finds in the input, which, once you look, are found to be causally efficacious. However, strictly speaking associationsim and anti-representationalism are independent dimensions. What makes Behaviorists distinctive among Empiricists is their rejection of representations. What unifies all Empiricists is their endorsement of associationism. Seen form this perspective, Gallistel and Fodor and Pylyshyn and Marcus have been arguing that representations are critical. The Google paper agrees. This still leaves associationism however, and position the Googlers embrace.[3]

So is this a step forward? Yes. It would be a big step forward if the information processing/representational model of the mind/brain became the accepted view of things, especially in the brain sciences. We could then concentrate (yet again) all of our fire on pernicious Empiricism so many Cog-neuro types embrace.[4] But, little steps my friends, little steps. This is a victory of sorts. Better to be arguing against Locke and Hume than Skinner![5]

That’s it. Take a look.



[1] Thx to Chris Dyer for bringing the paper to my attention. I put in the URL up rather than link to the paper directly as the linking did not seem to work. Sorry.
[2] Redolent of a competence/performance distinction, isn’t it?  The physiological bases of memory should not be confused with the physical bases for the deployment of memory.
[3] I should add that it is not clear that the Googlers care much about the cog-neuro issues. Their concerns are largely technological, it seems to me. They live in a Big Data world, not one where PoS problems (are thought to) abound. IMO, even in a uuuuuuge data environment, PoS issues will arise, though finding them will take more cleverness. At any rate, my remarks apply to the Google model as if intended as a cog-neuro one.
[4] And remember, as Gallistel notes (and PT emphasizes) much of the connectionism one sees in the brain sciences rests on thinking that the physiology has a natural associationist interpretation psychologically. So, if we knock out one strut, the other may be easier to dislodge as well (I know that this is wishful thinking btw).
[5] As usual, my thinking on these issues was provoked by some comments by Bob Berwick. Thx.

Monday, November 23, 2015

The concise Gallistel on how brains compute

Jeff Lidz sent me this great little piece by Randy Gallistel on his favorite theme: how most neuroscientists have misunderstood how brains compute. I’ve discussed Randy’s stuff in various FoL posts (here, here, and here). Here in just four lucid pages, Randy makes his main point again. If he is right (and the form of his argument seems impeccable to me), then much of what goes on in neuroscience is just plain wrong. Indeed, if Randy is right, then current neo-connectionist/neural net assumptions about the brain are about as accurate as 1950s-60s behaviorist conceptions were about the mind. In other words, at best of tertiary interest and, more likely, deserving to be completely forgotten.[1] At any rate, Randy here makes four main points.

First, that there is recent evidence (discussed here) strongly pointing to the conclusion that information can be stored inside a single neuron (rather than in connections of many neurons).

Second, that there is scads of behavioral evidence showing that brains store number values and that there is no way of storing numbers this in connection weights, thus implying that any theory of the brain that limits itself to this kind of hardware must be at best incomplete and at worst wrong.

Third, that there is a close connection between neural net “plasticity” conceptions of the brain and traditional empiricist conceptions of the mind (especially learning). In fact, Randy argues that these are largely flip sides of the same coin.

Fourth, that brains already contain all the hardware that is required to function like classical computers, the latter being the perfect complements for the computational cognitive theories that replaced behaviorism.

And all in four pages.

There is one argument that Randy hints at but doesn’t stress that I would like to add to his four. It is a conceptual argument. Here it is.

Whatever one thinks of cognition, it is clear that animals use large molecules like DNA and RNA for information processing. Indeed, this is now standard biological dogma. As Gallistel and King (here) illustrates, this system has all the capacities of a classical computer (addresses, read-write memory, variables, binding etc.). So here’s the conceptual argument: imagine that you had an animal with the wherewithal to classically compute hereditary information but instead of repurposing (exapting) this system for cognitive ends it developed an entirely different additional system for this purpose. In other words, it had all it needed sitting there but ignored these resources and embodied cognition in a completely different way. Does this seem plausible? Is this the way evolution typically works? Isn’t opportunism the main mover in the evolution game? And if it is, doesn’t this suggest that Randy’s conjecture must be right? In fact, wouldn’t it be weird if large chunks of cognition did not exploit that computational machinery already sitting there in DNA/RNA and other large molecules? In fact, wouldn’t the contrary assumption bear a huge burden of proof? Well, you know what I think!

Why is this not the common perception? Why is Randy’s position considered exotic? Here’s the one word answer: Empiricism! In the cog-neuro world this is the default view. There is little to empirically support this conception (see here for a review of the pas de deux between unsupported empiricism in psychology and tendentious reasoning in neural net neuroscience). Indeed, it largely flourishes when we know next to nothing about some domain of inquiry. However, it is the default conception of the mind. What Randy is pointing out (and has repeatedly pointed out and is right to point out) is that it is fatally flawed, not only as a theory of mind but also as a theory of the brain. And its flaws are conceptual as well as empirical. I can’t wait for the day that this becomes the conventional wisdom, though given the methodological dualism characteristic of the cog-neuro-sciences, I suspect that this day is not just around the corner. Too bad.



[1] Note that I say “deserving” of amnesia. This concedes the sad fact that neo-behaviorism is making a vigorous comeback within cognition. Yet another indication of the collapse of civilization.

Sunday, February 10, 2013

Another Impossibility Argument


Things that are not possible don’t happen.  Sounds innocuous huh?  Perhaps, but that’s the basis of POS arguments.  If the evidence cannot dictate the outcome of acquisition, and nonetheless the outcome is not random, then there must be something other than the evidence guiding the process. Substitute ‘PLD’ for ‘evidence’ and ‘UG’ for ‘something other than the evidence’ and you get your standard POS argument. There have been and can be legitimate fights about how rich the data/PLD is and how rich UG is, but the form of the argument is dispositive and that’s why it’s so pretty and so useful.  A good argument form is worth its weight in theoretical gold. 

That this is what I believe is not news for those of you who have been following this blog.  And don’t worry, at least for today, I am not going to present another linguistic POS argument (although I am tempted to generate some new examples so that we can get off of the Yes/No question data).  Rather, what I want to do is publicize another application of the same argument form, the one deployed in Gallistel/King (G/K) in favor of the conclusion that connectionist architectures are biologically impossible.

The argument that they provide is quit simple: connectionism requires too many neurons to code various competences.  They dub the problematic fact over which connectionist models necessarily stumble and fall ‘the infinitude of the possible’ (IoP). The problem as they understand it is endemic; computations in a connectionist/neural net architecture (C/NN) cannot be “implemented by compact procedures.” This means that such a C/NN cannot “produce as an output an answer that the maker of the system did not hard wire into the look-up tables. (261)” In effect, C/NNs are fancy lists (aka look-up tables) where all possibilities are computed out rather than being implicit in more compact form in a generative procedure. And this leads to a fundamental problem: the brain is just not big enough to house the required C/NNs.  Big as brains are, they are still too small to explicitly code all the required possible realizable cognitive states.

G/K’s argument is all in service of arguing that neuroscientists must assume that brains are effectively Turing-von Neumann (TvN) machines with addressable, symbolic, read/write memories. In a nutshell:

 … a critical distinction between procedures implemented by means of look-up tables and … compact procedures …is that the specification of the physical structure of a look-up table requires more information than will ever be extracted by the use of that table.  By contrast, the information required to specify the structure of a mechanism that implements a compact procedure may be hundreds of orders of magnitude less than the information that can be extracted using that mechanism (xi).

What’s the argument? Like POS arguments, it starts with a rich description of various animal competences. The three that play staring roles in the book are dead reckoning in ants, bee dancing and food caching behavior in jays.  Those of you who like Animal Planet will love these sections. It is literally unbelievable what these bugs and birds can do.  Here is a glimpse of jay behavior as reported in G/K (c.f. 213-217).

In summer, when times are good, scrub jays collect and store food in different locations for later winter feasting. They cache this food in as many as 10,000 different locations. Doing this involves remembering what they hid, where they hid it, when they hid it, whether they emptied it, if the morsel was tasty, how quickly the morsel goes bad, and who was watching them when they hid it. This is a lot of information. Moreover, it is very specific information, sensitive to six different parameters.  Moreover, the values for these parameters are indeterminate and thus the number of possible memories these jays can produce and access is potentially unbounded.  Though there is an upper bound on the actual memories stored, the number of potentially storable memories is effectively unbounded (aka, infinite).  This is the big fact and it has a big implication. In order to store these memories the jays need some sort of template that roughly says ‘stored X at Y at time Z, X goes bad in W days, X is +/- retrieved, X is +/- tasty, storing was +/- observed.’ This template requires a brain/mind that can link variables, value variables, write to memory and retrieve from memory so as to store useful information and access it when necessary.  Note, we can treat this template as a large sentence frame, much like ‘X weighs Y pounds’ and like the latter there is no upper bound on the number of possible realizations of this frame (e.g. John weighs 200 pounds, Mary weighs 90 pounds, Trigger weighs 1000 pounds etc.).  These templates combined with the capacity to put actual food type/time/place/ etc. values for the variables constitute “compact procedures” for coding the relevant actual information required. Notice how “small” it is relative to the number of actual instances of such templates (finite specification versus unbounded number of instances).

If this description is correct (G/K review the evidence extensively), here is what is neuronally impossible: to list all the potential instantiations of this kind of proposition and simply choose the ones that are actual. Why?

… the infinitude of the possible looms. There are many possible locations, many possible kinds of food, many possible rates of decay, many possible delays between caching and recovery – and no restrictions on the possible combinations of these possibilities. No architecture with finite resources can cope with this infinitude by allocating resources in advance to every possible combination. (217)

However, current neuroscience takes read/write memory (a necessary feature of a system able to code the above information) to be neurobiologically implausible. Thus, current neuroscience only investigates systems (viz. C/NNs) that cannot in principle handle these kinds of behavioral data.  What’s the principled reason for this inadequacy? Computation in C/NNs is not implemented by compact procedures. Rather, C/NNs are effectively elaborate look-up tables and so cannot “output an answer that the maker of the system did not hard wire into one of its look-up tables.” (261).

That’s the impossibility argument. If we assume that brains mediate cognition then it cannot be the case that animal brains are C/NN devices.

I strongly recommend reading these sections of G/K’s book. There is terrific detailed discussion of how many neurons it would take to realize a C/NN capable of dead reckoning. By G/K’s estimates (261) it would take all the neurons in an ant’s brain (ants are terrific dead reckoners) to realize a more or less adequate system of dead reckoning.

This is fun stuff, but I am no expert in these matters and though I tend to trust Gallistel when he tells me something, I am in no position to independently verify his calculations regarding the number of neurons required to implement a reasonable C/NN device.[1]  However, G/K’s conclusion should resonate with generativists.  Grammars are compact procedures for coding what is essentially an infinite number of possible sentences and UG is (part of ) a compact procedure for coding what is (at the very least) a very large number of possible Gs.[2] Thus, whatever might be true of other animals, human brains clearly capable of language cannot be C/NNs.[3] Why do I mention this. For two reasons:

First, there is still a large cohort of neuroscientists, psychologists and computationalists who try to analyze linguistic phenomena in C/NN terms. They look at pictures of brains, see interconnected neurons, and conclude that our brains are C/NNs and that our linguistic competence must be analysed to fit these “data.” G/K argue that this is exactly the wrong conclusion to draw. Note that the IoP is trivially obvious in the domain of language, so the G/K argument is very potent here. And that is cause for celebration as these same net-enchanted types are also rather grubby empiricists.

G/K discuss the unholy affinities between associationism and C/NN infatuation (c.f. chapter 11) (the slogan “what fires together wires together” reveals all).  Putting a stake through the C/NN worldview also serves to weaken the empiricist-associationist learning conception of acquisition.[4] I doubt it will kill it. Nothing can it appears. But perhaps it can intellectually wound it yet again (though only if the G/K material is taken seriously, which I suspect it won’t be either because the net-nuts won’t get it or they will simply ignore it). So an attack on C/NNs is also a welcome attack on empiricist/associationist conceptions and that’s always a valuable public service.

Second, this is very good news for the minimalistically inclined. Here’s why.  Minimalists are counting on animal brains being similar enough to human brains for there to be features of the former that can be used to explain some of the properties of FL/UG.  Recall the conceit: take the cognitive capacities of our ancestors as given and ask what we need to add to get linguistically capable minds/brains.  However, were animal brains C/NNs and ours clearly are not (recall how easy it is to launch the IoP considerations in the domain of language) then it is very hard to see how something along these lines could be true.  Is it really plausible that the shift to language brought with it an entirely new kind of brain architecture?  The question answers itself.  So G/K’s conclusions regarding animal brain architectures is very good news.

Note, we can turn this argument around: if as minimalisms requires (and seems independently reasonable) human brains are continuous with non-human ones and if the IoP requires that brains have TvN architectures, then language in humans provides a strong prima facie argument for TvNs in animals.  For the reasons that G&K offer (not unlike those Chomsky originally deployed) human linguistic competence cannot supervene on C/NN systems. And as G/K note, this has profound implications for current work in neuroscience, viz. the bulk of the work is looking in the wrong places for answers that cannot possibly serve. Quite a conclusion, but that’s what a good impossibility argument delivers.

Let me end with one last observation: there is a tendency to think that neuroscientists hold the high intellectual ground and cognitive science/linguistics must accommodate itself to its conclusions. G/K demonstrate that this is nonsense.  Cognition supervenes on brains. If some kind of brain cannot support what we know are the cognitive facts, then this view of the brain must be wrong.  Anything else would be old fashion dualism. Hmm, wouldn’t it be amusing if to defend their actual practice neuroscientists had to endorse hard core Cartesian dualism? Yet, if G/K are right, that’s what they are effectively doing right now.



[1] Note that the G/K argument is about physical implementation and so is additional (though related) to the arguments presented by Fodor and Pylyshyn or Marcus against connectionism. Not only does C/NN get the cognition wrong, it is also physically impossible given how many neurons it would take to implement a C/NN device to do even a half-assed job. 
[2] If UG is parametric with a finite number of parameters then the number of possible Gs will also be finite. However, it is virtually certain that the number of parameters (if there are parameters) is very high (see here for discussion) and so the space of possibilities is very very large. OF course, if there are no parameters, then there need be no upper bound on the number of possible Gs.
[3] Which does not mean that for some kinds of pattern recognition the brain cannot use C/NNs. The point is that these are not sufficient.
[4] The cognoscenti will notice that I am here carrying on my effort to change our terminology so that we use ‘learning’ to denote one species of acquisition, the species tightly tied to empiricism/associationism.