Comments

Showing posts with label Fodor. Show all posts
Showing posts with label Fodor. Show all posts

Sunday, November 20, 2016

Revisiting Gallistel's conjecture

I recently received two papers that explore Gallistel’s conjecture (see here for one discussion) concerning the locus of neuronal computation. The first (here) is a short paper that summarizes Randy’s arguments and suggests a novel view of synaptic plasticity. The second (here: http://www.nature.com/nature/journal/v538/n7626/full/nature20101.html)[1] accept Randy’s primary criticism of neural nets and couples a neural net architecture with a pretty standard external memory system. Let me say a word about each.

The first paper is by Patrick Trettenbrein (PT) and it appears in Frontiers in Systems Neuroscience. It does three things.

First, it reviews the evidence against the idea that brains store information in their “connectivity profiles” (2). This is the classical assumption that inter-neural connection strengths are the locus of information storage. The neurophysiological mechanisms for this are long term potentiation (LTP) and long term depression (LTD). LTP/D are the technical terms for whatever strengthens or weakens interneuron connections/linkages. I’ve discussed Gallistel and Matzel’s (G&M) critique of the LTP/D mechanisms before (see here). PT reviews these again and emphasizes G&M’s point that there is an intimate connection between this Hebbian “fire together wire together” LTP/D based conception of memory and associationist psychology. As PT puts it: “Crucially, it is only against this background of association learning that LTP and LTD seem to provide a neurobiologically as well as psychologically plausible mechanism for learning and memory” (88). This is why if you reject associationsim and endorse “classical cognitive science” and its “information processing approach to the study of the mind/brain” you will be inclined to find contemporary connectionist conceptions of the brain wanting (3).

Second, there is recent evidence that connection strength cannot be the whole story. PT reviews the main evidence. It revolves around retaining memory traces despite very significant alterations in connectivity profiles. So, for example, “memories appear to persist in cell bodies and can be restored after synapses have been eliminated” (3), which would be odd if memories lived in the synaptic connections. Similarly it has recently been shown that “changes in synaptic strength are not directly related to storage of new information in memory” (3). Finally, and I like this one the best (PT describes it as “the most challenging to the idea that the synapse is the locus of memory in the brain”), PT quotes a 2015 paper by Bizzi and Ajemian which makes the following point:

If we believe that memories are made of patterns of synaptic connections sculpted by experience, and if we know, behaviorally, that motor memories last a lifetime, then how can we explain the fact that individual synaptic spines are constantly turning over and that aggregate synaptic strengths are constantly fluctuating?

Third, PT offers a reconceptualization of the role these neural connections. Here’s an extended quote (5):

…it occurs to me that we should seriously consider the possibility that the observable changes in synaptic weights and connectivity might not so much constitute the very basis of learning as they are the result of learning.

This is to say that once we accept the conjecture of Gallistel and collaborators that the study of learning can and should be separated from the study of memory to a certain extent, we can reinterpret synaptic plasticity as the brain's way of ensuring a connectivity and activity pattern that is efficient and appropriate to environmental and internal requirements within physical and developmental constraints. Consequently, synaptic plasticity might be understood as a means of regulating behavior (i.e., activity and connectivity patterns) only after learning has already occurred. In other words, synaptic weights and connections are altered after relevant information has already been extracted from the environment and stored in memory.

This leaves a place for connectivity, but not as the mechanism of memory but as what allows memories to be efficiently exploited.[2] Memories live within the cell but putting these to good use requires connections to other parts of the brain where other cells store other memories. That’s the basic idea. Or as PT puts it (6):

The role of synaptic plasticity thus changes from providing the fundamental memory mechanism to providing the brain’s way of ensuring that its wiring diagram enables it to operate efficiently…

As PT notes, the Gallistel conjecture and his tentative proposal are speculative as theories of the relevant cell internal mechanisms don’t currently exist. That said, neuroiphsyiological (and computational, see below) evidence against the classical Hebbian view are mounting and the serious problems for storing memories in usable form in connections strengths (the bases of Gallistel’s critique) are becoming more and more well recognized.

This brings us to the second Nature paper noted above. It endorses the Gallistel critique of neural nets and recognizes that neural net architectures are poor ways of encoding memories. It adds a conventional RAM to a neural net and this combination allows the machine to “represent and manipulate complex data structures.”

Artificial neural networks are remarkably adept at sensory processing, sequence learning and reinforcement learning, but are limited in their ability to represent variables and data structures and to store data over long timescales, owing to the lack of an external memory. Here we introduce a machine learning model called a differentiable neural computer (DNC), which consists of a neural network that can read from and write to an external memory matrix, analogous to the random-access memory in a conventional computer. Like a conventional computer, it can use its memory to represent and manipulate complex data structures, but, like a neural network, it can learn to do so from data.

Note that the system is still “associationist” in that learning is largely data driven (and as such will necessarily run into PoS problems when applied to any interesting cognitive domain like language) but it at least recognizes that neural nets are not good for storing information. This latter is Randy’s point. The paper is significant for it comes from Google’s Deep Mind Project and this means that Randy’s general observations are making intellectual inroads with important groups. Good.

However, this said, these models are not cognitively realistic for they still don’t make room for the domain specific knowledge that we know characterizes (and structures) different domains. The main problem remains the associationism that the Google model puts at the center of the system. As we know that associationism is wrong and that real brains characterize knowledge independently of the “input,” we can be sure that this hybrid model will need serious revision if intended as a good cog-neuro model.

Let me put this another way. Classical cog sci rests on the assumption that representations are central to understanding cognition. Fodor and Pylyshyn and Marcus long ago agued convincingly that connectionism did not successfully accommodate representations (and, recall, that connectionist agreed that their theories dumped representations) and that this was a serious problem for connectionist/neural net architectures. Gallistel further argued that neural nets were poor models of the brain (i.e. and not only of the mind) because they embody a wrong concpetion of memory; one that that makes it hard to read/write/retrieve complex information (data structures) in usable form. This, Gallistel noted, starkly contrasts with more classical architectures. The combined Fodor-Pylyshyn-Marcus-Gallistel critique then is that connectionist/neural net theories were a wrong turn because they effectively eschewed representations and that this is a problem both from the cognitive and the neuro perspective. The Google Nature paper effectively concedes this point, recognizes that representations (i.e. “complex data structures) are critical  and resolves the problem by adding a classical RAM to a connectionist front end.

However, there is a second feature of most connectionist approaches that is also wrong. Most such architectures are associationist. They embody the idea that brains are entirely structured by the properties of the inputs to the system. As PT puts it (2):

Associationism has come in different flavors since the days of Skinner, but they all share the fundamental aversion toward internally adding structure to contingencies in the world (Gallistel and Matzel 2013).

Yes! Connectionists are weirdly attracted to associationism as well as rejecting representations. This is probably not that surprising. Once on thinks of representations then it quickly becomes clear that many of their properties are not reducible to statistical properties of the inputs. Representations have formal properties above and beyond what one finds in the input, which, once you look, are found to be causally efficacious. However, strictly speaking associationsim and anti-representationalism are independent dimensions. What makes Behaviorists distinctive among Empiricists is their rejection of representations. What unifies all Empiricists is their endorsement of associationism. Seen form this perspective, Gallistel and Fodor and Pylyshyn and Marcus have been arguing that representations are critical. The Google paper agrees. This still leaves associationism however, and position the Googlers embrace.[3]

So is this a step forward? Yes. It would be a big step forward if the information processing/representational model of the mind/brain became the accepted view of things, especially in the brain sciences. We could then concentrate (yet again) all of our fire on pernicious Empiricism so many Cog-neuro types embrace.[4] But, little steps my friends, little steps. This is a victory of sorts. Better to be arguing against Locke and Hume than Skinner![5]

That’s it. Take a look.



[1] Thx to Chris Dyer for bringing the paper to my attention. I put in the URL up rather than link to the paper directly as the linking did not seem to work. Sorry.
[2] Redolent of a competence/performance distinction, isn’t it?  The physiological bases of memory should not be confused with the physical bases for the deployment of memory.
[3] I should add that it is not clear that the Googlers care much about the cog-neuro issues. Their concerns are largely technological, it seems to me. They live in a Big Data world, not one where PoS problems (are thought to) abound. IMO, even in a uuuuuuge data environment, PoS issues will arise, though finding them will take more cleverness. At any rate, my remarks apply to the Google model as if intended as a cog-neuro one.
[4] And remember, as Gallistel notes (and PT emphasizes) much of the connectionism one sees in the brain sciences rests on thinking that the physiology has a natural associationist interpretation psychologically. So, if we knock out one strut, the other may be easier to dislodge as well (I know that this is wishful thinking btw).
[5] As usual, my thinking on these issues was provoked by some comments by Bob Berwick. Thx.

Monday, September 15, 2014

Computations. modularity and nativism

The last post (here) prompted three useful comments by Max, Avery and Alex C. Though they appear to make three different points (Max pointing to Fodor’s thoughts on modularity, Avery on indirect negative evidence and Alex C on domain specific nativism) I believe that they all end up orbiting a similar small set of concerns. Let me explain.

Max links to (IMO) one of Fodor’s best ever book reviews (here). The review brings together many themes in discussing a pair of books (one by Pinker, the other by Plotkin). It outlines some links between computationalism, modularity, nativism and Darwininan natural selection (DNS). I’ll skip the discussion on DNS here, though I know that there will be many of you eager to battle his pernicious and misinformed views (not!).  Go at it.  What I think is interesting given the earlier post is Fodor’s linking together computationalism, modularity and nativism.  How do these ideas talk to one another? Let’s start by seeing what they are.

Fodor takes computationalism to be Turing’s “simply terrific idea” about how to mechanize rationality (i.e. thinking). As Fodor puts it (p. 2):

…some inferences are rational in virtue of the syntax of the sentences that enter into them; metaphorically, in virtue of the ‘shapes’ of these sentences.

Turing noted that, wherever an inference is formal in this sense, a machine can be made to execute the inference. This is because…you can make them [i.e. machines NH] quite good at detecting and responding to syntactic relations among sentences.

 And what makes syntax so nice? It’s LOCAL. Again as Fodor puts it (p. 3):

…Turing’s account of computation…doesn’t look past the form of sentences to their meanings and it assumes that the role of thoughts in a mental process is determined entirely by their internal (syntactic) structure.

Fodor continues to argue that where this kind of locally focused computation is not available, computationalism ceases to be useful.  When does this happen? When belief fixation requires the global canvassing and evaluation of disparate kinds of information all of which have variable and very non-linear effects on the process. Philosophers call this ‘inference to the best explanation’ (IBT) and the problem with IBT is that it’s a complete and utter mystery how it gets done.[1] Again as Fodor puts it (p. 3):

[often] your cognitive problem is to find and adopt whatever beliefs are best confirmed on balance. ‘Best confirmed on balance’ means something like: the strongest and simplest relevant beliefs that are consistent with as many of one’s prior epistemic commitments as possible. But as far as anyone knows, relevance, strength, simplicity, centrality and the like are properties, not of single sentences, but of whole belief systems: and there’s no reason at all to suppose that such global properties of belief systems are syntactic.[2]

And this is where modularity comes in; for modular systems limit the range of relevant information for any given computation and limiting what counts as relevant is critical to allowing one to syntactify a problem and allow computationalism to operate.  IMO, one of the reasons that GG has been a doable and successful branch of cog sci is that FL is modular(ish) (i.e. that something like the autonomy of syntax is roughly correct).  ‘Modular’ means “largely autonomous with respect to the rest of one’s cognition” (p. 3). Modularity is what allows Turing’s trick to operate. Turing’s trick, the mechanization of cognition, relies on the syntacticifcation of inference, which in turn relies on isolating the formal features that computations exploit.

All of which brings us (at last!) to nativism.  Modularity just is domain specificity.  Computations are modular if they are “more or less autonomous” and “special purpose” and “the information [they] can use to solve [cognitive problems] are proprietary” (p. 3).  So construed, if FL is modular, then it will also be domain specific. So if FL is a module (and we have lots of apparent evidence to suggest that it is) then it would not be at all surprising to find that FL is specially tuned to linguistic concerns. And that it exploits and manipulates “proprietary information” and that its computations were specifically “designed” to deal with the specific linguistic information it worries about.  So, if FL is a module, then we should expect it be contain lots of domain specific computational operations, principles and primitives.

How do we go about investigating the if-clause immediately above?  It helps go back to the schema we discussed in the previous post. Recall the general schema in (1) that we used to characterize the relevant problem in a given domain, ‘X’ ranging over different domains.  (2) is the linguistic case.

(1)  PXD -> FX -> GX
(2)  PLD -> FL -> GL

Linguists have discovered many properties of FL.  Before the Minimalist Program (MP) got going, the theories of FL were very linguistically parochial. The basic primitives, operations and principles did not appear to have much to say about other cognitive domains (e.g. vision, face recognition, causal inference). As such it was reasonable to conclude that the organization of FL was sui generis. And to the degree that this organization had to be take as innate (which, recall, was based on empirical arguments about what Gs did) then to that degree we had an argument for innate domain specific principles of FL.  MP has provided (a few) reasons for thinking that earlier theories overestimated the domain specificity of FL’s organization. However, as a matter of fact, the unification of FL with other domains of cognition (or computation) has been very very very modest.  I know what I am hoping for and I try not to confuse what I want to be true with what we have good reason to be true. You should too. Ambitions are one thing, results quite another. How one might go about realizing these MP ambitions?

If (1) correctly characterizes the problem, then one way for arguing against a dedicated capacity is to show that for various values of ‘X,’ FX is the same. So, say we look at vision and language, then were FL = FV we would have an argument that the very same kind of information and operations were cognitively at play in both vision and language.  I confess, that stating things this baldly makes it very implausible that FL does equal FV, but heh, it’s possible. The impressive trick would show how to pull this off (as opposed to simply expressing hopes or making windy assertions that this could be done), at least for some domains. And the trick is not an easy one to execute: we know a lot about the properties of natural language Gs. And we want an FL that explains these very properties. We don’t want a unification with other FXs that sacrifices this hard won knowledge to some mushy kind of “unification” (yes, these are scare quotes) which sacrifices the specifics that we have worked so hard to establish (yes Alex, I’m talking to you). An honest appraisal of how far we’ve come in unifying the principles across modules would conclude that, to date, we have very few results suggesting that FL is not domain specific. Don’t get me wrong: there are reasons to search for such unifications and I for one would be delighted if this happens. But hoping is not doing and ambitions are not achievements. So, if FL is not a dedicated capacity, but is merely the reflection of more general cognitive principles then it should be possible to find FL being the same as some FX (if not vision, then something else) and that this unified FX’ (i.e. which encompasses FL and FX) can derive the relevant Gs with all their wonderful properties given the appropriate PLD. There’s a Nobel prize awaiting such a unification, so hope to it.[3]

It is worth noting that there is tons of standard variety psycho evidence that FL really is modular with respect to other cognitive capacities.  Susan Curtiss (here and here) reviews the wealth of double dissociations between language and virtually any other capacity you might be interested in. Thus, at least in one perfectly coherent sense, FL is a module and so a dedicated special purpose system. Language competence swings independently of visual acuity, auditory facility, IQ, hair color, height, voacab proficiency, you name it. So if one takes such dissociations as dispositive (and it is the gold standard) then FL is a module with all that this entails.

However, there is a second way of thinking about what unification of the cognitive modules consists in and this may be the source of much (what I take to be) confused discussion. In particular, we need to separate out two questions: ‘Is FL a module?’ and ‘Is FL contain linguistically proprietary parts/circuits?’ One can maintain that FL is a module without also thinking that its parts are entirely different from those in every other module.  How so? Well, FL might be composed from the same kinds of parts present in other modules, albeit put together in distinctive ways. Same parts, same computations, different wiring. If this were so, then there would be a sense in which FL is a module (i.e. it has special distinctive proprietary computations etc.), yet when seen at the right grain it shares many (most? All?) of its basic computational features with other domains of cognition. In other words, it is possible that FL’s computations are distinctive and dedicated, and that they are built from the same simple parts found in other modules. Speaking personally, this is how I now understand the Minimalist Bet (i.e. that FL shares many basic computational properties with other systems). 

This is a coherent position (which does not imply it is correct). At the cellular level our organs are pretty similar. Nonetheless, a kidney is not a heart, and neither is a liver or a stomach.  So too with FL and other cognitive “organs.”  This is a possibility (in fact, I have argued in places that this is also plausible and maybe even true). So, seen from the perspective of the basic building blocks, it is possible that FL, though a separate module, is nonetheless “just like” every other kind of cognition. This version of the “modularity” issue asks not whether FL is a domain specific dedicated system (it is!), but whether it employs primitive circuits/operations proprietary to it (i.e. not shared with other cognitive domains). Here ‘domain specific’ means uses basic operations not attested in the other domains of non-linguistic cognition.

Of course, the MP bet is easy to articulate at a general level. What’s hard is to show that it’s true (or even plausible).  As I’ve argued before, to collect on this bet requires, first, reducing FL’s internal modularity (which in turn requires showing Binding, movement, control, agreement, etc. are really only apparently different) and, second, showing that this unification rests on cognitively generic basic operations.[4] Believe me when I tell you that this program has been a hard sell.

Moreover, the mainstream Minimalist position is that though this may be largely correct, it is exactly wrong: there are some special purpose linguistic devices and operations (e.g. Merge), which are responsible for Gs distinctive recursive property. At any rate, I think the logic is clear so I will not repeat the mantra yet again.

This brings me to the last point I want to make: Avery notes that more often than not positive evidence relevant to fixing a grammatical option is missing from the PLD.  In other words, Avery notes that the PLD is in fact even more impoverished than we tend to believe. He rightly notes that this implies that indirect negative evidence (INE) is more important than we tend to think.  Now if he is right (and I have no reason to think that he isn’t), then FL must be chocked full of domain specific information. Why? Because INE requires a sharp specification of options under consideration to be operative.  Induction that uses INE effectively must be richer than induction exploiting only positive data.[5] INE demands more articulated hypothesis space, not less. INE can compensate for poor direct evidence but only if FL knows what absences it’s looking for! You can hear the dogs that don’t bark but only if you are listening for barking dogs. If Avery’s cited example is correct (see here), then it seems that FL is attuned to micro variations, and this suggests a very rich system of very linguistically specific micro parameters internal to FL. Thus, if Avery is right, then FL will contain quite a lot of very domain specific information and given that this information is logically necessary to exploit INE it looks like these options must be innately specified and that FL contains lots of innate domain specific information. Of course, Avery may be wrong and those that don’t like this conclusion are free (indeed urged) to reanalyze the relevant cases (i.e. to indulge in some linguistic research and produce some helpful results).

This is a good place to stop.  There is an intimate connection between modularity, computationalism, and nativism. Computations can only do useful work where information is bounded. Bounded information is what modules provide. More often than not the information that a module exploits is native to it. MP is betting that with respect to FL, there is less language specific basic circuitry than heretofore assumed. However, this does not imply that FL is not a module (i.e. part of “general intelligence”). Indeed, given the kinds of evidence that Curtiss reviews, it is empirically very likely that FL is a module. And this can be true even if we manage to unify the internal modules of FL and demonstrate that the requisite remaining computations largely exploit domain general computational principles and operations. Avery’s important question remains: how much acquisition is driven by direct and how much by indirect negative evidence? Right now, we don’t really know (at least not to the level of detail that we want). That’s why these are still important research topics.  However, the logic is clear, even if the answers are not.



[1] Incidentally, IBT is one of the phenomena that dualists like Descartes pointed to in favor of a distinct mental substance. Dualism, in other words, is roughly the observation that much of thought cannot be mechanized.
[2] It’s important to understand where the problem lies. The problem is not giving a story in specific cases in specific contexts. We do this all the time. The problem is providing principles that select out the IBT antecedent to a specification of the contextually relevant variables. The hard problem is specifying what is relevant ex ante.
[3] Successful unifications almost always win kudos. Think electricity and magnetism, the the latter two with the weak force, terrestrial and celestial mechanics, chemistry and mechanics. These all get their own chapters in the greatest hits of science books. And in each case, it took lots of work to show that the desired unification was possible. There is no reason to think that cognition should be any easier.
[4] I include generic computational principles here, so-called first factor computational principles.
[5] In fact, if I understand Gold correctly (which is a toss up), acquiring modestly interesting Gs strictly using induction over positive data is impossible.

Wednesday, August 27, 2014

Nativism, Rationalism and Empiricism-1

There are two different kinds of arguments for a Rationalist approach to the study of mind. The first, so far as I can tell, is virtually tautological. The second is quite substantive.  What are they? I’ll try to lay them out in a couple of posts. This one here discusses the “tautology.”

The tautological version is well laid out in the Royaumont conference papers (here) and how they relate to the innateness “controversy.”  I put the latter in quote marks because most of the participants (including and especially Fodor and Chomsky, the so-called hard core nativists) considered the idea that the mind has innate structure nothing a simple tautology. Indeed, this is how Fodor and Chomsky repeatedly refer to the “innateness hypothesis” (see e.g. 263, 268). It’s tautological that the mind (and brain) has structure and biases as without such there can be no induction whatsoever and it is taken for granted that biological systems are constantly inducing (viz. engaging in non-demonstrative inference).  This said, it is interesting to re-read the discussions for despite this general agreement, there is lots of intellectual toing and froing. Why? Because as Chomsky puts it (see Fodor’s version p. 268):

What is important is not just to see that something is a “tautology,” but also to see its import. (262)

What’s the import, as Fodor and Chomsky understood things?  That there is no “learning” without a set of given projectable predicates that undergird it. Or, more accurately, as Fodor puts it, “ the very idea of concept learning is confused” (143).  And the confusion? Two related, but importantly different, concepts have been run together; concept acquisition (CA) and belief fixation (BF). Regarding the former, we have no theory of how concepts are acquired. What we have are theories of BF, which are, at the most general level, inductive logics of various kinds, which, by their nature, presuppose a given set of projectable predicates and so cannot themselves serve as theories of CA.  Or as Fodor puts it:

…no theory of learning that anybody has ever developed is, as far as I can see, a theory that tells you how concepts are acquired; rather such theories tell you how beliefs are fixed by experiences – they are essentially inductive logics. That kind of mechanism, which shows how beliefs are fixed by experiences, makes sense only against the background of radical nativism. (144).

Fodor and Chomsky (and most of the other participants at the Royaumont conference I might add if the comments section is any indication) believe that the above is a virtual tautology. All theories of learning are selective (i.e. stories where the given hypothesis space proposes and the incoming experience disposes).[1] Where tautology ends and (some of the) hard work begins is to specify the set of projectable predicates that are in fact biologically/cognitively given (i.e. the shape and content of the hypothesis space that BFers actually bring to the “learning” problem).  To repeat, no given space of alternatives, no way for an inductive logic or theory of BF to operate. The account of what is given is (or is a very good part of) a theory of the relevant biases.[2]

Fodor and Chomsky pull several important consequences out of this tautology.

First, that many positions confidently explored in the cognitive literature are strictly speaking incoherent as expressed. Fodor discusses the “Piagetian” view that developmental conceptual change is a learning process in which learning replaces earlier conceptually weaker stages with subsequent conceptually stronger ones.  Fodor argues that this position is, very simply, conceptually untenable. It is not untenable that development involves a succession of stages where the ith stage is conceptually stronger than the ith-1 stage. Rather it is untenable that this development is a result of stronger concepts arising via induction (i.e. learning). Why, because for induction to be possible the conceptually stronger system has to be representable. But to be representable means that the concepts that represent it must be cognitively available (must already be in the hypothesis space). But if so, they cannot enter the hypothesis space by induction as they are already available for induction.  So, development cannot be a matter of CA via learning. 

Does this mean that we development cannot be a matter of stronger concepts being acquired over time? No. But it does mean that this process cannot be inductive (e.g. this scenario is compatible with “maturing” new concepts, just not learning new ones).  Note too, that this is compatible with treating development as a matter of new belief fixation. But recall that BF implies that the relevant concepts are given and available for computational use. Or, as Fodor puts it:

…a theory of conceptual plasticity of organisms must be a theory of how the environment selects among the innately specified concepts. It is not a theory of how you acquire concepts, but a theory of how the environment determines which parts of the conceptual mechanism in principle available to you are in fact exploited. (151)

In other words:

…fixation of belief along the lines of inductive logic…is one that assumes the most radical possible nativism: namely that any engagement of the organism with its environment has to be mediated by the availability of any concept that can eventually turn up in the organism’s belief. The organism is a closed system proposing hypotheses to the world, and the world then chooses among them in terms explicated by some inductive logic. (152)

To repeat, Fodor and Chomsky and virtually all the participants at the Royaumont conference take this to be tautological (as do I). The only theories of learning we have are theories of BF and these theories all presuppose that the stock of possible acquirable concepts is given.  Radical nativism indeed![3]

So far as I can tell, the logic that Fodor and Chomsky outlined well over 30 years ago has not changed. And, if this is correct, then the central problem in cognition, linguistics being a special case, is to adumbrate the relevant hypothesis space for any given domain. And the only way to do this is to investigate the acquirable in terms of the acquired and argue backwards. If BF is the name of the game, then what is presupposed had better suffice to deliver the concepts acquired, and once one looks carefully at what’s on offer, this simple requirement appears to rule out most of the most popular theories, or so Fodor and Chomsky (and I) would argue.

It is worth observing that this tautology was recognized by the great empiricist philosophers.  In this sense, the blank tablet metaphor generally associated with their theories of mind is unhelpful at best and misleading at worst.  The distinguishing mark of empiricism is not that the mind is unstructured (comes with no given hypothesis space) but that the dimensions of the hypothesis space are entirely sensory.  On this view, admissible concepts are either sensory primitive concepts or Boolean combinations of such.  This is a substantive theory, and, as Fodor notes, it has proven to be false.[4] Or as Fodor in his characteristic elegant way puts it:

I consider that the notion that most human concepts are decomposable into a small set of simple concepts –let’s say, a truth function of sensory primitives – has been exploded by two centuries of philosophical and psychological investigation. In my opinion, the failure of the empiricist program of reduction is probably the most important result of those two hundred years in the area of cognition. (268)

As many of you know, Fodor has argued that not only is there no possible reduction to a small number of sensory primitives, but that there is very little possible reduction at all, at least when it comes to our basic lexical concepts. I personally find his arguments against reductions rather strong.[5] However, whether one buys the conclusion, the form of the argument seems to me correct: if you want a restricted set of primitives then you are obliged to show how these can be used to build up what we actually observe. The empiricist restriction to a small base of sensory primitives failed to deliver the goods, therefore, it cannot be correct; it cannot be the case that our basic concepts are restricted to sensory primitives.

So, is nativism ineluctable? Yup.  So what’s the fight between Rationalists and Empiricists about? It’s about two things: (i) the shape of the hypothesis space: what are the primitive projectable predicates and how do they combine to deliver more complex predicates (e.g. what are the basic operations, primitives and principles of UG) (ii) how, given this space, are beliefs fixed (e.g. what is the relation between PLD and G)? Everyone is a nativist when it comes to CA. This is not controversial (or shouldn’t be). Empiricists are nativists that believe in a pretty spare hypothesis space. Rationalists are happy to entertain far more complex ones.  This difference has an impact on how one understands BF. I turn to this in the next post.



[1] The distinction between instructive and selective theories has a long history in the study of the immune system.  Here is a useful summary. Fodor’s point, which seems to me to be entirely correct (and obvious) is that all current theories of learning are selectionist and hence presuppose a fixed innate background of relevant alternatives.
[2] There may be room, in addition, for accounts of how to use incoming data to update the information that guides a learners movements across the given hypothesis space.  What kinds of evidential thresholds are there, how many competing hypothesis does one juggle at once, what are the functions that in/decrease a hypothesis’ credibility “score,” how many “kinds” of evidence are tabulated at once, does the credibility function treat all hypotheses the same or are some more privileged than others, etc.? These are all relevant concerns. But Fodor and Chomsky’s tautological point is that they all are secondary to the issue of what does the hypothesis space look like.
[3] This conclusion is still resisted by many. See for example, Gallistel’s review of Sue Carey’s book here.
            Others also seem to misunderstand the import of this. For example Amy Perfors (here) claims that Fodor’s point is “true but trivial” (132). This is taken to be a critique, but it is exactly Fodor’s point. As Perfors emphasizes: “…any conception which relies on not having a built-in hypothesis space is incoherent…” (128). This is a vigorous rewording of Fodor’s and Chomsky’s point.  It is curious how often one finds strongly worded criticisms of nativist positions followed immediately with these criticized positions offered as novel insights by the very same author, in the very same paper.
[4] Once again some have confused the issues at hand. Perfors (see above) is a good example. The paper contrasts Nativism and Empiricism (127). But if everyone is a nativist with respect to the requirement that for learning to be possible a hypothesis space must be given, then everyone must be a nativist, in Fodor and Chomsky’s sense.  The debate is not over whether we are nativists, but what kind of nativists we are (i.e. how rich a hypothesis space are we willing to tolerate).  The contrast is between Rationalists and Empiricists, the latter limiting the admissible predicates and operations to associationist ones. And, as Fodor notes (see immediately below), this is what’s wrong with Empiricism. It’s not the nativism, but the associationism that makes empiricism a failed program.
[5] I hate to pick on the Perfors paper (well, not really) but it demonstrates how cavalier critics can be when it comes to positions that they consider clearly incorrect.  The paper argues that Fodor’s critique can be finessed by simply understanding that one can have hierarchies of hypothesis spaces (132-3). Thus, contra Fodor, it is possible to treat elements of level N as decomposed of concepts of level N+1 and this gets all that Fodor criticizes but without the unwanted implicational consequences.  Maybe. But oddly the paper never actually illustrates how this might be done. You know, take a concept or two and decompose them and show how they operate to license the wanted inferences and block the unwanted ones. There are lots of concepts around to choose from. Fodor has discussed a bunch. But the Perfors paper does nothing even approaching this. It simply asserts that conceptual hierarchies gets one around Fodor’s arguments.  This is cheap stuff, really cheap. But sadly, all too common.