Comments

Showing posts with label hypothesis space. Show all posts
Showing posts with label hypothesis space. Show all posts

Friday, August 30, 2013

Yay!! I'm not the only one with a Reverand Bayes problem


Bob Berwick is planning to write something erudite about Bayes in a forthcoming post. I cannot do this, for obvious reasons. But I can throw oil on the fires. The following is my reaction to a paper that Ewan suggested that I read the critiques of. I doubt that his advice had the intended consequence. But, as Yogi Berra observed, it’s always hard to accurately predict how things will turn out, especially in the future. So here goes.

In an effort to calm my disquiet about Bayes and his contemporary acolytes, Ewan was kind enough to suggest that I read the comments to this B&BS (is it just me, or does the acronym suggest something about the contents?) target article.  Of course, being a supremely complaisant personality, I immediately did as bid, trolling through the commentaries and even reading the main piece and the response of the authors to the critics’ remarks so that I could appreciate the subtleties of the various parries and thrusts.  Before moving forward, I would love to thank Ewan for this tip.  The paper is a lark, the comments are terrific and, all in all, it’s just the kind of heated debate that warms my cold cynical heart.  Let me provide some of my personal highlights. But, please, read this for yourself. It’s a page-turner, and though I cannot endorse their views due to my incompetence, I would be lying if I denied being comforted by their misgivings. It’s always nice to know you are not alone in the intellectual universe (see here).

The main point that the authors Jones and Love (J&L) make is that, as practiced, there’s not much to current Bayesian analyses, though they are hopeful that this is repairable (in contrast to many of the commentators who believe them to be overly sanguine e.g. see Glymour’s comments. BTW, no fool he!). Indeed, so far as I can tell, they suggest that this is not a surprise for Bayes as such is little more than a pretty simple weighted voting scheme to determine which among a set of given alternatives best fits the data (see L&J’s 3).  There is some brouhaha over this characterization by the law firm of Chater, Goodman, Griffiths, Kemp, Oaksford and Tenebaum (they charge well over $1000 per probable hour I hear), but J&L stick to their guns and characterization (see p. 219) claiming that the sophisticated procedures that Chater et. al. advert to “introduces little added complexity” once the mathematical fog is cleared (219).

So, their view is that Bayesianism per se is pretty weak stuff. Let me explain what I take them to mean. J&L note (section 3 again) that there are two parts to any Bayesian model, the voting/counting scheme and the structure of the hypothesis space. The latter provides the alternatives voted on and a weighting of the votes (some alternatives are given head starts). Now, Bayes’ Rule (BR) is a specification of how votes should be allocated as data comes in.  The hypothesis space is where the real heavy lifting is done. In effect, in J&L’s view (and they are by no means the most extreme voices here as the comment sections show) BR, and modern souped up versions thereof, add very little of explanatory significance to the mix. If so, J&L observe, then most of the psychological interest of Bayesian models resides in the structure of the assumed hypotheses spaces, i.e. whatever interesting results emerge from a Bayesian model, stem not from the counting scheme but the structure of the hypothesis space.  That’s where the empirical meat lies:

All a Bayesian model does is determine which of the patterns or classes of patterns it is endowed with is most consistent with the data it is given. Thus, there is no explanation of where those patterns (i.e. hypotheses) come from. (220)

This is what I meant by saying that, in J&L’s view, Bayes, in and of itself, amounts to little more than the view that “people use past experience to decide what to do or expect in the future” (217). In and of itself Bayes does not specify or bound the class of possible or plausible hypothesis spaces and so in and of itself it fails to make much of a contribution to our understanding of mental life. Rather, in and of itself, Bayesian precepts are anodyne: who doesn’t think that experience matters to our mental life?

This view is, needless to say, heartily contested. Or so it appears on the surface. So Chater et. al. assert that:

By adopting appropriate representations of a problem in terms of random variables and probabilistic dependencies between them, probability theory and its decision theoretic extensions offer a unifying framework for understanding all aspects of cognition that can be properly understood as inference under uncertainty: perception, learning, reasoning, language comprehension and production, social cognition, action planning, and motor control, as well as innumerable real world tasks that require the integration of these capacities. (194)

Wow! Seems really opposed to J&L, right?  Well maybe not. Note the first seven words of the quote that I have conveniently highlighted. Take the right representation of the problem add a dash of BR and out pops all of human psychology.  Hmm. Is this antithetical to J&L’s claims?  Not until we factor out how much of the explanation in these domains comes from the “appropriate representations” and how much from the probability add on.  Nobody (or at least nobody I know) has any problem adding probabilities to mentalist theories, at least not in principle (one always wants to see the payoff). However, if we were to ask where the hard work comes in, J&L argue that it’s in choosing the right hypothesis space and not in probabilizing up a given such space.  Or this is the way it looks to J&L and many many many, if not most, of the other commentators.

Let me mote one more thing before ending. J&L also pick up on something that bothered me in my earlier post. They observe more than a passing resemblance between modern Bayesians and earlier Behaviorism (see their section 4).  They assert that in many cases, the hypotheses that populate Bayesian spaces “are not psychological constructs …but instead reflect characteristics of the environment. The set of hypotheses, together with their prior probabilities, constitute a description of the environment by specifying the likelihood of all possible patterns of empirical observations (e.g. sense data)” (175). J&L go further and claim that in many cases, modern Bayesians are mainly interested in just covering the observed behavior, no matter how it is done. Glymour dubs this “Osiander’s Psychology,” the aim being to “provide a calculus consistent with the observations” and nothing more.  At any rate, it appears that there is a general perception out there that in practice Bayesians have looked to “environmental regularities” rather than “accounts of how information is represented and manipulated in the head” as the correct bases of optimal inference.

Chater et. al. object to this characterization and allow “mental states over which such [Bayesian] computations exist…” (195).  This need not invalidate J&L’s main point however.  The problem with Behaviorism was not merely that it eschewed mental states, but that it endorsed a radical form of associationism. Behaviorism is the natural end point of radical associationism for the view postulates that mental structures largely reflect the properties of environmental regularities. If this is correct, then it is not clear what adding mental representations buys you. Why not go directly from regularities in the environment to regularities in behavior and skip the isomorphic middle man. 

It is worth noting that Chater et. al. seem to endorse a rough vision of this environmentalist project, at least in the domains of vision and language. As they note “Bayesian approaches to vision essentially involve careful analysis of the structure of the visual environment” and “in the context of language acquisition” Bayesians have focused on “how learning depends on the details of the “linguistic environment,” which determines the linguistic structures to be acquired” (195).

Not much talk here of structured hypothesis spaces for vision or language, no mention of Ullman-like Rigidity Principles or principles of UG. Just a nod towards structured environments and how they drive mental processing. Nothing prevents Bayesians from including these, but there seems to be a predisposition to focus on environmental influences.  Why?  Well, if you believe that the overarching “framework” question is how data (i.e. environmental input) moves you around a hypothesis space, then maybe you’ll be more inclined to downplay the role of the structure of that space and highlight how input (environmental input) moves you around it. Indeed, a strong environmentalism will be attractive to you if you believe this. Why? Well, given this assumption then mental structures are just reflections of environmental regularities and, if so, the name of the psychological game will focus on explaining how data is processed to identify these regularities.  No need to worry about the structure of hypothesis spaces for they are simple reflections of environmental regularities, i.e. regularities in the data.

Of course, this is not a logically necessary move. Nothing in Bayes requires that one downplay the importance of hypothesis spaces, but one can see, without too much effort, why these views will live comfortably together. And it seems that Chater et. al., the leading Young Bayesians, have no trouble seeing the utility of structured environments to the Bayesian project. Need I add that this is the source for the expressed unease in my previous post on Bayes (here).

Let me reitterate one more point and then stop.  There is no reason to think that the practice that J&L describe, even if it is accurate, is endemic to Bayesian modeling. It is not. Clearly, it is possible to choose hypothesis spaces that are more psychologically grounded and then investigate the properties of Bayesian models that incorporate these.  However, if the more critical of the commentators are correct (see Glymour, Rehder, Anderson a.o.) then the real problem lies with the fact that Bayesians have hyped their contributions by confusing a useful tool with a theory, and a pretty simple tool at that.  Here are two quotes expressing this:

Rehder [i.e. in his comment, NH] goes as far as to suggest viewing the Bayesian framework as a programming language, in which Bayes’ rule is universal but fairly trivial, and all of the explanatory power lies in the assumed goals and hypotheses. (218)

…that all viable approaches ultimately reduce to Bayesian methods does not imply that Byesian inference encompasses their explanatory contribution. Such an argument is akin to concluding that, because the dynamics of all macroscopic physical systems can be modeled using Newton’s calculus, or because all cognitive models can be programmed in Python, calculus or python constitutes a complete and correct theory of cognition. (217)

So, in conclusion: go read the paper, the commentaries and the replies. It’s loads of fun.  At the very least it comforts me to know that there is a large swath of people out there (some of them prodigiously smart) that have problems not dissimilar to mine with the old Reverend’s modern day followers.  I suspect that were the revolutionary swagger toned down and replaced with the observation that Bayes provides one possibly useful way for exploring how to incorporate probabilities into the mental sciences, nobody would bat an eye.  I’m pretty sure that I wouldn’t. All that we would ask is what one should always ask: what does doing this buy us?

Monday, February 25, 2013

Fodor on Concepts


There has been a bit of a kerfuffle in the thread to What’s Chomsky Thinking Now concerning Fodor’s claim that all of our concepts are innate.  Unfortunately, with the exception of Alex Drummond, most who have participated in the discussion appear unacquainted with Fodor’s argument.  It has two parts, both of which are interesting. To help focus the indignation of his critics, I will outline them below as a public service.  Before starting however, let me share with you my own personal rule of thumb in these matters, one that I learned from Kuhn’s discussions of Aristotelian physics: when someone very smart says something that you think is obviously dumb then go back and reread it for it may be that you have thoroughly misunderstood the point. I am acquainted with Jerry Fodor. Let me assure you he is very smart. So let’s start.

As noted Fodor has a two pronged argument.  The first part (an excellent very short form of which can be found here 143ff) is an observation about learning as a form of inductive logic. Fodor distinguishes between theories of concept acquisition and theories of belief fixation. The latter is what theories of learning are about. Learning theories have nothing general to say about concept acquisition because, being inductive theories, they presuppose the availability of a set of basic concepts without which the inductive learning mechanism cannot get off the ground.  If this all sounds implausible, considering Fodor’s example will make his intent clear.

Consider someone learning a new word miv in a classical learning context.  The subject is shown cards, some of which are miv and some non-miv. The subject is given a cookie whenever s/he correctly identifies the miv cards and is hit by a bolt of lightening when s/he fails (we want the reinforcement here to be very unambiguous).  What does the subject do, according to any classical learning theory, s/he considers a hypothesis of the form “X is miv iff X is…”, the blank being filled in with a specification of the features that are criterial for being a miv.  The data is then used to assess the truth of the hypotheses with various values of “…”.  So if miv means “red and round” then the data will tend to confirm  “X is miv iff X is red and round” and disconfirm everything else. This much Fodor takes to be obvious. If learning is a form of inductive inference (and, as he notes, there is no other theory of learning), then it takes the indicated form. 

Fodor then asks where do the hypotheses that are tested come from? In other words, where do the fillers of “…” come from?  They are GIVEN. Inductive theories presuppose that the set of alternatives that the data filter are provided up front. Given a hypothesis space, the data (environmental input) can be used to assign a number (a probability) of how well that hypothesis fits the data.  What inductive theories don’t do is provide the hypothesis space.  Another way of making the same point is that what inductive logics (i.e. learning theories) do is explain how given some input the user of that logic should/does navigate the hypothesis space: where’s the best place to be given that the data has been such and so.  However, if this is what inductive logics do (and, I cannot repeat this enough, all learning theories are species of inductive logics), then the field of concepts used by the inductive logic cannot themselves be fixed by the inductive logic.  Or as Fodor puts it (147):

You have to be nativistic about the conceptual resources of the organism because the inductive theory of learning simply doesn’t tell you anything about that – it presupposes it – and the inductive theory of learning is the only one we’ve got.

So, Fodor’s argument amounts to pointing out what everybody should be nodding in agreement with: no induction without a hypothesis space.  If the inductive theory is a theory of learning, then the hypothesis space must be innate and that means that all the concepts used to define it must be innate as well.  As I said, this part of the argument is apodictic, cannot be gainsaid and, in fact, never has been. Even Quine, a rather extreme associationist, agreed that everyone is a nativist to some degree for without some nativism (enough to define the hypothesis space) there can be no induction and hence no learning. Fodor’s point is to emphasize this point and use it against theories that suggest that one can bootstrap one’s way form less conceptually complex systems of “knowledge” to more complex ones.  If this means that one can expand one’s hypothesis space by learning and ‘learning’ means induction then this is impossible.[1] 

None of this should be surprising or controversial. Controversy arises with respect to Fodor’s second prong of the argument. He takes the concepts words tag to be effectively atomic. Another way of making this point in the domain of language is that there is no lexical decomposition, or at least very very little. Why is this assumption so important? Because the relation between the input and the atomic features of the hypothesis space is causal, not inductive. You see a red thing and +red lights up. Pure transduction.  Induction proceeds given this first step: count how many of the lit features are red+round vs red+not-round, vs green+round etc.  So, for atomic features/concepts the relation between their “lighting up” and the environment is not one of learning (one doesn’t learn to light them up) it’s just a brute fact (they light up).  So, and this is an important point, to the degree that most of our words denote atomic concepts (i.e. to the degree that there is no lexical decomposition) to that degree there is no interesting inductive theory of concept acquisition. Note, this does not preclude their being a possibly interesting causal theory, e.g. maybe being exposed to a miv is causally responsible for triggering the concept miv or maybe being exposed to a dax is causally responsible, or maybe being exposed to a miv right after birth is or while being snuggled by your mother etc. The causal triggers might conceivably be very complex and finding them may be very difficult. However, with resepct to atomic features, one can only discover brute causal connections, not inductive ones. Fodor’s point is that we should not confuse them as they are very different. Recently Fodor has speculated that prototypes are causally implicated in causally triggering concepts, but he insists, rightly given his strong atomicity, that this relation is not inductive (See here).

To recap, the logic of the first argument is that primitive concepts cannot be “learned” as they are presupposed for learning to take place. This allows the possibility that one “learns” to combine these primitives in various ways and that’s what concept acquisition is.  Concept acquisition is just learning to form complex concepts. Fodor is conceptually happy with this possibility. It is logically possible that concept “acquisition” amounts to defining new concepts in terms of the primitive ones. As applied to words (which I am assuming denote concepts), it is logically possible that most/many words are complex definitions. Logically possible? Yes. Actually the case? No, or that’s what Fodor has been arguing for a very long time.

His arguments are almost always of the same form: someone proposes some complex definition for a term and he shows that it doesn’t work.  Indeed, very few linguists, psychologists or philosopher have managed to provide any but a handful of purported definitions. ‘Bachelor’ may mean unmarried man, but as Putnam noted a long time ago, there are not many words like it.

Fodor is actually in a good position to understand this point for he along with Katz once investigated reducing meanings to feature trees. David Lewis derided this “markerese” approach to semantics (another instance of be careful what you hurl as it may boomerang back at you (see Paul on Harman on Lewis here)), but what really killed it was the realization that virtually all words bottomed out in terminal faetures referring to the very concept that the featural semantics was intended to explicate. So, e.g. the markerese representation for ‘cat’ ended up having a terminal CAT. This clearly did not move explanation forward, as Fodor realized.

So is Fodor right about definitions?  Well, I am slightly less skeptical than he is about the virtues of decomposition, however, this said, I cannot find good examples showing him to be wrong. As the first part of his argument is unassailable, then those that don’t like the conclusion that ‘carburetor’ is innate (i.e. a primitive of our conceptual hypothesis space) had better start looking for ways of defining these words in terms of the available primitives.  If past history is any guide, they will fail.  Definitions in terms of sense data have come and (happily) gone and cluster concepts, once considered seriously, have long been abandoned. There is a little industry in linguistics working on argument structure in the Hale-Keyser (HK) framework, but, at least from where I sit, Fodor has drawn significant blood in his debates with HK aficionados. Suffice it for now to repeat, that this is where the action must be if Fodor is to be proven incorrect and the ball is clearly not in his court.  It is easy to show that he is wrong, viz. show that most/many words denote complex concepts.  How to show Fodor is wrong is easy. Showing that he is has proven to be far more challenging.[2]

So that’s the argument. The first step is clearly correct. All the action concerns the second.  One further point: there has been a lot of discussion in the thread that Fodor is advocating a nutty kind of nativism that eschews learning from the environment. As should be clear, this is simply false. If word learning is belief fixation then it can be as inductivist as you like. However, if word learning is concept acquisition then the question entirely revolves around the nature of the primitives concerning which everyone must take as innate and hence not acquired. Fodor’s bottom line is that hypothesis spaces are not acquired but presupposed and that as a matter of fact there is far less definition one might have supposed. That’s the argument; fire away!


[1] Alex Clark mentioned Sue Carey’s recent book that appeared to consider this bootstrapping possibility. Gallistel reviewed her book making effectively this point that induction/learning cannot expand a hypothesis space (here). To repeat, all that such theories show is how to most effectively navigate this space given certain data.
[2] One interesting avenue that Paul has been exploring revolves around Frege’s notion of definition.  For Frege definition changed a person’s cognitive powers. This is really interesting. Paul’s work starts from Jeff Horty’s discussion of Frege’s notion (here and considers how to extend it to theories of meaning more generally (c.f. here and here).