Comments

Showing posts with label small world situations. Show all posts
Showing posts with label small world situations. Show all posts

Thursday, April 7, 2016

Yang on Bayes 2

Here is part 2. I have included the previous paragraph so that you can get a running start into the discussion. Let me remind you that Charles is culpable for the content of what follows only insofar as reading his stuff stimulated me to think as I did. In other words, very culpable.

It is worth noting that Bayes makes these substantive assumptions for principled reasons. Bayes’ started out as a normative theory of rationality. Bayes was developed as a formalization of the notion “inference to the best explanation.” In this context the big hypothesis space, full data, full updating, optimizing rule structure noted above are reasonable for they undergird a the following principle of rationality: choose that theory which among all possible theories best matches the full range of possible data. This makes sense as a normative principle of inductive rationality. The only caveat is that Bayes seems to be too demanding for humans. There are considerable computational costs to Bayesian theories (again CY notes these) and it is unclear how good a principle of rationality is if humans cannot apply it. However, whatever its virtues as a normative theory, it is of dubious value in what L.J. Savage (one the founders of modern Bayesianism) termed “small world” situations (SWS).

What are SWSs? They are situations where the hypothesis space, the space of options over which the probabilities are defined, is small. In SWSs the central problem is to describe the structure of the “small world.” All else is secondary. So in SWSs useful idealizations will help focus on these structures. And that is the problem with Bayes. As Pat Suppes notes (CY quotes from this paper p. 47, my emphasis NH)):

…any theory of complex problem solving cannot go far simply on the basis of Bayesian decision notions of information processing. The core of the problem is that of developing an adequate psychological theory to describe, analyze and predict the structure imposed by organisms on the bewildering complexities of possible alternatives facing them. The simple concept of an a priori distribution over these alternatives is by no means sufficient and does little toward offering a solution to any complex problem.

…understanding the structures actually used is important for an adequate descriptive theory of behavior….As…inductive logic comes to grips with more realistic problems, the overwhelming combinatorial possibilities that arise in any complex problem will make the need for higher-order structural assumptions self-evident.

In short, the hard problem is to find the structure of the SWSs (to describe the restricted hypothesis space) and Bayes brings little (nothing?) of relevance to the table for solving this problem. Bayes does not prevent considerations of this problem, but it does not promote them or make them central concerns either. And I would argue that by not recognizing (or, worse, radically deemphasizing) the SWS character of most psychological processes, Bayes deflects attention from the hard problem, the one that must be solved if there is to be any progress, onto secondary concerns. That’s the standard cost of a misidealization; it leads you to look in the wrong place.

Let me put this another way, echoing Suppes and CY. The hard problem is “the overwhelming combinatorial possibilities that arise in any complex problem.” The Bayes idealization abstracts away from this at every step: its default assumptions are big hypothesis spaces, large data, panoramic update and optimization. Each of these causes problems. If the cognitive reality (at least in language, but I suspect everywhere) is small worlds, very limited data, myopic (and dumb) updating and satisficing decision then Bayes is just misleading. Moreover, starting where Bayes does means that the bulk of the work will not be done by the Bayes part of any account, but by the algorithms that massage these assumptions out. And indeed, this is what we find when Bayesians respond to their critics. Here’s an example of this move that CY discusses.

CY notes that when confronted with problems Bayesians usually go Marrian. So, for example, when it is observed that Bayes is inconsistent with probability matching, the retort is that this is only so for the level 1 theory. The level 2 algorithms allow for probability matching when certain further constraints are added to the optimizing function. However, if the critique of Bayes is that the idealization is the wrong one, then the fact that one can patch up the problem by saying more is not so much a sign of health but possible confirmation of the initial wrong turn. Of course you can patch anything up with further machinery. This is never in doubt (or almost never). The question is not whether this is possible, but whether the patching up is consistent with the spirit of the initial idealization. If it is, then no problem. If it isn’t, then it argues against the idealization. CY argues that how Bayes handles probability matching goes against the Bayesian spirit of the initial idealization (see the CY discussion of Steyvers et al p 12).

There are many critiques of Bayes that go along the same lines. Many note that Bayes stories have an ad hoc flavor to them (see here and references in linked paper). One critic, Clark Glymour (who, BTW, is no schnook (see here)), has described the work as “Ptolemaic” (see here) and not in a good way (is there a good way?). In the cases discussed, the ad hoc flavor comes from specific assumptions added to the basic Bayes account to accommodate the recalcitrant data, or so the critics argue. It is above my pay grade to evaluate whether the critics are correct in each case. However, this is what we expect if the basic Bayes account is based on a misidealization of the relevant problems.

In fact, this is precisely how we now understand the failure of Ptolemaic astronomy. Think epicycles. Why the epicyles? Because the whole Ptolemaic project starts from two completely wrong idealizations; that the universe is geocentric and that orbits are circular. Starting here, the astronomical data demands epicycles and equants and all the other geometric paraphernalia. Change to a heliocentric system and allow elliptical orbits and all this apparatus disappears. The ad hoc apparatus is inevitable given the initial idealization of the problem.  Given the starting assumptions it is possible to “model” all planetary motion (in fact, we now know that any possible set of orbits can me modeled). However, it is equally clear that in this case tracking the data is not a virtue. Wrong idealization, wrong theory no matter how well it can be made to fit the data.

Let me end this (again) overly long post, with a few observations.

First, arguing against an idealization is very difficult. Why? Because it is often possible to “fix” the problems that a misidealization generates by adding further bells and whistles. Thus, it is not generally possible to argue against an idealization empirically. Rather the argument is necessarily more subtle: the accounts delivered are non explanatory, they involve too much ad hoc machinery, they are too complex, they don’t fit well with other things we want etc. Despite their difficulty, critiques of idealizations are one of the most important kinds of scientific arguments. They are hardly ever dispositive, but they are extremely important precisely because they expose the basic lay of the land and identify the hard problems that need solutions. They are also the kinds of arguments that can make you a scientific immortal (think Newton on contact mechanics and Einstein in the aether).

Second, CY presents an interesting Marr like argument against Bayes. It proposes that level 1 theories that have transparent relations to level 2 accounts are better than those that don’t (see the prior three posts on Marr for some discussion). CY argues that Bayes level 1 and 2 theories are conceptually caliginous because of the initial misidealizaation. Starting from the wrong place will make it hard to transparently relate level 1 computational theories with level 2 theories of algorithms and representations. CY makes such an argument as follows.

It is well known that Bayes procedures taken transparently are computationally intractable. CY observes that to date, even the non-transparent ones that have been proposed are pretty bad (i.e. very slow). This is not a surprise in a Marrian setting if the idealization pursued by Bayes is wrongheaded. Indeed, in a Marr setting one might argue that not being able to generate “nice” algorithms (transparent ones being particularly nice) is a sign that the computational level theory is heading in the wrong direction. The fact that CY’s proposed algorithm is both tractable and easily implementable and that Bayesian proposals are not is a strong argument against the adequacy of the latter idealization in a Marr setting given the Marrian desideratum that computational level theories inform and constrain level 2 algorithmic theories.

CY’s argument goes further. CY’s proposed theory provides a dynamics for acquisition: it not only can tell you how the data determines the G you end up with, but it also tells you how to traverse the set of alternatives available as the data comes in. So, the theory tells you how hypotheses change over real time. Bayes does not do this. It just specifies where you end up given the data available. Bayes is iterative, but not incremental. CY’s theory is both.

Interestingly, CY derives this dynamic incrementality via a transparency assumption regarding the Elsewhere Principle. This is something that a Bayes account cannot possibly achieve precisely because it is well known to be intractable when understood transparently. In other words, CY can do more than Bayes because its understanding of the level 1 problem allows for a very neat/transparent mapping to a level 2 account. This is unavailable for Bayes because the transparency assumption in a Bayes theory leads to intractable computations. So, wrong idealization, less transparency and so no obvious dynamics.

Third, the Bayes idealization embodies a confusion linguists should be familiar with. Remember the child as little linguist hypothesis?[1] It’s the idea that the way to think of how LADs acquired Gs is on a par with how linguists discover the properties of particular Gs. This, we know, is wrong on at least two grounds: first, linguists when they operate consider a wide range of theoretical alternatives for any given "problem" and second linguistic consider full linguistic data to solve their problems while children only consider a small subset of potentially useful data, the PLD. Thus, in two important ways, children inhabit a “smaller world” than linguists do and consider much scantier data in reaching their conclusions than linguists do.

Bayes and the child-as-little-linguist hypothesis make the same mistake. LADs are entirely unlike linguists, and a good thing too. The difference is that the LADs hypothesis space of possible Gs given PLD is very small. In other words, they inhabit a small linguistic world. Linguists qua scientists do not. The aim of linguistics is to figure out what this small world looks like. Sadly, linguists’ access to this small world is not aided by the same built in advantages that the LAD comes equipped with, so it’s standard inference to the best explanation for GGers. That’s why for us the process is slow and laborious while kids acquire their Gs roughly by the age of 5.

So, both Bayes and the child-as-little-linguist hypothesis mis-describe the main acquisition problem by assimilating it to a rational decision problem thereby abstracting away from the critical features of situation: rich small hypothesis space, sparse, misleading and uninformative data, and, quite likely, no optimizing (see CY p 23 on the principle of sufficiency).

Let me put this one more way: the inferences that LADs make in acquiring their Gs are not rational. They consider too few possible hypotheses and utilize very little data in getting to the G they choose. True LADs use data and choose among Gs but the problem is not really like an ideal inductive procedure. Moreover, the ways that it is not like an ideal inductive procedure are what allow the whole procedure to successfully apply. It is precisely by radically narrowing the hypothesis space that it is possible for the LAD to choose the “right” G despite lousy data. The process, it seems, need not be rational to be effective. Just the opposite one might say.

Fourth, CY has a very good discussion of the likely irrelevance of the subset principle to acquisition. This is important (and maybe I will return to it sometime in the future) because the fact that Bayes can derive the subset principle is taken to be a feather in its cap. Roughly speaking, if Bayes then subset principle and because subset principle therefore good for Bayes. But, CY argues that there is good reason to think that the subset principle is not operataive in acquisition. If so, no argument for Bayes.

Note that this is interesting regardless of whether Bayes is right. The subset principle is often invoked so seeing how problematic it is has its own charms. BTW, one of the big problems with the subset principle concerns its tractability. It is computationally very complex as CY notes. If so, if Bayes derives it, it might not be good news for Bayes.

Ok, enough. Run and get CY. It’s an excellent paper and very important. If nothing else, I hope that it focuses the debate over Bayes. The problem is not the details, but the idealization. The Bayes stance runs against what linguists believe to be the basic lay of the cognitive land. If we are right, then they are likely wrong. This is a place where the different starting points are worth sharpening. CY shows us how to do this, and that’s why it is such an important paper.



[1] See the following: Valian, V., Winzemer, J., & Erreich, A. (1981). A little linguist model of syntax learning. Language acquisition and linguistic theory, 188-209.

Tuesday, April 5, 2016

Yang on Bayes 1

This is the first of two posts on a recent paper by Charles Yang. Once again, the topic got away from me. I break it down into two parts to prevent you running away too scared to even look. I suspect that my clever maneuver won’t help. But I try.

One more caveat: this is my understanding of Charles’ paper. He should only be held responsible for what I say because he allowed me to read it, and that might be a culpable act.

Charles Yang has a new paper (CY) forthcoming in Language Acquisition and it is a must read. Aside from sketching a very powerful critique of contemporary Bayesianism as applied to linguistic problems (I will return to this critique momentarily), it also does something that I never thought I would witness in my lifetime; it makes quantitative predictions in a linguistic domain. And by “quantitative” I do not mean giving p-values or confidence intervals. I mean numerical predictions about the size of measurable effect. That’s quantitative! So, run, don’t walk to the above link and read the damn thing!

Ok, now that I’ve discharged my kudosing responsibilities, I want to discuss a second feature of CY. It offers an excellent critique of current Bayesian practice as applied to linguistic issues. As you all know, Bayes is a big player nowadays. In fact, I am tempted to say that Bayes has stepped into the position that Connectionism once held as the default theoretical framework in the cognitive sciences in general and the cognition of language in particular. As you may also recall, I have expressed reservations about Bayes and what it brings to the linguistics table (see here, here, here, here). However, it was not until I read CY that I clearly understood what bugs me about the Bayes framework as applied to my little domain of interests. I want to “share” this aha moment with you.

The conclusion can be put briskly as follows: Bayes is not so much wrong as wrong-headed. From a Generative Grammar (GG) perspective, it’s not the details that are off the mark (though they can be) but the idealization implicit in the framework that is (i.e. if you accept as I do that GG has correctly identified the “computational” problems (see here) then Bayes is of little relevance and maybe worse). And because this is so, there is very little we can learn from Bayesian modelings of linguistic problems. There is both (i) not enough there there and (ii) what there there is points in the wrong direction.

Let me put this point another way: all theories idealize. Empirically successful theories built on good idealizations.[1] CY’s argument is that Bayes is a bad idealization when applied to matters linguistic. If it is right (and, surprise surprise, I believe it is) then Bayes not only adds little, it positively misleads and misdirects. Why? Because bad idealizations cannot be empirically redeemed. It’s ok to be wrong (you can build on error given a good framing of the problem). But if a theory is wrongheaded it will impede progress.  Whereas an adequate conception of the problem (which is what idealizations embody) tolerates empirical missteps. You can’t data your way out of a misconceived idealization. Why not? Because if you’ve got the problem wrong then data coverage will come with a big price, ad hoc assumptions whose main purpose is to make up for the basic misconception. Hence, a misframing of the problem leads to explanatory sterility and misplaced efforts. That’s the claim. Let’s see how Bayes, misidealizes when considered form a GG perspective.

A Bayes model consists of 3 moving parts: (i) a hypothesis space delimiting the range of options (possibly weighted), (ii) a specification of the data relevant to the choosing of the right hypothesis, (iii) an update rule saying how to evaluate a given hypothesis given certain evidence. This 3-step procedure can iterate so that data can be evaluated incrementally.[2] What makes a Bayes account distinctive is not this 3-step process. After all, what do (i-iii) say beyond the uncontroversial truism that data is relevant to hypothesis acceptance?  No, what makes the Bayes picture distinctive is how it conceives of the hypothesis space, the relevant data and the update rule. This is what gives Bayes content and this CY argues is where Bayes misfires. It misidentifies the size of the hypothesis space, the amount of relevant data and the nature of the update rule. How exactly? As follows:

a.     Bayes misidentifies the size of the hypothesis space. In particular, Bayes takes as the default assumption that the hypothesis space is large. This makes the basic cognitive problem one of picking out the right hypothesis from a large number of (wrong) possibilities. This is a mis-idealization if within linguistic domains (e.g. acquisition) only a (very) small number of candidates are ever being evaluated wrt the input at any one time.  If the candidate set is small, then the “hard” problem is providing the right characterization of the restricted hypothesis space, not figuring out how to find the right theory in a large space of alternatives.[3]

b.     Bayes mischaracterizes the nature of the data exploited. In particular, Bayes idealizes the learning problem as trying to figure out how to use lots of complex information to find the right hypothesis in a large space. However, CY notes, G acquisition proceeds by using only a small part of the "relevant" data at any given time and there is not much of it. More specifically, the PLD that the child uses is severely restricted (sparse, degenerate and inadequate). Thus, in contrast to the full range of linguistic data that linguists use to find a native speaker’s G, kids only use a severely restricted data set when making zeroing in on their Gs. Thus, for the child, the problem is not finding a G in a large haystack of Gs using a very big pitchfork (that’s the linguist’s problem), but is more like skewering one G from among a small number of scattered Gs using a fragile toothpick (the PLD being multiply deficient). If this is correct, then the hard problem is again finding the structure of the G space so that pretty “weak” evidence (roughly sparsely scattered examples of main clause phenomena)[4] suffices to fix the right G for the LAD.

c.     Bayes uses all the data to evaluate all the hypotheses at every iterative step. The way that Bayes models work is that every hypothesis in the space is evaluated wrt all of the data at any given point (i.e. cross situational learning). So, add new data and we update every hypothesis wrt that data. We might describe this as follows: the procedure takes a panoramic view of the updating function. Contrast this to a procedure where only one theory (or two) is ever seriously being considered at any one step and that alternatives are considered only when the favored one(s) fails (see here). On this second view, evaluation of hypotheses is severely myopic with virtually all but a very few number of alternatives ever being considered. Moreover, the myopia might be yet more severe: the alternatives are not chosen among the most highly valued alternatives but randomly. So, if H1 fails then the procedure does not opt for H2 because it is the next best theory to that point, but the rule is to just pick another hypothesis at random. So, not only is the procedure myopic, it is pretty dumb. Blind and dumb and thus not very rational.

d.     Bayes misunderstands the decision rule to be an optimizing function.  A symptom of this is misunderstanding is the problem Bayes has in accounting for the widespread phenomenon of probability matching (PM). Indeed, Bayes doesn’t explain PM. It explains it away. Why? Because Bayes alone cannot explain it. Bayesian inference is inconsistent with PM. Left to its own devices, Bayes implies that agents will select the option that maximizes the posterior rather than split the difference probabilistically between many different options. But the latter is what PM does (PM implies splitting the difference in accord with the probability of the options). To deal with this (and PM is pervasive), Bayes adds further assumptions (some might argue ad hoc assumptions (see below)) to allow the maximizing rule to result in probability matching. If so, this suggests that the Bayesian idealization without further supplementation points one in the wrong direction (see here for discussion). Much preferred would be an update rule that allows PM. Such rules exist and CY discusses them.

It is worth noting that Bayes makes these substantive assumptions for principled reasons. Bayes’ started out as a normative theory of rationality. Bayes was developed as a formalization of the notion “inference to the best explanation.” In this context the big hypothesis space, full data, full updating, optimizing rule structure noted above are reasonable for they undergird a the following principle of rationality: choose that theory which among all possible theories best matches the full range of possible data. This makes sense as a normative principle of inductive rationality. The only caveat is that Bayes seems to be too demanding for humans. There are considerable computational costs to Bayesian theories (again CY notes these) and it is unclear how good a principle of rationality is if humans cannot apply it. However, whatever its virtues as a normative theory, it is of dubious value in what L.J. Savage (one the founders of modern Bayesianism) termed “small world” situations (SWS).




[1] For example, the view that there is an infinite number of natural language objects is an idealization. It is certainly false that humans can deal manage sentences with, say, 40 levels of embedding. So there is some upper bound on the number of sentences native speakers can effectively manage. Thus we have no behavioral evidence that human linguistic capacity is in fact infinite in the required sense. However, this is not really relevant to the utility of the idealization. What matters is whether it makes the central problem vivid (and tractable) and it does. The basic fact about human linguistic competence is that we can use and understand sentences never before encountered. The question is how we extend beyond what we have been exposed to do this (i.e. the projection problem). It matters not a whit if the domain of our competence extends to an infinite set of objects or to just a very very very large one. Or even small one for that matter. What matters is that we project beyond what we have heard to sentences original to us. If we can do this, we need rules. The infinity assumption vividly highlights the need for rules or Gs in any account of linguistic competence. Thus, in that sense, it is a fine idealization even if, perhaps, false.

I should add, that I am not sure that it is false in the relevant sense. After all, humans have numerical competence even though I doubt that we do well with integers 100,000,000 digits long. My point is that even if false, the idealization is an excellent one for it highlights the problem that needs solving and that whatever solution we come up for making this assumption will extend naturally to the more realistic one. Solving the projection problem in the “larger” domain will serve to solve it in the smaller.
[2] Note, that a procedure is iterative does not mean that it is incremental in the interesting sense that we want from our acquisition models. Incremental in the latter means that the more data we get then the closer we get to the true theory. It means that as we get more data we don’t bounce around the hypothesis space from G to G. Rather as we get more data we smoothly hone in on the right G. It is possible for a theory to be iterative without being incremental. Dresher and Kaye have good discussions of this and the assumption (idealized) that there is instantaneous learning within parameter setting models. The latter idealization makes sense if in fact there is no smooth relation between more data and closing in on the right G (though this is not its only virtue). As we would like our theories to be incremental figuring out how to make this so in, for example, a parameter setting model, is a very interesting question (and the one that Dresher and Kaye focus on).
[3] Moreover, finding the right answer in a restricted space need not be identical to finding the right answer in a large space. Finding your keys on your desk (notice I said “your” not “my”) is not obviously the same problem as finding a needle in a haystack.
[4] CY has an excellent discussion of just how sparse things can be. Standard PoS arguments note how deficient it is relative to much of the knowledge attained.