Comments

Showing posts with label Bayesian. Show all posts
Showing posts with label Bayesian. Show all posts

Tuesday, November 19, 2013

Bayesian claims?

Last week I had two Bayes moments: the first was a vigorous discussion of the Jones and Love paper with computational colleagues with a CS mind set and the second was a paean delivered by Stan Dehaene in his truly excellent Baggett Lectures this year (here). These conversations were quite different in interesting ways. 

My computational colleagues, if I understood them correctly (and this is a very open question) saw Bayes as providing a general formal framework useful to the type of problems cognitive neuroscientists encounter. Importantly, on this conception, Bayes is not an empirical hypothesis about how minds/brains compute but an empirically neutral notation whose principle virtue is that allows you to state matters more precisely than you can in simple English. How so? Well minds/brains are complicated and having a formal technology that can track this complexity in some kind of normal form (as Bayes does) is useful. On this view, Bayes has all the virtues (and charms?) of double entry bookkeeping.[1] On this view, every problem is amenable to a Bayesian analysis of some kind so there is no way for a Bayes approach as such to be wrong or inapposite, though particular proposals can be better or worse. In short, Bayes so considered is more akin to C++ (an empirically neutral programming language) than to Newton’s mechanics (a description of real world forces).

There is a contrasting view of Bayes that is prevalent in some parts of the cog-neuro world. Here Bayes is taken to have empirical content. In contrast to the first view, its methods can be inappropriate for a given problem and a Bayes approach as such can even be wrong if it can be shown that its design requirements are not met in a given problem domain.  On this view, Bayes is understood to be a description of cog-neuro mechanisms and describing a system as Bayesian is to attribute to that system certain distinctive properties.

These two conceptions cannot be more different. And I suspect that the move between these two conceptions has muddied the conceptual landscape.  Empirically minded cognitive neuroscientists are attracted to the second conception for it commits empirical hostages and makes serious claims (I review some of these below). CS types seem attracted to the first conception precisely because it provides a general notation for dealing with any computational problem and it does so by being general and thus bereft of any interesting empirical content.[2] Whichever perspective you adopt, it’s worth keeping them separate.

I just read a very good exposition of the second conception of Bayes-as-mechanism by O’Reilly, Jbabdi and Behrens (OJB) (here).[3] OJB is at pains (i) to show what the basic characteristics of a Bayes system are, (ii) to illustrate successful cases where the properties so identified gain explanatory purchase within cog-neuro and (iii) to show where the identified properties are not useful in explaining what is going on. Step (iii) illustrates that OJB takes Bayes to making a serious empirical claim.  Here are some details, though I recommend that you read the paper in its entirety as it offers a very good exposition of why neuroscientists have become interested in Bayes, and it’s not just because of (or even mainly due to) its computational explicitness.

OJB identifies the following characteristics of a Bayesian system (BS):

1.     BSs represent quantities in terms of probability density functions (PDFs). These represent an observer’s uncertainty about a quantity (1169).
2.     BSs integrate information using precision weighting (1170). This means that the information is combined sensitive to their relative reliability (as measured probabilistically): “It is a core feature of Bayesian systems that when sources of information are combined, they are weighted by their relative reliability (1171).”
3.     BSs integrate new information and prior information according to their relative precisions. This is analogous to what occurs when several sources of new information are combined (e.g. visual and haptic info). As OJB puts it: “the combined estimate is partway between the current information and the prior, with the exact position depending on the relative precision of the current estimate and the prior (1171).”
4.     BSs represent the values of all parameters in the model jointly, viz. “It is a central characteristic of fully Bayesian models that they represent the full state space (i.e. the full joint probability distribution across all parameters) (1171).”

In sum: a BS represents info as PDFs, does precision weighted integration of old and new info, and fully represents and updates all info in the state space. If this is what a BS is, then there are several obvious ways that a given proposal can be non-BS. A useful feature of OJB is that it contrasts each of these definitional features with a non-BS alternative. Here are some ways that a proposed system can be non-BS.

·      It represents quantities as exact (contra 1). For example, in the MSTG paper here, learners represented their word knowledge with no apparent measure of uncertainty in the estimate.
·      It combines info without sensitivity to its reliability (contra 2). Thus, e.g. in the MSTG paper information is not combined probabilistically by considering the weighting of the prior and input. Rather contrary info leads one to drop the old info completely and arbitrarily choose a new candidate.
·      It uses a truncated parameter spaces or does not compute values for all alternatives in the space. Again, in MSTG paper the relevant word-meaning alternatives are not updated at all as only one candidate at a time is attended to.

The empirical Bayes question then is pretty straightforward conceptually: to what degree do various systems act in a BS manner and when a system deviates along one of the three dimensions of interest, how serious a deviation is it? Once again, OJB offer useful illustrations. For example, OJB notes that multi-sensory integration looks very BSish. It represents incoming info PDFly and does precision weighted integration of the various inputs. Well almost. It seems that some modalities might be weighted more than “is optimal” (1172). However, by and large, the BS model the central features of how this works. Thus, in this case, the BS model is a reasonable description of the relevant mechanism.

There are other cases where the BS idealization is far less successful. For example, it is well known that “adding parameters to a model (more dimensions to the model) increases the size of the state space, and the computing power required to represent and update it, exponentially (1171).” Apparently problems arise even when there are only “a handful of dimensions of state spaces” (1175). Therefore, in many cases, it seems that behavior is better described by “semi-Bayesian models,” (viz. with truncated state spaces) or “non-Bayesian models” (viz. in which some parameters are updated and some ignored) (1175).  Or models in which “variance-blind heuristics” substitute for precision weighted integration or “rather than optimizing learning by integrating information over several trials, participants seem to use only one previous exemplar of each category to determine its ‘mean’ (1175).”

OJB describe various other scenarios of interest all with the same aim: to show how to take Bayes seriously as a substantive piece of cog-neuro science. It is precisely because not everything is Bayesian that arguing that a mechanism is Bayes that one might get some explanatory insight from the classification. OJB take Bayes to be a useful description for some class of mechanisms, ones with the three basic characteristics noted above: PDF representations, precision weighted integration and fully jointly specified state space.

OJB points out one further positive feature of taking Bayes in this way: you can start looking for neural mechanisms that execute these functions, e.g. how neurons or populations of neurons might allow for PDF like representations and their integrations. In other words, a mechanistic interpretation of Bayes leads to an understandable research program, one with real potential empirical reach.

Let me end here. I understand the OMB version of Bayes. I am not sure how much it describes linguistic phenomena, but that is an empirical question, and not one that I will be well placed to adjudicate. However, it is understandable. What is less understandable are versions that do not treat Bayes as a hypothesis about mental/neural mechanisms. If Bayes is not this, then why should we care? Indeed, what possible reason can there be in not taking Bayes in this mechanistic way?  Is it the fear that it might be wrong so construed?  Is being wrong so bad?  Isn’t it the aim of cog-neuro to develop and examine theories that could be wrong? So, my question to non-mechanistic Bayesians: what’s the value added?





[1] This sounds more dismissive than it should perhaps. The invention of double entry bookkeeping was a real big deal and if Bayes serves a similar function, then it is nothing to sneeze at. However, if this is what practitioners take its main contribution to be, they should let us know in so many words.
[2] I believe that one can see an interesting version of these contrasting views in the give and take between Alex C and John Pate in the comments section here.
[3] The paper (How can a Bayesian approach inform neuroscience) is behind a paywall. Sorry. If you are affiliated with a university you should be able to get to it pretty easily as it is a Wiley publication. 

Sunday, November 10, 2013

Computational Linguistics: Too Computational for Linguistics?

Even though I have recently moved on from the bed of nails that is the current job market to a cushy tenure track job, I still find myself reading the job announcements on LinguistList on a daily basis. There's of course all kinds of professional reasons for doing so, but the actual driving force behind this minor obsession of mine is more twisted. For you see, I have an existentialist streak that allows me to derive perverse amounts of joy from things that should cause me grief, worry, pain, and outbreaks of homicidal rage. And job searches for computational linguists got all of that aplenty.

Monday, August 5, 2013

My Problems with Reverend Bayes


The title of this post is doubly misleading. First, I have no problem with the Reverend. He’s never treated me badly, most likely because he’s been dead for some time. It’s the modern day application of his rule that disturbs me. Second, my “problem” is more a discomfiture than a full-blown ache. I don’t know enough (though I wish I did, really) to be distressed. However, uncomfortable I am (say this as Yoda would), and have decided to post about this unease in the hope that a kindly Bayesian will take pity on me and put my worries to rest. So, what follows tries to articulate the queasy feelings I have about the Bayesian attitude towards cognitive problems, especially in the domain of language, and I put them here, not because I am confident in the following remarks, but because I want to see if the way I am thinking about these matters makes any sense.[1]  So let’s start.

First, these divagations have been prompted by a nice comment by Mark Johnson and reply by Charles Yang and comment by Aaron White here. The comments lead me to think about my worry in the following way. Be warned, it takes a little time to get there.

As you know, I am obsessed with the contrast between rationalist (R) and empiricist (E) conceptions of mind. This contrast has often, IMO, been misunderstood as a contrast between nativist and non-nativist conceptions. However, this cannot be correct, for any theory of acquisition, even the most empiricist ones requires natively given structures to operate, i.e. every theory of “learning” needs a given (i.e. native) hypothesis space against which data provided by the environment is evaluated.  So a better way of marking the R/E contrast is in how articulate the hypothesis space in any given domain is. Es take as their 0th assumption that the space is pretty wide and fairly flat and that learning, (i.e. the sophisticated use of environmental information) guides navigation through the space. Rs take as their 0th assumption that cognitive spaces are pretty narrow and highly structured and that environmental influences, though not unimportant, play a secondary role in explaining why/how acquirers get from where they start to where they end up.  If there are lots of roads from A to B then finding the best one can be a very complicated task. If there is just one or two, choosing correctly is not nearly as complicated. This I take to be the main R/E difference: both concede the importance of native structure and both leave a role for environmental input. The difference lies in the relative weight each assigns to these different factors. Es bet that the action lies with good ways of evaluating the environmentally provided information. Rs lay their money on finding the narrow set of articulated options.  With this as background, here’s my problem with Reverend Bayes' descendants.[2]

I believe that Bayesian methods generally favor the first conception of the acquisition problem. I am pretty sure that this is not required, i.e. there is nothing in Bayesianisn per se that requires this. However, one reason to use fancy counting methods (and they can be fancy indeed (think Dirchlet)) is the belief that how one counts is largely causally responsible for where one ends up. In other words, Bayesian methods are interesting to the degree that the set of options is wide. If so the real trick is to figure out how to efficiently navigate this space in response to environmentally provided information. Thus, Bayesian affinities resonate harmoniously with E-like background assumptions. Consequently, if one believes like I do that a good deal (most?) of the interesting causal action lies with constrained and articulated shape of the hypothesis space then one will look on Bayesian predilections with some suspicion. Put more diplomatically, if there is a trade off between how tight and structured the space of options is and how complex and sophisticated the learning procedure is then Rs and Es will place their research bets in different places even if both features (i.e. the shape of the space and the nature of the learning theory) are agreed to be important.

If this is so, then the problem with Bayesiansism from my Rish point of view is that it presupposes an answer to (and hence begs) the fundamental question of interest: how structured is the mind? 

You can get a good taste of this from Perfors et. al.’s paper (here). What it shows is that if one starts with three grammatical options, a linear grammar, a simple right branching grammar and a phrase structure grammar (PSG), then there is information in the linguistic input that would favor choosing PSGs and that an (ideal) Bayesian acquisition device could use this information to converge on this grammar.  This conclusion gets used to argue that linguistic minds need not specify that the choice of hierarchical grammars by LADs (viz. PSGs) is “innate” for a Bayesian learning mechanism suffices to reach this same conclusion without ruling out the other options nativistically. Putting aside whether anyone argued for what Perfors et. al. argued against (see here for a very critical discussion), what is useful for present purposes is that Perfors et. al. illustrate how Bayesians trade Bayesian learning for restrictive hypothesis spaces. Indeed, in my experience (limited though it is) I’ve noticed that one thing Bayesians never mind doing is throwing in another option into the space of possibilities, secure in the knowledge that, in the limit, the Bayesian learner will get to the right one no matter how big the space is.

To repeat something I said earlier, it is possible that Bayesian methods of counting have advantages even in highly structured hypothesis spaces of the kind that syntacticians like me are predisposed to think exist in the domain of language. However, this is what needs showing, in my view. However, and this lies behind some of my worry, one can view the SYTG paper discussed in a prior post (here) as arguing that in such a context such methods are not at all helpful. Indeed, they point in the wrong direction. Berwick gave similar arguments at the last LSA for morphological acquisition. In both kinds of cases, it seems that in the more restricted hypothesis domains that they consider much simpler counting procedures do a whole lot better, and for principled reasons. So, even though there is nothing incompatible between the Reverend’s rule and structured minds, there is an affinity between Eism and Bayesianism that shows up both in theory (how the computational problem is posed) and practice (in concrete proposals for dealing with particular problems).

So, why do the Reverend’s acolytes bother me? Reason/feeling number 1: because their conception of acquisition is largely environmentally driven and my Rish sensibilities based as they are in what happens in the domain of language leads me to think that this is wrong. And not just a little wrong, but deeply wrong. Wrong in conception, not wrong in execution. Or, to put this in Marr’s terms, Bayesians in their E-nishness have misconstrued the computational problem to be solved. It’s not how do we use environmental input to navigate a big flat space but how do we use such data to make relatively simple structured choices.

The Marr segue above leads to a second (and largely secondary) source of my unease. Bayesians often describe their theories as Level 1 computational theories in Marr’s sense (see here in reply to here). Here’s Mark Johnson, for example, from the above linked to comment. I interpret “probabilistic” here as “Bayesian.”

A probabilistic model is what Marr called a "computational model"; it specifies the different kinds of information involved and how they interact.

There is a good reason for why Bayesian’s endorse this view of their proposals; interpreted algorithmically these theories often seem to be computational disasters (e.g. here and here). Suffice it to say that there seems to be general agreement that Bayesian analyses do not lend themselves to easy transparent  algorithmitization. Oddly, as Stabler notes here, in discussing the efficient parsability of MG definable languages, this is not the case for standard minimalist grammars:

In CKY and Earley algorithms, the operations of the grammar (em and im) [internal and external merge, NH] are realized quite directly [my emphasis, NH] by adding, roughly, only bookkeeping operations to avoid unnecessary steps (p.8).

Indeed, in my limited experience most generative parsing models starting from the Marcus parser onward have had relatively transparent relations between grammars and parsers, and this was considered a virtue of both (see here).  Indeed, since grammars to get used and seem to be used relatively effectively, it would be copacetic if there were a nice simple relation between competence grammars and parsing grammars (between grammars and algorithms that are used to parse incoming “language”).  However, from what I can gather, this is something of a problem for Bayesians, as it seems that the move from a level 1 competence theory to a level 2 algorithmic accounts won’t be particularly simple or transparent as the simple transparent ones seem to be a computational mess. 

My earlier post on word acquisition (here) touches on this, noting that the procedures that SYTG ran to test the Bayesian story was a pain to run and that this is a general feature of Bayesian accounts. However, I am sure that the issue is very complex and that I have probably misunderstood matters (hence my setting down my worries to act, I hope, as useful targets).

Let me end on one thing that I like about Bayesian approaches. It seems that they have a good way of dealing with what Mark calls chicken and egg problems. Bayesians have ways of combining two difficult problems to make the solving of each easier. This looks like it would very often be a useful thing to be able to do. And to the degree that Bayesians offer a compelling way of doing this, this is a good thing. A question: is this kind of solution to the chicken-egg problem limited to Bayesian kinds of analysis or is it a property that simpler counting systems could encode as well? If it is a distinctive property of Bayesianism, that would seem to be a very nice feature.

Let me abruptly end here and let the target practice (and enlightenment) begin. 



[1] This is one of the nicest things about blogs. In contrast to articles where you need arguments that start from reasonable premises and go in a coherent direction, in a blog post it is possible to ruminate out load and hope that with a little help from your “friends” a little more clarity might be forthcoming.
[2] It is a curious consequence of the Bayesian position that, at first blush, they postulate a whole lot more native givens than Rs typically do. I suspect that this is related to what Gallistel and King call the “infinitude of the possible.” The Bayesian way is to load the hypothesis space with LOTS of options and winnow them down using environmental information. This puts a lot into the space. If the space is given (i.e. innate) then this approach has the curious property of loading the mind with a lot of stuff, much more than Rs typically consider there.  So, in a curious sense, Es of this stripe are far more “nativist” than Rs are.

Sunday, December 2, 2012

A False Truism?


It’s my strong impression that everyone working on language, be they generativist partisans or enemies, takes it for granted that linguistic performance (if not competence) involves a heavy dose of stats. So, for example, the standard view of language learning is basically Bayesian, i.e. structured hypothesis space with grammars ordered by some sort of simplicity metric the “winner” being the simplest one consistent with the incoming data, the procedure involving a comparison of alternatives ranked by simplicity and conformity with the data.  This is indistinguishable from the set up in Aspects chapter 1.
Same with language processing, where alternative hypotheses about the structure of the incoming sounds/words/sentences are compared and assessed; the simplest one best fitting the data carrying the parsing day.[1] Thus, wisdom has it that language use requires the careful, gradual and methodical assessment of alternatives, which involves iteratively trading off some sophisticated measure of goodness of fit against some measure of simplicity to eventually get to a measure of believability. Consequently, nobody really wonders anymore whether statistical estimation/calculation is a central feature of our cognitive lives but which particular versions are correct.  Call this the “Stats Truism.”  Jeff Lidz recently sent me a very interesting paper that suggests that this unargued for presupposition, one incidentally that I have shared (do share?), should be moved from the truism column to the an-assumption-that-needs-argument-and-justification pile. Let me explain.

Kids acquire grammars very quickly. By about age 5 a kid’s grammar is largely set.  What’s equally amazing is that by age six, kids typically have a vocabulary of 6,000-8,000 words, which translates into an acquisition rate of roughly 6-8 words per day from the day they were born. This is a hell of a rate! (remember, for the first several years kids sleep half the day (if their parents are lucky) and cry for the other half (do I remember that!)). Psychologists have investigated how they do this and the current wisdom is that they employ some kind of fast mapping of sounds to concepts (so much for Quine’s derision of the museum myth!). Medina, Snedeker, Trueswell and Gleitman (MSTG) study this fast mapping process and what they discover is a serious challenge for the Stats Truism. Here’s why. They find that kids learn words more or less as follows: they quickly jump to a conclusion about what a word means and don’t let go!  They don’t consider alternative possibilities, they don’t revise prior estimates of their initial guess, and they don’t much care about the data after their initial guess (i.e. they don’t guess again when given disconfirming evidence for their initial hyporthesis, at least for a while).[2] Or, in MSTG’s own words:

1.     “Learners hypothesize a single meaning based on their first encounter with a word (3/6).” Thus, learning is essentially a one trial process where everything but the first encounter is irrelevant.
2.     “Learners neither weight nor even store back-up alternative meanings (3/5).” Thus, there is no (fancy (i.e. Bayesian) or crude (i.e. simple counting)) hypothesis comparison/testing going on as in word learning kids only ever entertain a single hypothesis.
3.     “On later encounters, learners attempt to retrieve their single hypothesis from memory and test it against a new context, updating only if it is disconfirmed. Thus they do not accrue a “bests final hypothesis by comparing multiple …semantic hypotheses (3/6).” Indeed, MSTG note that “a false hypothesis once formed blocks the formation of new ones (4/6).”

In sum, kids are super impulsive, narrow-minded, pig-headed learners, at least for words (though I suspect parents may not find this discovery this so surprising).

MSTG note that if they are correct (and the experiments are cool (and pretty convincing) so look them over) this constitutes a serious challenge to standard statistical models in which “each word-meaning hypothesis [is] based on the properties of all past learning instances regardless of the order in which they are encountered [and] numerous hypotheses [are held] in mind (with changing weights) until some learning threshold is reached (3/6).” SMTG’s point: at least in this domain the Stats Truism isn’t.

A further interesting feature of this paper is that it suggests why the Truism doesn’t hold.  MSTG argue that the standard experimental materials for investigating word learning in the lab don’t scale up. In other words, when materials that more adequately reflect real world situations are used the signature properties of statistical learning disappear.

The main difference between real world situations and the lab context is well summed up in an aphorism I once heard from Lila (the G in MSTG): “A picture is worth a thousand words, and that’s the problem.”  MSTG observes that in real life situations learners cannot match “recurrent speech events to recurrent aspects of the observed world because “the world of words and their contexts is enormously [i.e. too?-NH] complex (1/6).” As MSTG note, most word learning investigations abstract away from this complexity, either by assuming stylized learning contexts or assuming that a kid’s attention is directed in word learning settings so that the noise is cancelled out and the relevant stimulus is made strongly salient. MSTG provide pretty good evidence that these assumptions are unrealistic and that the domain in which words are learned is very busy and very noisy. As they note:

The world of words and their contexts is enormously complex. Few words are taught systematically…[I]n most instances, the situations of word use arise adventitiously as adults interact socially with novices. Words are heard buried inside multiword utterances and in situations that vary in almost endless ways…so that usually a listener could not be warranted in selecting a unique interpretation for a new item.

The important finding is that in these more realistic contexts, the signature properties of statistical learning disappear (not mitigated, not reduced, disappear).

MSTG suggest two reasons for this. First, within any given context too many things are plausibly relevant so picking out exactly what’s important is very hard to do.  Second, across contexts what’s relevant can change radically, so determining which features to consider is very difficult. In effect, there is no relevance algorithm either within or across contexts that the child can use to direct its attention and memory resources. Together these factors overwhelm those capacities that we successfully deploy in “stripped down laboratory demonstrations (1/6).”[3] Said less coyly: the real world is too much of a "blooming buzzing confusion" for statistical methods to be useful.  This suggestion sounds very counterintuitive, so let’s consider it for a moment.

It is generally believed that the virtue of statistical models is that they are not categorical and for precisely this reason they are better able to deal with the gradient richness of the real world (see here for example and here for discussion).  What MSTG’s results suggest is that crude rules of thumb (e.g. the first is best) are better suited to the real world than are sophisticated statistical models. Interestingly, others have made similar suggestions in other domains. For example, Andrew Haldane (here) invidiously compares bank supervision policies that depend on large numbers of weighted variables that are traded off against one another in statistically sophisticated ways with policies that use a single simple standard such as the leverage ratio. As he shows, the simple single standard blunt models outperform the fancy statistical ones most of the time. Gerd Gigerenzer (here) provides many examples where sophisticated statistical models are bested by rather ham fisted heuristic algorithms that rely on one-good-reason decision procedures rather than decision rules that weigh, compare and evaluate many reasons. Robert Axelrod observes how very simple rules of behavior (tit for tat) can triumph in complex interactive environments against sophisticated rules. It seems that sometimes less is more, and simple beats sophisticated.  More particularly, when the relevant parameters are opaque or it is uncertain how options should be weighted (i.e. when the hypothesis space is patchy and vague) statistical methods can do quite a bit worse than very simple very unsophisticated categorical rules.[4] In this setting, MSTGs results seem lees surprising.

MSTG makes yet one more interesting point.  It seems that the Stats Truism is true precisely when it operates against a rich, articulated, well-defined, set of options (“relatively” small helps too).[5] As MSTG note “statistical models have proven adequate for properties of language for which the learner’s hypothesis space is known to consist of a very small set of options (5/6).” I will leave it to your imagination to consider what someone who believes in a rich UG that circumscribes the space of possible grammars (me, me, me!) would make of this.

I would like to end with a couple of random observations.

First if anything like this is on the right track it further dooms the kind of big data investigations that I discussed here. What’s relevant is potentially unbounded. There’s no algorithm to determine it. To use Knight-Keynes lingo (see note 3), much of inquiry is uncertain rather than risky. If so, it’s not the kind of problem that big data statistical analysis will substantially crack precisely because inquiry is uncertain, not merely risky.[6]

Second, statistical language learning methods will be relevant in exactly those domains where the mind provides a lot of structure to the hypothesis space.  Absent this kind of UG-like structuring, sophisticated methods will likely fail. Word learning is an unstructured domain. Saussurian arbitrariness (the relation between a concept and the sound that tags it being arbitrary) implies that there isn’t much structure to word the learning context and, as MSTG show, in this domain, statistical learning methods are worse than useless, they are wrong. So, if you like statistical language learning and hope to apply it fruitfully you’d better hope that UG is pretty rich.

Third, there are two reasons for doubting that MSTG’s point will be quickly accepted. First, though simple, Baye’s Law looks a lot more impressive than ‘guess and stick’ or ‘first is best.’  There is a law of academia that requires that complicated trump simple. Simple may be elegant, but complexity, because it is obviously hard and shows off mental muscles (or at least appears to), enhances reputations and can be used to order status hierarchies. Why use simple rules when complex ones are available? Second, truisms are hard to resist because they do seem so true. As such the Stats Truism will not soon be dislodged. However, MSTG’s argument should at least open up our minds enough to consider the possibility that it may be more truthy than true.



[1] I believe that this is currently the dominant view in parsing (but remember I am no expert here). Earlier theories (see here) did not assume that multiple hypotheses were carried forward in time but that once decisions were made about the structure, alternatives were dropped. This was used to account for garden path phenomena and was motivated theoretically by the parsing efficiency of deterministic left corner parsers. 
[2] MSTG have an interesting discussion of how kids recover from a wrong guess (of which there are no doubt many). In effect, they “forget” what they guessed before.  Wrong guesses are easier to forget, correct ones stick in memory longer.  Forgetting allows this “first-is-best” process to reengage and allows the kid to guess again.
[3] Remember Cartwright’s observation (here for discussion) that Hume’s dictum is hard to replicate in the world outside the lab? Here’s a simple example of what she means.
[4] Sophisticated methods triumph when uncertainty can be reduced to risk. We tend to assume that these two concepts are the same. However, two very smart people argued otherwise: Frank Knight and J.M. Keynes argued for the importance of the distinction and urged that we not assimilate the two. The crux of the difference is that whereas risk is calculable, uncertainty is not. Statistical methods are appropriate in evaluating risk but not in taming uncertainty.
[5] Gigerenzer and Haldane also try to identify the circumstances under which more sophisticated procedures come into their own and best the simple rules of thumb.
[6] A point also made by Popper here. We cannot estimate what we don’t know. There are, in the immortal words of D. Rumsfeld “unknown unknowns.”