Comments

Showing posts with label Gigerenzer. Show all posts
Showing posts with label Gigerenzer. Show all posts

Thursday, April 23, 2015

Where the estimable Mark Johnson corrects Norbert (sort of)

For various reasons, Mark J could not post this as a comment on this post. As he knows much more than I do about these matters, I thought it a public service to lift these remarks from the comments section to make them more visible. I think that this is worth reading in conjunction with Charles' recent post (here). At any rate, this is interesting stuff and I don't disagree much (or have not found reasons to disagree much) with what Mark J says below. I will of, course,  allow myself some comments later. Thx Mark. 

***

This was originally written as a comment for the "Bayes and Gigerenzer" post, but a combination of a length restriction on comments and my university's not enabling blog posts from our accounts meant I had to email Norbert directly.

As Norbert has remarked, Bayesian approaches are often conflated with strong empiricist approaches, and I think this post does that too.  But even within a Bayesian approach, there are powerful reasons not to be a "tabula rasa" empiricist.  The "bias-variance dilemma" is a mathematical statement of something I've seen Norbert say in this blog: learning only works when the hypothesis space is constrained.  In mathematical terms, you can characterise a learning problem in terms of its bias -- the range of hypotheses being considered -- and the variance or uncertainty with which you can identify the correct hypothesis.  There's a mathematical theorem that says that in general as the bias goes down (i.e., the class of hypotheses increases) the variance increases.

Given this, I think a very reasonable approach is to formulate a model that includes as much relevant information from universal grammar as we can put into it, and perform inference that is as close to optimal as we can achieve from data that is as close as possible to what the child receives.  I think this ought to be every generative linguist's baseline model of acquisition!  Even with an incomplete model and incomplete data, we can obtain results of the kind "innate knowledge X plus data Y can yield knowledge Z".

But for some strange (I suspect largely historical) reason, this is not how Chomskyian linguists think of computational models of language acquisition.  Instead, they prefer ad hoc procedural models.  Everyone agrees there has to be some kind of algorithm which children use to learn language.  I know there are lots of pyschologists who are sure they have a good idea of the kinds of things kids can and can't do, but I suspect nobody really has the faintest idea of what algorithms are "cognitively plausible".  We have little idea of how neural circuitry computes, especially over the kinds of hierarchical representations we know are involved in language.  Algorithms which can be coded up as short computer programs (which is what most people have in mind when they say simple) might turn out to be neurally complex, while we know that the massive number of computational elements in the brain enable it to solve computationally very complex problems.  In vision -- a domain we can sort-of study because we can stick electrodes into animals' brains -- it turns out that the image processing algorithms implemented in brains are actually very sophisticated and close to Bayes-optimal, backed up with an incredible amount of processing power.  Why not start with the default assumption that the same is true of language?

It's true that in word segmentation, a simple ad hoc procedure -- a simple greedy learning algorithm that ignores all interword dependencies -- actually does a "reasonable job" of segmentation, and that improving either the algorithm's search procedure or making it track more complex dependencies actually decreases the overall word segmentation accuracy.  Here I think Ben Borschinger's comment has it right - sometimes a suboptimal algorithm can correct for the errors of a deficient model if the errors of each go in the opposite way.  We've known since at least Goldwater's work that an inaccurate "unigram" model that ignores inter-word interactions will prefer to find multi-word collocations and hence undersegment.  On the other hand, a naive greedy search procedure tends to over-segment, i.e., find word boundaries where there are none.  Because the unigram model under-segments, while a naive greedy algorithm over-segments, the combination actually does better than approaches where you just improve only the search procedure or only the model (by incorporating inter-word dependencies) since now you have "uncancelled errors".

Of course it's logically possible that children use ad hoc learning procedures that rely on this kind of "happy coincidence", but I think it's unlikely for several reasons.

First, these procedures are ad hoc -- there is no theory, no principled reason why they should work.  Their main claim to fame is that they are simple, but there are lots of other "simple" procedures that don't actually solve the problem at hand (here, learn words).  We know that they work (sort of) because we've tried them and checked that they do.  But a child has no way of knowing that this simple procedure works while this other one doesn't, so the procedure would need to be innately associated with the word learning task.  This raises Darwin's problem for the ad hoc algorithm (as well as other related problems: if the learning procedure is really innately specified, then we ought to see dissociation disorders in acquisition, where the child's knowledge of language is fine, but their word learning algorithm is damaged somehow).

Second, ad hoc procedures like these only partially solve the problem, and there's usually no clear way to extend them to solve the problem fully, so some other learning mechanism will be required anyway.  For example, the unigram+greedy approach can find around 3/4 of tokens and 1/2 of the types, and there's no obvious way to extend it so it learns all the tokens and all the types.  But children do eventually learn all the tokens and all the types, and we'll need another procedure for doing this.  Note that the Bayesian approach that relies on more complex models does have an account here, even though it currently involves "wishful thinking": as the models become more accurate by including more linguistic phenomena and the search procedures become more accurate, the word segmentation accuracy continues to improve.  We don't know how to build models that include even a fraction of the linguistic knowledge of a 3 year old, but the hope is that eventually these models would achieve perfect word segmentation, and indeed, be capable of learning all of a language.  In other words, there isn't a plausible path by which the ad hoc approach would generalise to learning all of a language, while there is plausible path for the Bayesian approach that relies on more and more accurate linguistic models.

Finally -- and I find it strange to be saying this to a linguist who is otherwise providing very cogent arguments for linguistic structure -- there really are linguistic structures and linguistic dependencies, and it seems weird to assume that children use a learning procedure that just plain ignores them.  Maybe there is a stage where children think language consists of isolated words (this is basically what a unigram model assumes), and the child only hypothesises larger linguistic structures after some "maturation" period.  But our work shows that you don't need to assume this; instead, a single model that does incorporate these dependencies combined with a more effective search procedure actually learns words from scratch more accurately than the ad hoc procedures.

Norbert sometimes seems very sure he knows what aspects of language have to be innate.  I'm much less sure myself of just what has to be innate and what can be learned, but I suspect a lot has to be innate (I think modern linguistic theory is as good a bet as any).  I think an exciting thing about Bayesian models is that they give us a tool for investigating the relationship between innate knowledge and learnability.  For example, if we can show that a model with innate knowledge X+X' can learn Z from data Y, but a model with only innate knowledge X fails to learn Z, then probably innate knowledge X' plays a role in learning Z.  I said probably because someone could claim that the child's data isn't just Y but also includes Y' and from model X and data Y+Y' it's possible to infer Z.   Or someone might show that a completely different set of innate knowledge X'' suffices to learn Z from Y.  So of course a Bayesian approach won't definitely answer all the questions about language acquisition, but it should provide another set of useful constraints on the process.


Tuesday, April 21, 2015

Bayes and Gigerenzer

Gigerenzer identifies an important assumption made in “rational” theories of cognition, of which Bayes is one variety. He calls it, following Carnap, the principle of full information (see here and here for discussion). I’ll call it “Carnap’s Principle” (CP). CP is one of the features that makes rational theories of cognition rational. How? Well, CP requires taking all the relevant data into account. Not some of it, not the stuff you might like, but all of it, even the troublesome bits. Agents are rational, at least in an ideal Level-1 Marr’s sense, to the degree that they cleave to CP.[1] Of course, this is not always doable, for the best may be unachievable, demanding resources (memory, computational speed) that are biologically/cognitively unavailable. Theories of constrained optimization accommodate this fact by considering possible ways of approximating the Level-1 “best” solution given the actual computational resources at hand. The spirit of CP (take account of all the relevant data) is respected in algorithms that take account of all the relevant information it is computationally possible to take account of. The way this gets into Bayes stories is that first we try to figure out what would be best given all the data and then superimpose algorithms on this that respect the resource problems; algorithms being “better” the closer they approximate the ideal without being too resource demanding. Put another way, Carnap’s full information principle is part of the theory of competence, and resource constraints are part of the theory of performance. Performance approaches the ideal to the degree that it respects CP.[2]

Gigerenzer observes that constrained optimization approaches that incorporate CP as an ideal support an interesting counterfactual: were there no cost to using all the data, then using all the available data would be best (‘best’ here means would result in better (best?) outcomes). Or, to put this another way, ignoring information exacts a cost that would be mitigated were all the information actually used. In a word, the rational thing to do is also the best thing to do were we but able to do it, but, sadly, because of resource constraints we cannot do what’s rational, which is why we seem not to be acting optimally when we behave (see here for one exposition). 

Gigernezer challenges CP. He argues that there are times when using all the available  information leads to worse outcomes than ignoring some of it does. One might describe this in a way that heightens paradox as follows: acting rationally can be counter-productive or studied ignorance can be empowering. On this view, there are times when the “best” theory is one that ignores relevant data in making its calculations. And ‘best’ here means that using all the data even when so using it is NOT computational onerous systematically yields worse results than using only part of it. In other words, Gigerenzer’s point is that there are situations where abstracting from resource costs, it is more rational to use less data than all of the data if by ‘rational’ one means gets better results.

Given this possibility, there are two kinds of situations to consider which result in two different kinds of “failure” to achieve an outcome: (i) failure resulting from not using all the information and (ii) failure resulting from actually using all the relevant information. The world is a complicated place it seems and the best strategy intimately depends on the circumstances.

It’s worth noting that CP is an add-on to Bayes. By this I mean that one can use Bayesian methods both over an informationally restricted domain or over one that includes all the data.  Bayesian methods are agnostic wrt this (though CP helps underwrite the rationality assumption that is part of Bayes methodological justification (the rational analysis meme)).  The question then is really one about the right characterization of the relevant data: which should be used and which ignored to get the best results. Rational theories assume that all usable data should in fact be used. Less is NEVER more! Gigernezer disagrees.

Gigerenzer has offered several empirical examples that back this reasoning up. What I want to consider here are whether there any linguistiky examples? Here I offer for your delectation two that seem to illustrate Gigerenzer’s observations.

I reviewed one a while ago when discussing Medina, Snedeker, Trueswell & Gleitman’s model of word learning and a modest revision of this model by Stevens, Yang, Trueswell and Gleitman (SYTG) (here, here, and here). A key feature of these models is that they reject “cross situational learning,” this being “the tabulation of multiple, possibly all, word-meaning associations across learning instances.” In other words, a key feature of these models is that they did very many fewer calculations than they could have done. They thus exploited a vey reduced subset of the relevant information available. Usefully, SYTG compares the results of richer uses of cross situational learning with their model, which makes use of a very restricted subset of the available information. The interesting result is that the models that more approximate the CP ideal do worse than the one that is far more myopic. In other words, less is more in this particular case, suggesting that even were one able to compute all the relevant alternatives (though it should be noted that it’s very computationally costly to do so), it would be ill-advised for it leaves the learner worse off. The problem then is not that the computation is too hard to do (though it is that as well), but that doing it even were it computationally cheap offers worse outcomes than doing a more half-ass job.[3] Let’s hear it for sloth!

I’ve recently read another illustration of this logic, this time as pertains to learning word segmentation (here). The paper is by Kasia Hitczenko (currently at UMD, yeah!!) and Gaja Jarosz (at Yale) (H&J). This was Kascia’s UG thesis work (and yes, they young are more talented than we were when we were their age. Thank the lord that I don’t have to compete for grad school admission now). At any rate, the paper compares two methods of word segmentation acquisition; ideal versus constrained learner models. The key is that the former abstract from memory limitations that the latter incorporate.

H&J investigate two models of word segmentation acquisition that incorporate the same probability model but differ in that one is incremental (and hence more realistic) and consequently executes a more limited set of computations. In particular, in the batch learner, the estimation procedure is “calculated over the whole corpus, while in the incremental algorithm, it is calculated over individual utterances” (3). The former calculation is far more extensive in that it evaluates goodness of fit over all the possible information in the corpus, unlike the incremental learner, which does a less thorough set of comparisons. Surprisingly, the more limited incremental learner does much better than the more rational batch learner. Moreover, and this is the important thing that H&J shows, what makes the incremental learner better is its memory limitations (i.e. the fact that unlike the batch learner it does not make all the comparisons it is possible to make). H&J demonstrate this by adding two kinds of memory limitations to the batch learner model and then showing that with these restrictions in place (forcing the batch learner to do a less thorough job), the batch learner does as well as the incremental learner with regard to segmentation. In other words, it appears that the ideal learner does less well than a more myopic one but not because of resource constraints. Rather, as H&J says concerning the batch models: “providing this model with access to all of the available information prevents it from segmenting accurately.” This is a fact about the segmentation procedure not about resource limitations.

This last point is important. It is common knowledge that ideal learning is often not feasible. The calculations are just too intractable. However, here, as in the earlier word learning models discussed above, the problem is not computational cost. The result is that actually doing the fuller computation leaves the learner worse off than not doing it all. In other words, abstracting away from computational cost, the more “rational” model, the one closer to endorsing CP, does worse than the more limited one. Thus, this is another example, in the linguistic domain, where less really is more.

This result is of interest not only to Bayesians, but to those like me who’ve always liked thinking of acquisition problems in terms of ideal speaker hearers (ISH).  ISHs have perfect unboundedly capacious memories and are able to take all the data in at once (instantaneous learning). What the results above seem to indicate is that ISH assumptions might actually make acquisition harder, as disregarding some of the data may be required to acquire a language at all. And this is interesting, for it suggests that for some matters (e.g. when statistical calculations become important), the ISH idealization might be misleading precisely because it incorporates Carnap’s Principle as an ideal.

We all know that idealizations are literally speaking false in that they abstract away from what we know to be many (possibly important) factors. What Gigerenzer has pointed out is that they can be false in two ways, one more benign than the other, IMO.

The first way is that they identify abstract away from resource constraints that we know to be important. LADs do not have perfect arbitrarily capacious memories. We know this. And thus real learners might fail to realize the properties we ascribe to ideal learners for precisely this reason. The utility of the idealization remains, however, as we assume that the closer the real learner’s resources approach the ideal conditions, the closer the real learner’s acquisition gets to the ideal limit case.

The second way to fail is to misunderstand that less can be more; in particular that having resource limitations may be important factors in making acquisition possible  (or at least much easier). This is less intuitive for it suggests that there is a resource curse; you can have too much of a good thing.

So, it seems that Gigerenzer’s observations have empirical utility, at least in the domain of word acquisition. So next time you are forgetful, remember there may be an up side. Given my current early onset proclivities, I sure hope so.


[1] Your favorite neighborhood Bayesian makes this assumption, which is one of the reasons that it is a species of rational analysis.
[2] I should add here, that this is not the way I think that linguists use, or ought to use, the competence/performance distinction. Competence is not some ideal target that performance tries to hit. Rather, competence theories are theories of what a cognizer knows and performance theories are accounts of how this knowledge is put to use. What one knows is not in any reasonable sense a “target” that performance theories aim to hit.
[3] The way that SYTG argues for this is that they actually do the costly computation and show that it is less successful than the one that eschews cross situational learning. So, even though the CP compliant computations are indeed costly, this is not the reason that they are less successful in the word learning conditions explored. Of course, that they are computationally expensive is another reason to reject them. But it is an additional reason. It’s the counterfactual that is Gigerenzer interesting.

Sunday, December 2, 2012

A False Truism?


It’s my strong impression that everyone working on language, be they generativist partisans or enemies, takes it for granted that linguistic performance (if not competence) involves a heavy dose of stats. So, for example, the standard view of language learning is basically Bayesian, i.e. structured hypothesis space with grammars ordered by some sort of simplicity metric the “winner” being the simplest one consistent with the incoming data, the procedure involving a comparison of alternatives ranked by simplicity and conformity with the data.  This is indistinguishable from the set up in Aspects chapter 1.
Same with language processing, where alternative hypotheses about the structure of the incoming sounds/words/sentences are compared and assessed; the simplest one best fitting the data carrying the parsing day.[1] Thus, wisdom has it that language use requires the careful, gradual and methodical assessment of alternatives, which involves iteratively trading off some sophisticated measure of goodness of fit against some measure of simplicity to eventually get to a measure of believability. Consequently, nobody really wonders anymore whether statistical estimation/calculation is a central feature of our cognitive lives but which particular versions are correct.  Call this the “Stats Truism.”  Jeff Lidz recently sent me a very interesting paper that suggests that this unargued for presupposition, one incidentally that I have shared (do share?), should be moved from the truism column to the an-assumption-that-needs-argument-and-justification pile. Let me explain.

Kids acquire grammars very quickly. By about age 5 a kid’s grammar is largely set.  What’s equally amazing is that by age six, kids typically have a vocabulary of 6,000-8,000 words, which translates into an acquisition rate of roughly 6-8 words per day from the day they were born. This is a hell of a rate! (remember, for the first several years kids sleep half the day (if their parents are lucky) and cry for the other half (do I remember that!)). Psychologists have investigated how they do this and the current wisdom is that they employ some kind of fast mapping of sounds to concepts (so much for Quine’s derision of the museum myth!). Medina, Snedeker, Trueswell and Gleitman (MSTG) study this fast mapping process and what they discover is a serious challenge for the Stats Truism. Here’s why. They find that kids learn words more or less as follows: they quickly jump to a conclusion about what a word means and don’t let go!  They don’t consider alternative possibilities, they don’t revise prior estimates of their initial guess, and they don’t much care about the data after their initial guess (i.e. they don’t guess again when given disconfirming evidence for their initial hyporthesis, at least for a while).[2] Or, in MSTG’s own words:

1.     “Learners hypothesize a single meaning based on their first encounter with a word (3/6).” Thus, learning is essentially a one trial process where everything but the first encounter is irrelevant.
2.     “Learners neither weight nor even store back-up alternative meanings (3/5).” Thus, there is no (fancy (i.e. Bayesian) or crude (i.e. simple counting)) hypothesis comparison/testing going on as in word learning kids only ever entertain a single hypothesis.
3.     “On later encounters, learners attempt to retrieve their single hypothesis from memory and test it against a new context, updating only if it is disconfirmed. Thus they do not accrue a “bests final hypothesis by comparing multiple …semantic hypotheses (3/6).” Indeed, MSTG note that “a false hypothesis once formed blocks the formation of new ones (4/6).”

In sum, kids are super impulsive, narrow-minded, pig-headed learners, at least for words (though I suspect parents may not find this discovery this so surprising).

MSTG note that if they are correct (and the experiments are cool (and pretty convincing) so look them over) this constitutes a serious challenge to standard statistical models in which “each word-meaning hypothesis [is] based on the properties of all past learning instances regardless of the order in which they are encountered [and] numerous hypotheses [are held] in mind (with changing weights) until some learning threshold is reached (3/6).” SMTG’s point: at least in this domain the Stats Truism isn’t.

A further interesting feature of this paper is that it suggests why the Truism doesn’t hold.  MSTG argue that the standard experimental materials for investigating word learning in the lab don’t scale up. In other words, when materials that more adequately reflect real world situations are used the signature properties of statistical learning disappear.

The main difference between real world situations and the lab context is well summed up in an aphorism I once heard from Lila (the G in MSTG): “A picture is worth a thousand words, and that’s the problem.”  MSTG observes that in real life situations learners cannot match “recurrent speech events to recurrent aspects of the observed world because “the world of words and their contexts is enormously [i.e. too?-NH] complex (1/6).” As MSTG note, most word learning investigations abstract away from this complexity, either by assuming stylized learning contexts or assuming that a kid’s attention is directed in word learning settings so that the noise is cancelled out and the relevant stimulus is made strongly salient. MSTG provide pretty good evidence that these assumptions are unrealistic and that the domain in which words are learned is very busy and very noisy. As they note:

The world of words and their contexts is enormously complex. Few words are taught systematically…[I]n most instances, the situations of word use arise adventitiously as adults interact socially with novices. Words are heard buried inside multiword utterances and in situations that vary in almost endless ways…so that usually a listener could not be warranted in selecting a unique interpretation for a new item.

The important finding is that in these more realistic contexts, the signature properties of statistical learning disappear (not mitigated, not reduced, disappear).

MSTG suggest two reasons for this. First, within any given context too many things are plausibly relevant so picking out exactly what’s important is very hard to do.  Second, across contexts what’s relevant can change radically, so determining which features to consider is very difficult. In effect, there is no relevance algorithm either within or across contexts that the child can use to direct its attention and memory resources. Together these factors overwhelm those capacities that we successfully deploy in “stripped down laboratory demonstrations (1/6).”[3] Said less coyly: the real world is too much of a "blooming buzzing confusion" for statistical methods to be useful.  This suggestion sounds very counterintuitive, so let’s consider it for a moment.

It is generally believed that the virtue of statistical models is that they are not categorical and for precisely this reason they are better able to deal with the gradient richness of the real world (see here for example and here for discussion).  What MSTG’s results suggest is that crude rules of thumb (e.g. the first is best) are better suited to the real world than are sophisticated statistical models. Interestingly, others have made similar suggestions in other domains. For example, Andrew Haldane (here) invidiously compares bank supervision policies that depend on large numbers of weighted variables that are traded off against one another in statistically sophisticated ways with policies that use a single simple standard such as the leverage ratio. As he shows, the simple single standard blunt models outperform the fancy statistical ones most of the time. Gerd Gigerenzer (here) provides many examples where sophisticated statistical models are bested by rather ham fisted heuristic algorithms that rely on one-good-reason decision procedures rather than decision rules that weigh, compare and evaluate many reasons. Robert Axelrod observes how very simple rules of behavior (tit for tat) can triumph in complex interactive environments against sophisticated rules. It seems that sometimes less is more, and simple beats sophisticated.  More particularly, when the relevant parameters are opaque or it is uncertain how options should be weighted (i.e. when the hypothesis space is patchy and vague) statistical methods can do quite a bit worse than very simple very unsophisticated categorical rules.[4] In this setting, MSTGs results seem lees surprising.

MSTG makes yet one more interesting point.  It seems that the Stats Truism is true precisely when it operates against a rich, articulated, well-defined, set of options (“relatively” small helps too).[5] As MSTG note “statistical models have proven adequate for properties of language for which the learner’s hypothesis space is known to consist of a very small set of options (5/6).” I will leave it to your imagination to consider what someone who believes in a rich UG that circumscribes the space of possible grammars (me, me, me!) would make of this.

I would like to end with a couple of random observations.

First if anything like this is on the right track it further dooms the kind of big data investigations that I discussed here. What’s relevant is potentially unbounded. There’s no algorithm to determine it. To use Knight-Keynes lingo (see note 3), much of inquiry is uncertain rather than risky. If so, it’s not the kind of problem that big data statistical analysis will substantially crack precisely because inquiry is uncertain, not merely risky.[6]

Second, statistical language learning methods will be relevant in exactly those domains where the mind provides a lot of structure to the hypothesis space.  Absent this kind of UG-like structuring, sophisticated methods will likely fail. Word learning is an unstructured domain. Saussurian arbitrariness (the relation between a concept and the sound that tags it being arbitrary) implies that there isn’t much structure to word the learning context and, as MSTG show, in this domain, statistical learning methods are worse than useless, they are wrong. So, if you like statistical language learning and hope to apply it fruitfully you’d better hope that UG is pretty rich.

Third, there are two reasons for doubting that MSTG’s point will be quickly accepted. First, though simple, Baye’s Law looks a lot more impressive than ‘guess and stick’ or ‘first is best.’  There is a law of academia that requires that complicated trump simple. Simple may be elegant, but complexity, because it is obviously hard and shows off mental muscles (or at least appears to), enhances reputations and can be used to order status hierarchies. Why use simple rules when complex ones are available? Second, truisms are hard to resist because they do seem so true. As such the Stats Truism will not soon be dislodged. However, MSTG’s argument should at least open up our minds enough to consider the possibility that it may be more truthy than true.



[1] I believe that this is currently the dominant view in parsing (but remember I am no expert here). Earlier theories (see here) did not assume that multiple hypotheses were carried forward in time but that once decisions were made about the structure, alternatives were dropped. This was used to account for garden path phenomena and was motivated theoretically by the parsing efficiency of deterministic left corner parsers. 
[2] MSTG have an interesting discussion of how kids recover from a wrong guess (of which there are no doubt many). In effect, they “forget” what they guessed before.  Wrong guesses are easier to forget, correct ones stick in memory longer.  Forgetting allows this “first-is-best” process to reengage and allows the kid to guess again.
[3] Remember Cartwright’s observation (here for discussion) that Hume’s dictum is hard to replicate in the world outside the lab? Here’s a simple example of what she means.
[4] Sophisticated methods triumph when uncertainty can be reduced to risk. We tend to assume that these two concepts are the same. However, two very smart people argued otherwise: Frank Knight and J.M. Keynes argued for the importance of the distinction and urged that we not assimilate the two. The crux of the difference is that whereas risk is calculable, uncertainty is not. Statistical methods are appropriate in evaluating risk but not in taming uncertainty.
[5] Gigerenzer and Haldane also try to identify the circumstances under which more sophisticated procedures come into their own and best the simple rules of thumb.
[6] A point also made by Popper here. We cannot estimate what we don’t know. There are, in the immortal words of D. Rumsfeld “unknown unknowns.”