Comments

Tuesday, November 18, 2014

Why Morphology? (part deux)

As Chomsky has repeatedly emphasized, natural language (NL) has two distinctive features; it’s hierarchically recursive and it contains a whole bunch of lexical items (LI) with rather distinctive properties (when compared to what one finds in animal communication systems (e.g. they are not stimulus bound and they are very conceptually labile)). A natural question that arises is whether these two features are related to one another; does the fact that NLs are hierarchically recursive (actually generated by Gs that are recursive and produce hierarchical linguistic objects, or the corresponding I-language is such, but let’s forget the niceities for now) have any causal bearing on the fact that NL lexicons are massive or vice versa?[1] The only proposal that I know of that links these two facts together causally is Lila Gleitman’s hypothesis that vocab acquisition leverages syntactic knowledge. The syntactic bootstrapping thesis is that Gs facilitate (and hence accelerate) vocab acquisition in LADs. By “vocab” I mean “open” class contentful items like ‘shoe’ and ‘rutabaga’ rather than “closed” class items like ‘in’ or ‘the’ or the plural ‘-s’ morpheme. NLs have a surprisingly large number of open class “words” when compared to other animal systems (at least 3 orders of magnitude larger) and it is reasonable to ask why this is so.[2]

Gleitman’s syntactic bootstrapping thesis provides a possible answer: without syntactic leverage, acquiring words is a very very arduous process (see here, here, and here for discussion about how first words are acquired), and as only humans have syntax, only they will have large vocabs. Oh, btw, by acquisition I intend something pretty trivial; tagging a concept with a label (generally phonological, but not exclusively so, think ASL). I don’t think that this is all there is to vocab acquisition,[3] but it turns out that even this is surprisingly difficult to accomplish, (contrary to intimation to the contrary in the philo of language literature, see Quine on the ‘museum myth’) and it requires lots of machinery to pull off (e.g. a way of generating labels, namely, something like a phonology and a way of identifying the things that need tagging).

I mention all of this because I just recently heard a lecture by Anne Christophe (see slides here) that goes into some details about the possible mechanics behind this leveraging process that bears on an earlier question I have been wondering about for a long time: why do NLs have so much phonologically overt morphology?  For any English speaker, morphology seems like a terrible idea (it’s a mess and a pain to learn, you should hear my German). However, Christophe argues that morphology and closed class items serve to facilitate the labeling process that drives open class vocab acquisition in humans. In other words, her work (and here I mean that of the work from her lab as there are many contributors here as the slides make clear) sketches the following picture: closed class items (and I including morphological stuff here) provide excellent environments for the identification and tagging of open class expressions. And as closed class items are in fact very closely tied to (correlated with) syntactic structure (think Jabberwocky!!), morphology, broadly construed, is what enables kids to leverage syntax to build vast lexicons. If correct, this forges a close causal connection between the two distinctive properties NLs display; large vocabs and syntactic structure. Let’s call this the Gleitman-Christophe thesis (GCT).

What’s the evidence? Effectively, morphology provides relatively stable landmarks between which content words sit thus allowing for easy identification and tagging. In other words, morphology allows the LAD to zero in on content words by locating them in relatively fixed morpho-syntactic frames. And very young LADs can use this information for they are known to be very good at acquiring this fixed syntactic scaffolding using non-linguistic distributional (and statistical) methods (see slide 18).[4] Closed class items are acquired early (they are ubiquitous, frequent and short) and it has been shown that kids can and do exploit them to “find” content words. Christophe reports on new work that shows how useful this can be for categorizing words into initial semantic classes (e.g. distinguishing Ns (canonically things) from Vs (canonically eventish)).  She then goes on to describe how acquiring a few words can further enhance the process of acquiring yet more vocab. The first few words act as “seeds” that once acquired serve to further leverage the acquisition process. In sum, Christophe describes a process in which prosody (which Christophe discusses in the first part of the talk), morphology and syntax greatly facilitate word acquisition, which in turn enables yet more word acquisition. We thus get a virtuous circle anchored in morpho-syntax, which is in turn anchored in epistemologically prior pre-linguistic capacities to find fixed phono-morpho-syntactic landmarks that allow LADs to quickly fix on new words.[5] This all provides good evidence in favor of the GCT. It also allows one to begin to formulate one answer to my earlier question: Why Morphology?

So what’s morphology for? (This is a dumb question really, for it could be for lots of things. However, thinking functionally is natural for ‘why’ questions. See below for a just so story). Well among other things It is there to support large vocabs and this, I would suggest is a very big deal.  I once suggested the following thought experiment: I/you am/are going to Hungary (I/you speak NO Hungarian) and am/are made the following offer: I/you can have a vocab of 50k words and no syntax whatsoever or a perfect syntax and a vocab of 10 words. Which would I/you prefer? I/(you?) would take door number 1. One can get a long way on 50k words. Nor is this only for purposes of “communication” (though were there an indirect tie between communicative enhancement and morpho-syntax I would be ok with that). Phenomenologically speaking, tagging a concept has an effect not unlike making an implicit assumption explicit. And explicitness is a very good way to enhance thought (indeed, it feels like it allows one to entertain novel thoughts). Having a word for something allows it to be conceptually accessible and salient in a way that having the concept inchoately does not. In fact, I am tempted to say, that having a word for something can change the way you think (i.e. it can affect cognitive competence, not just enhance performance).[6] So, tagging concepts really helps and tagging a lot of concepts really really helps. Thus, if you are thinking of putting in an Amazon order for a syntax, I would suggest asking for one that also supports large scale vocab acquisition (tagging concepts) and the GCT argues that such a syntax would come with phonologically overt morphology and phonologically overt closed class items that our pre-linguistic proclivities can exploit to build large lexicons quickly. In other words, if you want an NL that is good for thinking (and maybe also for communication) get one that has the kinds of relations between morphology and syntax that we actually find.

Note that this view of morphology leaves (completely?) open the question of how morphology functions inside UG. It is consistent with this view that G operations are driven by feature checking requirements, some of which become realized overtly in the phonology (this is characteristic of early Minimalist proposals). It is also consistent with the view that they are not (e.g. that they are mere by-products of grammatical operations rather than drivers thereof (this is what we find in later EPP based conceptions of grammatical operations in later minimalism). It is consistent with the idea that morphology exists to fit syntax to phonology (readjustment rules), or that it’s not (i.e. it’s a functionally useless ornamentation).  All the GCT requires is that there be reliable correlations between overt markers and the syntax so as to allow the LAD to leverage the syntax in order to acquire a content rich lexicon.

If this is right, then it might serve as the start of an account of why there is so much overt morphology and/or closed class items (very frequent items that function to limn syntactic structure) in NLs.  In fact, it suggests that there should be no NL that eschews both, though what the mix needs to be can be quite open. So Chinese and English don’t have much morphology, but they have classifiers (Chinese) or Determiners and Verbal morphology (English) and this can serve to do the lexicon building job (Anne C tells me that all the stuff done on French discussed in the slides replicates in Mandarin).

As an aside, IMO, one of the puzzles of morphology is why some languages seem to have so much (Georgian) and some seem to have so little (English, Chinese). If morphology is that important in the grammar either functionally (e.g. for processing) or grammatically (e.g. it drives operations) then why should some NLs express so much of it overtly and some have almost none at all that is visible. The GCT offers a possible way out: morphology per se is not what we should be looking at. Rather it is morphology plus closed class items; items that give you fixed frames for triangulating on content “words.” There needs to be a sufficiency of these to undergird vocab acquisition, but the mix need not be crucial (in fact it is not even clear how much of both is sufficient or if there may be a cost in having too much. Or even if these queries make any sense).

Let me end here. NLs are stuffed with what appears to be “useless” (and as an English speaker, cumbersome) morphology. And useless it may be from a purely grammatical point of view (note I say may, leaving the question open). But GCT suggests that overt grammar dependent fixed points can be very useful for building lexicons given our pre-linguistic capacities. And given the virtues of a good sized lexicon for thinking and communicating, a syntax that can support this should have advantages over one that doesn’t (that’s the just so story, btw). If correct, and the data collected so far is non-trivial, this is nice to know for it serves to possibly bridge two big facts about NLs, facts that to date seem (or more accurately, seemed to me) to be entirely independent of one another.





[1] Nothing I say touches on the fact that NL lexical items have very distinctive properties, at least when compared to symbols in animal systems. In other words, the lexicon presents two puzzles: (i) why is it so big? (ii) why do human lexical items function so differently from non-human ones. What follows tries to say something about the first, but leaves the second untouched. Chomsky has discussed some of these distinctive features and Paul Pietroski has a forthcoming book that discusses the topic of lexicalization. 
[2] So far as I know, no animal communication system has any closed class expressions, which, if they have no syntax, would not be a surprise given Gleitman’s conjecture.
[3] Paul Pietroski has a terrific forthcoming book on what lexicalization consists in “semantically” and I intend my views on the matter to closely track his. However, for present purposes, we can ignore the semantic side of the word acquisition process and concentrate solely on the process of phonetically tagging LIs.
[4] That kids are really good at identifying morphology has always surprised me. This is far less the case in second language acquisition if my experience is anything to go on. At any rate, it seems that kids rarely make “errors” of commission morphologically. If they screw up, which is surprisingly infrequently, it manifests as errors of omission. Karin Stromswold has work from a while ago documenting this.
[5] This might also provide a model for how to think of the thematic hierarchy.  Is this part of UG or not? Well, one reason for thinking not is that it is very hard to define theta roles so that they apply across a broad class of verbs. Dowty showed how hard it is to define ‘agent’ and ‘patient’ etc. and Grimshaw noted that it is largely irrelevant for syntactic concerns anyhow. When does it matter? Real, theta roles matter for UTAH. We need to know where a DP starts its derivational life.  If this is its sole role, then what one needs theta roles for is not semantic interpretation in general but for priming the syntax.  In this case, all one needs are a few thematically well behaved verbs (like ‘eat,’ ‘hug,’ ‘hit,’ ‘kiss,’) and to get the syntax off the ground. Once flying, we don’t need thematic information any more for there arise other ways, some of them being morpho-syntactic, to figure our where a DP came from (think case, or fixed word order position or agreement patterns). At any rate, like that morphology case, the thematic hierarchy need not be part of UG to be linguistically very important to the LAD.
[6] Think of this as a very weak (almost truistic) version of the Sapir-Whorf hypothesis.

Monday, November 17, 2014

The First Hilbert Question: Adriana Belletti

This is the first of our Hilbert-Question (HQ) posts. Adriana Belletti kicks the series off with the following query:
How much of detected developmental stages can be attributed to the internal complexity of formal/grammatical factors and how much to general mechanisms? Are there really general mechanisms?
There are structures and constructions that appear to be particularly hard in acquisition (Crain & Thornton 1998, Guasti 2002, 2007 for a background on issues in theoretically oriented language acquisition studies). Typically, this is the case cross-linguistically, in domains in which constructions in different languages are closely comparable (e.g relative clauses, passive …). For instance, it has been claimed that Passive is delayed in several languages and this has been attributed to late development of the different aspects of the formal computation involved (A-chain, Borer & Wexler 1987; by-phrase Fox & Grodzinsky 1998, smuggling à la Collins 2005 as in Hyams and Snyder 2005). However, recent and less recent (e.g. Demuth 2010, Crain et al. 1987/2009) contributions indicate that under appropriate conditions – strictly formal, discourse related and possibly linked to the richness in the input – passive is not that hard for even young children, across languages.  Similarly, in the acquisition of complex object relatives/A’-dependencies recent evidence has shown that not all types of object A’-dependencies are hard for young children. Appeal has been made to the principle of locality regulating intervention (Rizzi 1990, 2004) to account for the selected difficulty (Friedmann, Belletti, Rizzi (2009). This leads to the following more general query: do general mechanisms ever play any comparable role in providing refined explanations of these developmental milestones?

References:

Borer, H. & K. Wexler (1987) “Maturation in Syntax”, in Parameter Setting, T.Roeper     E.Williams eds. Reidel, 123-172
Crain, S. & R. Thornton (1998) Investigations in Universal Grammar, MIT Press
Crain, S., R. Thornton, K. Murasugi (1987/2009) “Capturing the evasive passive”, Language        Acquisition, 16.2, 123-133
Collins, C. (2005) “A Smuggling Approach to the Passive in English”, Syntax, 8.2, 81-120
Demuth, K., F. Moloi, M.Machobane (2010) “Three year-olds’ comprehension, production and generalization of Sesotho passive”, Cognition, 115.2, 238-251
Fox, D. & Y.Grodzinsky (1998)  “Children’s passive: A view from the by-phrase”, Linguistic      Inquiry, 29: 311-332
Friedmann, N., Belletti, A., Rizzi, L., 2009. Relativized relatives: Types of intervention in the acquisition           of A-bar dependencies. Lingua 119, 67–88
Guasti, M.T. (2002) Language Acquisition, MIT Press
Guasti, M.T. (2007) L’acquisizione del linguaggio, Raffaello Cortina Editore
Hyams, N. & W.Snyder (2005) “Young children never smuggle: Reflexive clitics and the universal            freezing hypothesis”, GALANA, University of Hawaii
Rizzi, L.  (1990) Relativized Minimality, MIT Press.

Rizzi, L.  (2004) “Locality and the left periphery”, in A. Belletti ed., Structures and beyond: The cartography    of syntactic structures, Vol. 3. OUP, 223–251

Tuesday, November 11, 2014

How to participate in our Hilbert Fest

I did not provide information about what to do if you have a question you want to pose. Put it in a word document of at most 3 pages and send it to me:
Nhornste@umd.edu

Hubert, Henk and I will then vet it, get back to you and we go on from there. Can't wait.

Size Matters

I post here for Bob Berwick.

Size Matters: If the Shoe Fits

Everyone’s taught in Linguistics 101 that Small is Beautiful: At the end of the day, if you wind up writing one grammar rule per example, your grammatical theory amounts to zip – simply a list of the examples, carrying zero explanatory weight.  Why zero? Because one could swap out any part of the data for something entirely different, even random examples from, say, Martian and wind up with a rule set just as long.  Computer scientists immediately recognize this for what it is: Compression. And most linguists know the classic illustration – English auxiliary verb sequences, from Logical Structure of Linguistic Theory to Syntactic Structures (1955/1975) and then to Aspects (1965), buffed to high polish in Howard Lasnik’s Syntactic Structures Revisited (2000), chapter 1.  Forgetting morphological affixes for the moment, there are 8 possible auxiliary verb sequences – 1 with zero aux verbs; 3 with 1 aux verb (some form of modal, have, or be); 3 with two aux verbs, and one with all 3 aux verbs, so 8 possible aux patterns in all. Assuming 1 rule per pattern, wind up with 8 rules, a description as long as the original data we started with. Worse, 8 rules could describe any sequence of three binary-valued elements (modal or not, have or not, be or not)[1].  However, if we assume that our representational machinery includes parentheses that can denote optionality, we can boil the 8 rules down to just one:  Aux à (Modal)(Have)(Be). In this way Compression = generalization. But, Aspects notes, whether we’ve got parentheses-as-optionality at our disposal or not is entirely an empirical issue. To again draw from the classic conceit, if we were instead Martians, and Martian brains couldn’t use parentheses but only something like ‘cyclic permutation’, then the Martians couldn’t compress the English auxiliary verb sequences the same way – what Martians see as small or large depends on what machinery they’ve got in their heads.[2]

However, this example misses – or perhaps tacitly puts aside – one important point: Size matters in another way.  If your shoes fit, they fit for two reasons: they are not too small, but they are also not too large.  As Goldilocks said, they are just right.  Which is simply to say that a good grammar not only has to compress as much as possible the data it’s trying to explain, wringing out the right generalizations, but it also has to fit the data as closely as possible. A grammar shouldn’t generate sentences (and sentence structures) that aren’t in the data.  Otherwise, a very, very small grammar, with just 3 or 4 symbols, e.g., S → Σ*, would ‘cover’ any set of sentences, and we’d be done. But that would be just like saying that Yao Ming’s size 18 basketball shoes fit everyone.  Of course they do, but for most people they are way too large.  Likewise, this one rule overgenerates wildly. The measure we want should reward both small grammars and tightly fitting ones, in some balanced way. Surprisingly, something resembling this appears among the earliest documents in contemporary generative grammar:[3]

“Given the fixed notation, the criteria of simplicity governing the ordering of statements are as follows: that the shorter grammar is the simpler, and that among equally short grammars, the simplest is that in which the average length of derivation of sentences is least” (Chomsky, 1949; 1951, p. 3).
So what do we find here? The first part – the shorter the grammar, the simpler (and better) – sounds pretty close to “Small is Beautiful”.  But what about the second part, measuring how the grammar “fits” the data? That’s less clear. One might argue that the “average derivation length of sentences” could serve as a proxy for goodness of fit.  This isn’t entirely crazy if we can identify longer derivations somehow with fit, but it isn’t clear how to do this, at least from this passage.[4] Then there’s the two-part aspect – first we collect the shortest grammars, and only then do we check to find the ones that have the shortest derivation lengths covering the data. So the two criteria aren’t balanced. Some improvement is called for.

It was the late computer scientist Ray Solomonoff (1956, 1960) who first figured out how to patch all this up – ironically it seems, on hearing Chomsky talk about generative grammar in at the IEEE Symposium on Information Theory, 1956, where Chomsky first publicly presented his “Three models for the description of language.” Ray realized that probabilistic grammars would work much better than the arithmetic formulas he had been using to study inductive inference.  The result was the birth of algorithmic information theory, which concerns the shortest description length of any sort of representation – including grammars (now called Kolmogorov complexity).[5]  The key insight is the modern realization that information, computation, and representation are all intertwined in an incestuous, deep ménage à trois: now that everyone has PCs, we are all familiar that it takes a certain number of “bits” to write down or encode, well, this blog, coding the letters a through z as binary strings, the way your computer does. But as soon as we talk about the number of bits required for descriptions, we have already entered into the realm of Shannon’s information theory, which considers exactly this sort of problem[6].  Probabilities are then necessary, since, very informally, the amount of information associated with a description is equal to the negative logarithm of its probability. Finally, since encoding and decoding require some kind of algorithm, computation’s in the mix as well. 

What Solomonoff discovered is that we should try to find the shortest description of a language wrt some family of grammars, where crucially the description includes the description length of the grammar for the language as well. This attempts to minimize two factors: (1) the length in bits of the data (some finite subset of a language) as encoded (generated) by that grammar[NH1]  plus (2) the length in bits of the description of the grammar itself (the algorithm to encode/decode the language). That’s the two-part formulation we’re looking for (and now we can more clearly see that “average derivation length,” assuming the exchangeability of derivation length for probabilities, with longer derivations corresponding to lower probability of generation, is in fact a kind of proxy for coding the fit of a grammar to a language).[7]  Now, if a dataset is truly random and contains no patterns at all then its description will be about as long as the data itself–there’s no grammar (or more generally no theory) that can describe it more compactly than this. But if the data does contain regularities, like the Aux verb examples, then we can wring out of the data lots of bits via a grammar. The single Aux rule takes only 10 symbols to write down (9 on the right-hand side of the rule, and S on the left-hand side), while listing all the aux sequences explicitly takes 20 symbols.[8] 

In short, minimizing factor (2) aims for the shortest grammar in terms of bits, but this must also take into account factor (1), minimizing the length of the encoded data. So there is an explicit trade-off between factors (1) and (2) that penalizes short but overly profligate grammars, or, conversely, long and overly restrictive grammars. A grammar such as “S → Σ*” is only 3 symbols long (where the alphabet includes all English words) and certainly covers the Auxiliary verb examples, but each auxiliary verb example is again generated via a separate, explicit expansion and the grammar wildly overgenerates. While the size of this grammar plus the data it covers might be short for the first one or two examples, after just a few more cases the single-rule grammar we displayed earlier beats it hands down. Worse, since this grammar also generates all the invalid aux strings, it will also code them just the same – just as short – as valid strings, that is, in a relatively small number of bits. And that signals overgeneralization. Similarly, an overly-specific grammar will also lose out to the single-rule aux grammar.  If the over-specific grammar generates exactly the first four auxiliary verb sequences and no more, it would have to explicitly list the remaining 4 examples, and the total length of these would again exceed the length of the single-rule grammar in bits. In this sense, minimizing description length imposes an “Occam’s Razor” principle that tries for the grammar that best compresses the data[NH2]  and the grammar.

The Solomonoff-Kolmogorov metric can be cashed out in several ways.  One familiar approach is the Bayesian approach to language learning, as used, e.g., originally by Horning (1969).  This in fact was Solomonoff’s tack as wellIt’s easy to see why. Given some data D and a grammar G ranging over some family of possible grammars, we try find the most likely G given D, i.e., the G that maximizes p(G|D). Rewriting this using Bayes’ theorem, this is the same as trying to find the grammar G in our family of grammars that maximizes p(G)×p(D|G)/p(D).   Since p(D) is fixed, this is the same as simply trying to find the G that maximizes the numerator, p(G)×p(D|G) – the prior probability of G, with smaller G’s corresponding to higher prior probabilities, times the likelihood of the data given the grammar – how well the data is fit by the grammar. Exactly what we wanted.

The Solomonoff-Kolmogorov measures can also be closely linked to the so-called Minimum Description Length principle (MDL) of Rissanen (1983), an apparent rediscovery of the Solomonoff-Kolmogorov approach, which also states that the best theory is the one that minimizes the sum of the length of the description of the theory in bits, plus the length of the data as encoded by the theory.[9]

So does this shoe fit?  How has this notion fared in computational linguistics? The real problem for all description length methods is that while they provide a reasonable objective function as a goal, they don’t tell you how to search the space of possible hypotheses. In general, the Solomonoff-Kolmogorov-MDL approach formally establishes a familiar trade-off between rules and memorized patterns: if any sequence like kick the bucket appears frequently enough, then it saves many bits to record it via (generative) rules, rather than simply listing it verbatim in memory.  And the more often a collocation like this appears, the more savings accrue. Such methods have been adopted dozens of times over the past 50-odd years, often directly in models of language acquisition, and often in a Bayesian form (see, e.g., Olivier, 1968; Wolff, 1977, 1980 1982; Berwick, 1982, 1985; Ellison, 1991; Brent and Cartwright, 1993; Rissanen and Ristad, 1994; Chen, 1995; Stolcke, 1994; Dowman, 1997; Villavincencio, 2002; Briscoe, 2005; Hsu & Chater. 2010; Hsu et al., 2011; and several others).  Most recently, MDL has been revived in an especially insightful way by Rasin and Katzir, 2014, “On evaluation metrics in OT.” At least to my mind, Rasin and Katzir’s MDL approach make a good deal of sense for evaluation metrics generally, and is well worth considering seriously as the 50th anniversary of Aspects approaches. It’s available here: http://ling.auf.net/lingbuzz/001934.

One blog post clearly can’t do justice to all this work, but, for my money, the MDL highwater mark so far remains Carl De Marcken’s seminal Unsupervised Language Acquisition (1996 doctoral dissertation, MIT), available here:
http://www.demarcken.org/carl/papers/PhD.pdf; and summarized (and extended) in his 1996 ACL paper, Linguistic Structure as Composition and Perturbation, here: http://www.demarcken.org/carl/papers/acl96.pdf. Starting out with only positive-example strings of phonemes pseudo-transcribed from speech (so in particular, no ‘pauses’ between words), Carl’s unsupervised MDL method crunches through the Wall Street Journal and Brown corpuses, building generatively stored hierarchical “chunks” at several linguistic levels and “compiling out” common morphological, syntactic, and even syntactic-semantic patterns compositionally for reuse, e.g., from “to take off” as [vp <v><p><np>]◦[v take]◦[p off]◦[np ], all the way to “cranberry {RED BERRY}” as c◦r◦a◦n◦berry{BERRY+RED} (here the open circle ◦ is a composition operator that’s just string concatenation in the simplest case, but hierarchical composition for syntax, i.e., merge). As Carl notes, “no previous unsupervised language-learning procedure has produced structures that match so closely with linguistic intuitions.”  So, just about 20 years ago, we already had a learning method based on Solomonoff-type induction that really worked – on full-scale corpora too.  What happened?  Carl’s approach was never really surpassed. A few people (most notably John Goldsmith) embraced the MDL idea for their own work and a few more picked up on Carl’s compositional & probabilistic hierarchical MDL-“chunks” approach. But for the most part it seems as though the AI and computer science & learning community fell into its usual bad habit of failing to stand on the shoulders of past success – hard to tell whether one should chalk this up to a familiar need to always appear sui generis, or to ensure that the MDL wheel would be faithfully rediscovered every generation to the same great surprise, or to just plain scholarly laziness. Whatever. Here’s hoping it won’t be another 20 years before we discover MDL yet again. It remains to be seen whether we can cobble together a better fitting shoe[NH3] .





[1]Actually ternary if we assume that any of the 3 aux forms can go in each position.
[2]But maybe not so much.  One might rightly ask whether there could be a universal representation or coding method that could cover any computational agent’s representational powers – Martian or otherwise. The answer is, “Yes, but.”  Solomonoff  (1960) was the first to show that a universal prior does exist – but that it is noncomputable. Yet this too can be fixed. For additional discussion, see Berwick, 1982; 1985; and Li & Vitányi, 2008.
[3]Contemporary because, as Kiparsky (2007) points out, the Pāṇini grammarians (400 BCE) also developed a simplicity principle for their version of generative grammar based on the idea that a rule should only be posited when it “saves” more feature specifications than it “costs” to memorize a form literally. A special thanks to Ezer Rasin for pointing me to Kiparsky’s recent lecture about this, here:
 http://web.stanford.edu/~kiparsky/Papers/paris.pdf
[4]But this can be done; exactly this measure has been used by computer science theorists; see, e.g., Ginsberg & Lynch, 1977. Derivation complexity in context-free grammar forms. SIAM J. Computing, 6:1, 123-138.  I talk about this a bit more in my thesis (Berwick, 1982; 1985).
[5]The Russian mathematician Kolmogorov independently arrived at the same formulation in the early 1960s, as did Chaitin (1969). See Li and Vitányi, Introduction to Kolmogorov Complexity and its Applications, Springer, 2008, for an excellent exposition of this field.
[6]A bit more carefully, following Shannon’s account (1948), if there are n symbols in our alphabet, and the grammar is m symbols long, transmitting a “message” corresponding to our grammar would take about mlog2 n bits  to “transmit” this grammar to somebody else
[7]This is not as easy to do as it might appear – it took Solomonoff several tries.  See Li and Vitányi 2008, pp. 383-389 for details about this.
[8]More properly, the we ignore the rewrite arrow and any details about the coding length for different symbols, Optimal coding schemes would alter this slightly.
[9]Again, this is not as straightforward; see Li & Vitanyí, pp. 383-389 for a discussion of the additional conditions that must be imposed in order to ensure the interchangeability between Kolmogorov complexity and MDL.


 [NH1]I assume that there are better and worse ways to do this as well? So more compact descriptions of the data are better?  Or this not so? Maybe an example wrt aux system would help the illiterate like me.
 [NH2]So best compresses the PLD and the G?
 [NH3]Really nice. You might want to do a follow up showing how some of this worked in an example or two. It’s really neat stuff. Excellent.