Comments

Showing posts with label explanation. Show all posts
Showing posts with label explanation. Show all posts

Thursday, January 15, 2015

Truth and beauty

I should have been warned off of Aeon because of the Evans’ fiasco, but the magazine is not really all that culpable. Its content just reflects the stuff that is “out there” and the editors seem to largely act to make topical stuff available for easy reading in one place. This is how I came to read this. It just popped into my mailbox recently one morning and I couldn’t resist.  It’s about the relation between truth and beauty, or more accurately, how and whether aesthetic judgments about a theory are or should be taken to be marks of its truth.  The author Philip Ball thinks not. I’m not sure that I agree, but it got me thinking, so I thought I would try and write something about it. I’ve done this before (here) and some of what I say rehearses earlier themes. Here it is.

Mr. Ball notes that many scientists (Einstein, Dirac, Weyl (three real biggies), Greene and Arkani-Hamed are mentioned by name) have believed that theoretical beauty is a mark of truth. Some can be quoted as saying that theoretical beauty even trumps empirical verification. Mr. Ball takes these quotes at face value. Greene suggests that what Einstein and Weyl (the two culprits) intended is that beauty should count as having some evaluative power  re truth but that “Ultimately, theories are judged how they fare when faced with cold, hard, experimental facts” (p. 3 Ball quoting Greene).[1] I’m not sure what Einstein or Weyl meant. But I personally am interested in how beauty is relevant in theory evaluation, if it is relevant at all. This is, after all, a theme talked up in early minimalism, which I personally found useful. So is beauty linked to truth? And if so, how?

Before approaching these questions, it is worth trying to pin down what it is: what are the marks of theoretical beauty?  As Ball notes, scientists themselves are all over the place on this. For some, it’s like porn to supreme court justices (viz. know it when they see it). But Ball’s essay contains at least one attempt at a specification. Arkani-Hamed (AH) (he is, btw, an important young hot shot (here)) associates theoretical beauty with a certain kind of inevitability:

“There are very few principles and there’s no other possible way they could work once you understand them deeply enough” (p. 4, Ball quoting Arkani-Hamed).

Note three things about AH’s proposal: beauty is the combination of (i) simplicity of the Occam variety (few principles) coupled with (ii) a certain kind of modality (no other possible way they could work) and (iii) is not surface visible in that it takes intellectual effort to perceive (“once you understand them deeply enough”). The third is of particular interest wrt Ball’s many comments on the issue, for much of his discussion, I believe, revolves around the fact that beauty is hard to discern, that there can be lots of disagreement and that it is elusive.  All fair points, but not really immediately relevant to AH’s point.  AH’s view seems to be that finding beauty requires hard intellectual work. It may be the eye of the beholder, but only an eye that’s put in the hours to train itself to see properly. So beauty need not be (and generally won’t be) immediately evident on inspection. Consequently, the fact that what’s beautiful is not immediately discernable or that it is sensitive to cultural mores does not seem all that relevant to AH’s point. That said, let’s concentrate on points (i) and (ii): simple and inevitable.

Occam’s razor is now part of the accepted methodological wisdom. All things being equal (ATE), simpler theories are better than more complex ones. And how is “simple” measured? By the number of axioms or assumptions (how to enumerate these is not trivial but I will put that to one side). So a theory with three axioms trumps one with four and one with four beats one with five ATE.  As Ball notes, things are seldom equal, so applying Occam in the wild is never easy. But, it is a principle that all are ready to accept in some form and we should ask why?  Why is simpler better, or, more specifically, truer?

In the good old 16th and 17th centuries (if not earlier) there was theological justification for this assumption. We live in God’s universe governed by his/her laws. God is an elegant thinker and would never do in a complicated manner what could be done simply (Why not? Is God lazy?).[2] Thus, God’s character guarantees a universe governed by simple laws.[3]

Many of us find this justification hard to accept nowadays, but luckily for us, there are other ways of buying into this viewpoint. Thus, simpler theories are generally more epistemologically solid than are profligate ones. Think of a stool carrying a weight. If three legged, each leg supports 1/3 the load. If four legged, each supports 1/4. Take weight as evidence and legs as axioms and you can see how fewer axioms means greater evidence for each the fewer there are. Bayesians can even formalize this insight, and they have tended to make a big deal out if it. So, something like Occam is something even the theologically fussy can sign onto as a virtue of theory

Improtantly, for my purposes, this epistemological virtue carries metaphysical kudos. If we assume that theories that are better empirically supported are more likely to be true, then simpler theories ATE are more likely to be true. I’m not sure why we think that better supported theories are more likely to be true, but we do think this and this provides at least one link between theoretical simplicity and truth.

Of course, things are much much more complicated. Comparing different theories can be very hard. Indeed, there is no general recipe for how to compare different axioms, definitions, primitives etc. And there seems to be no absolute measure of “simplicity,” or at lest none that garners general assent. But, though there is no general measure, there are lots of local ones and the principle can and does have teeth in many places. Examples: why assume an aether if all goes well without it? Why assume two conceptions of mass if you can get away with one? And, most relevant to us: Why assume internal a syntactic theory with levels if they are not required empirically? Why assume three rule types if one suffices? We all know the drill, and a good one it is, for this kind of simplicity ties elegance together with a notion of “independent evidence,” and every scientist worth her/his salt loves independent evidence (consilience rules!!!).

So, we value simple theories as pointing towards truth and if simplicity is part of beauty, this partly partly explains why we value beauty if it can be had (which, as Ball notes, is not always clear given that things are seldom equal). 

There is a further reason to value simple theories: they are easier to explore. Simple theories are often readily intelligible (as things with fewer moving parts often are). So, methodologically, if one’s aim is to find out what’s what then using an instrument (theory) that is understandable (simple) to probe things is ATE better than one that is opaque.  Simple ideas are generally more manageable, so they enjoy a kind of methodological or epistemological advantage. But note, this does not imply that they are more often true, though perhaps (weasel word!) a case can be made that they are good ways at getting to truth. If so, beauty may not mark truth but it may be the best road to it.

There is still another way to link simplicity and truth; via explanation. Here’s what I have in mind. What’s the “ugliest” possible theory of anything? I would say it’s a list. Why? Because it has zero explanatory value.[4] The virtue of simple theories is that they seem to carry more explanatory oomph.  And what better mark of a theory’s truth than its explanatory power? What we want out of a theory is not only that it cover the data, but that it cover it in such as way as to explain why the data is the way it is.

This locution, “why the data is the way it is” has two related sub-parts with slightly different foci. It means explaining not only why we have the data we happen to have, but also the data we will have and could have. And it means explaining why we don’t have other than the data we have (viz. why the data we don’t see is missing).[5]  The main problem with a list, then, is that it does not tell you how to expand itself so as to include what it must and exclude what it must. A simple theory, being general, does. Simple theories can explain and part of being beautiful is having explanatory oomph, which modally relates to what is possible.[6]

Note the explanatory property of simple theories already introduces the modal feature of beauty that AH noted. Why X? Because X is the only way things could have been. Theories explain the actual in terms of the possible.  This suggests the following thought: what makes a theory truly beautiful is that it perfectly explains the actual in terms of the possible. Or, put another way, in the best case, the fit between what is possible and what is actual is perfect. All that can be observed has been and all that has been observed can be. There is no spillover.[7]

A joke I’ve told before (see here) illustrates the logic, I think: why are there 1-hump camels and 2-hump camels but not 3,4…N-humped camels? Because there are concave camels and convex camels and that’s all there is.  The “joke” illustrates how changing our conception of humps (look at the curves not the humps) allows us to exhaust the kinds of humps+valley combos we expect to find. Looking at the number of humps invites the question of why are there at most 2. Why 2, and not 3,4,…N? The stopping point seems arbitrary. Looking at the curves and seeing them as concave and convex seems to exhaust the space of options (simple curves being one or the other). And this restricted space also coincides with what we find. The possible hump+valley options perfectly fit the attested realizations (viz. 2 and only 2). Thus, once we think in terms of curves, it seems clear that things could not have been otherwise hump/camel-wise (but see comments to above link for problems with this “joke”)..

Now, before I get lots of comments taking the “joke” apart (again, see earlier post), let me try to illustrate what I have in mind with a more relevant syntactic example (there are several others in the earlier post). Chomsky has famously argued that properly understood Merge yields both phrase structure and movement.  What is the proper way to understand it? Well, merge is an operation that takes two linguistic items A and B and puts them together (forms a set). What As and Bs? Chomsky says that there are two possibilities: (i) neither A nor B contains the other or (ii) one of A or B contains the other. In the first case, Merge(A,B) yields a constituent like {A,B}. In the other it yields a constituent like {B,{A… B}}. Thus, if (i) and (ii) exhaust the options then it looks like every possible instance of a very simple combination rule (merge) results in just the two products that we find prominently in Gs (viz. products of PS rules and movement rules). What makes Chomsky’s story attractive is that it looks like the two principle properties of Gs (hierarchy and displacement) follow exhaustively from the possible ways two inputs can fall under the rule. Either the to-be-combined form a part whole relation or they do not. Thus, a simple exhaustive account of the options yields (all and only?) the attested structural dependencies.[8]  

I don’t know about you, but if Chomsky is right here, then I consider this to be a very nice kind of account. Why do Gs have both PS rules and movement rules? Because the simplest conception of the combination operation has these two “kinds” of rules (and only these) as consequence.  

There are other examples of this kind of thinking in the syntax literature as I discussed here. Nor is this restricted to syntax. It is also a staple of phonological feature theories, where the aim is to adumbrate all and only the possible linguistic sounds and, if I understand Heinz and Idsardi’s Science paper, the range of possible phonological processes.

I think that there is another way of describing what lends such stories AH’s feeling of inevitability. When properly framed, a theory serves to close off further questions. Here’s what I mean. A really satisfying account brings a kind of question closure with it. The concave/vex theory of camels explains why the question of 3 humped camels is, in some sense, ill-formed. Chomsky’s account of structure dependence discussed here not only removes linear processes from the grammar but explains why they could not have been there. Properly understood, such rules cannot be stated within the theory and that’s why they don’t exist. Such rules don’t merely fail to exist, they cannot exist for they cannot be stated. Thus, they are not merely contingently absent, they are necessarily absent. Indeed, when the theory is properly understood, one sees that their possibility is actually inconceivable. It’s this sense of theoretical beauty that AH’s remark pointed to, I am suggesting.

So, is beauty a mark of truth? I think so. The problem is that which conception of beauty is the right one is generally what is theoretically up for grabs. The theorist’s challenge is to provide beautiful accounts; simple stories that exhaustively and completely adumbrate the options that completely describe what in fact happens. Such stories are simple and exhaustive and hence beautiful.  So, does beauty count? Sure. Is it the only virtue? No. But IMO it is one often worth sacrificing a few data points to. But you knew I would say that, right?



[1] I am not sure what Greene meant here. “Ultimately” is a very long time. As Keynes noted, ultimately (in the long run) we are all dead. I suspect what Greene is doing here is CYAing a bit. Among themselves, scientists like to talk about the intangibles that count towards theory evaluation. To the public, they like to appear hard-headed so that they can crap on the artsy-fartsy emotive types, especially those with religious or spiritual inclinations. Talking about experimental facts serves this aim well.  In truth, everyone wants new data to back up interesting claims and everyone wants theories that are not clunky. How these different features get weighted at any particular time is an art, not the product of algorithm.  So, here Greene, a well known purveyor of beauty in theory, is trying to cover his public posterior. 
[2] In this regard God is the anti-Thomas Mann, of whom Peter Gay said: “Mann did not like to be simple if it was at all possible to be complicated.”
[3] I never quite understood why these thinkers thought they knew God’s aesthetic preferences or work habits. But, it seems that they could and did.
[4] For the statistically inclined compare histograms with the statistics that describe them. The former represents actual data points. It does not explain them in any way. The latter carries some explanatory force as it not only “accounts” for the histograms but (in the best case) describes where a possible data point can and cannot fall.
[5] This has obvious relationship to the issue of negative data near and dear to a linguist’s heart.
[6] The link between explanatorieness and truth gets us into deep waters very quickly: why after all should the universe be comprehensible? Damn if I know. Descarte’s benevolent deity? Darwin? Dumb luck?
[7] Needless to say, this cannot be understood as every possible instance has been observed, rather an instance of every type has been observed and no instance of an impossible type has been. 
[8] I am not sure whether this argument is in fact accurate. Prima facie, there are more kinds of grammatical rules (e.g. construal, agreement, deletion).  My own view is that the right next step is to try and unify construal, etc. with movement in some way. Doing this would then derive all possible G operations from Merge as Chomsky wants to do. However, I believe that my desires here are idiosyncratic outliers theoretically. However, unless this is done, Chomsky’s “derivation” is not complete or exhaustive and so, not “perfect.”

Monday, August 12, 2013

Explaining Camels


There are certain kinds of explanations, which when available, are particularly satisfying.  What makes them such is that they not only explain the facts in front of you, but do so in ways that make the facts inevitable.  How do they do this? Well, one way is by rendering the non-extant alternatives not merely false, but inconceivable.  A joke/riddle that I like to tell my students displays the quality I have in mind.

One physicist/mathematician asks another: why are there 1-humped camels and 2-humped camels, but no N-humped camels, N>2? Answer: Because camels are convex or concave; no other models available.

I love this answer. It’s perfect. How so? By shifting the relevant predicates (from natural numbers to simple curves) the range of possible camels is reduced to two, both of which are attested!  Concave/convex, exhaust the space of options.  And once you think of things in this way, it is clear why 1 and 2 are the only possible values for N.

Let me put this another way: one gets a truly satisfying explanation if one can embed it in concepts that obviate further why questions. How are they obviated? By exhausting the range of possibilities. Why are there no 3-humped camels? Because 3-humped camels are neither concave nor convex and these are the only shapes camels can come in.

The joke/riddle has a second useful attribute. I displays what I take to be a central aim of theoretical research: to redescribe the conceivable alternatives in such as way as to restrict the range of available alternatives to what one actually sees. The aim of theory is not merely to cover the data, but to explain why the data falls in the restricted range it does, and this requires carefully observing what doesn’t happen (what Paul Pietroski calls ‘negative’ facts (e.g. here)).

So, do we have any of these kinds of explanations within syntax? I think we do, or at least there have been attempts to provide such. Let me illustrate.

One current example is Chomsky’s proposed account for why grammatical operations are structure dependent. This is in Problems of Projection (which I would link to, but it is behind a Lingua paywall with an exorbitant price so I suggest that you just get a copy from someplace else). Here’s what we want to explain: given that rules that move T-to-C (as in Y/N questions in English) target the “highest” Ts and not the linearly closest Ts (i.e. leftmost), why must they target these Ts, (i.e. why can’t they target the linearly most proximate Ts)?

The answer that Chomsky gives is that grammatical operations cannot use notions like linear proximity, because linguistic objects are not linearly specified until Spell Out, i.e. the final rule of the syntax. So why can’t grammatical operations be linearly dependent (i.e. non-structure dependent)? Because the syntactic manipulanda (i.e. phrase markers) contain no linear (i.e. left-right order) information. Thus, if grammatical rules manipulate phrase markers and these don’t contain linear information then there is no way to state linearly-dependent rules over these objects.  In other words, why do such rules not appear to exist? Because they can’t be specified for the objects over which the grammatical rules operate, and that’s why linear dependent syntax rules don’t exist.

Or, put positively: why are all syntactic rules structure dependent? Because that’s the only way they can be. In other words, once the impossible options are eliminated all that’s left coincides with what we find.  In this way, the actual is explained via the possible and explanatory oomph is attained. Indeed, I suspect (believe!) that the best way to explain anything is by showing how the plausible alternatives are actually conceptually impossible when thought about in the right way.

Here’s another minimalist example. It comes from current conceptions of Spell Out and how they’ve been used to account for phase impenetrability (i.e. the prohibition against forming dependencies across a phase head).  Here’s the question: why are phases impenetrable? Answer: because Spell Out “sends” phase head complements to the interfaces thereby removing their contents from the purview of the syntax/computational system. In effect, dependencies across phase heads are not possible because complements of phase heads (and hence their contents) are not syntactically “there” to be related.

Note the similarity to the first account: just as linear information is not there to be exploited and hence only structure dependent operations are stateable, so too phase complement information is not there and so it cannot be exploited. In both cases, the “reason” the condition holds is that it cannot fail to hold. There really is only one option when properly conceptualized.

Here’s another example from an earlier era: one of the most interesting arguments in favor of dispensing with constructions as grammatical primitives came from considering a conundrum relating to examples like (1c).

(1)  a. John is likely to have kissed Mary
b. John was seen/believed by Mary
c. John is seen/believed to have kissed Mary

The puzzle is the following in the context of a construction-based conception of grammatical operations (i.e. a view of FL in which the basic operations are construction based rules like Passive, Raising etc).[1] In (1a), John moves from the post verbal position to the subject position via Raising. In (1b) the operation that moves John to the top is Passivization. Question: what is the rule that moves John in (1c)?  Is this Passive or Raising?  There is no determinate answer.

Eliminating constructions gives a simple answer to the otherwise unanswerable question: it’s neither, as these kinds of rules don’t exist. ‘Move alpha’ is the sole transformation and it applies in producing both Raising and Passive constructions. Of course, if this is the only (movement) transformation, then the question is it Raising or Passive dissolves. It’s move alpha and only move alpha. As the earlier question (Raising or Passive?) had no good answer, a conception of grammar where the question dissolves has its charms.

One last example, this one from an undergrad research thesis by Noah Smith (of CMU fame; yes he was once a joint ling/CS student). He wrote when Jason Merchant’s work on ellipsis was first emerging and he asked the following question: Given that Merchant has shown that ellipsis is deletion and not interpretation, why can’t it be the latter?  He rightly (in my view) surmised that this could not be a data driven fact as the relevant data for determining this was very subtle. For Jason it amounted to some case and preposition stranding correlations in sluiced and non-sluiced constructions. On the reasonable assumption that these fall outside the PLD, the fact that ellipsis was deletion could not have been a data driven outcome. Say this is correct. Noah asked why it had to be correct; why ellipsis had to be deletion and could not be interpretation a la Edwin Williams (i.e. ellipsis amounted to filling in the contents of null phrase markers with null terminals at LF).[2] Noah’s answer? Bare Phrase Structure (BPS). BPS replaces the earlier combo of phrase structure + lexical insertion rules. This has the effect of eliminating the distinction between the content and position of a lexical item. As such, he argued, the structures that the interpretive theory of ellipsis presupposed (phrases with no lexical terminals) were conceptually unavailable and this leaves the deletion analysis as the only viable option. So why is ellipsis deletion rather than interpretation? Because the interpretive theories required structures that BPS rendered impossible. I confess to always liking this story.

One caveat before concluding: I am not here proposing that the proposed explanations above are correct. I have some questions regarding the Spell Out explanation of the PIC for example and there are empirical challenges to Merchant’s evidence in favor of deletion analyses of ellipsis. However, the kinds of proposals mentioned above are interesting and important for, if correct, they explain (rather than describe) what we find. And explanation is (or should be) what scientific inquiry aims for.

To end: one of the aims of theoretical work is to find ways of framing questions in such a way that all and only the conceptually possible answers are actualized.  This requires finding a vocabulary that not only accommodates/describes the actual, but renders the non-attested impossible, i.e. unstateable. This makes the consideration of systematic absences (viz. negative data) central to the theoretician’s task. Explanation lies with the dogs that don’t bark, the things that though logically possible, don’t occur. Theories explain what happens in terms of what can happen. This means keeping one’s eyes firmly focused on what we don’t find, the actual simply being the residue once the impossible has been pared away.


[1] This is not the place to go into this, but note that rejecting constructions as grammatical primitives does not imply that constructions might not be derived objects of possible psycholinguistic interest.  I discuss this a bit (here) in the last chapter.
[2] There was an interesting and animated debate about ellipsis between Edwin Williams and Ivan Sag, the former defending an interpretive conception (trees with null terminals filled in at LF) vs the latters deletion analysis (similar to Merchant’s contemporary approach).