Comments

Thursday, July 18, 2013

Why Morphology?


For quite a while now, I’ve been wondering why natural language (NL) has so much morphology.  In fact, if one thinks about morphology with a minimalist mind set one gets to feeling a little like I. I. Rabi did regarding the discovery of the muon. His reaction? “Who ordered that?”.  So too with morphology; what’s it doing and why is there both so much of it in some NLs (Amer-Indian languages) and so little of it in others (Chinese, English)? One thing seems certain, look around and you can hardly miss the fact that this is a characteristic feature of NLs in spades!!

So what’s so puzzling? Two things. First, it’s absent from artificial languages, in contrast to, say, unbounded hierarchy and long distance dependency (think operator-variable binding). Second, it’s not obviously functionally necessary (say to facilitate comprehension). For example, there is no obvious reason to think that Chinese or English speakers (there is comparatively little morphology here) have more difficulty communicating with one another than do Georgian or Athabaskan speakers, despite the comparative dearth of apparent morphology. In sum, morphology does not appear to be conceptually or functionally necessary for otherwise we (I?) might have expected it to be even more prevalent than it is. After all if it’s really critical and/or functionally useful then one might expect it to be everywhere, even in our constructed artificial languages.  Nonetheless, it’s pretty clear that NLs hardly shy away from morphological complexity.

Moreover, it appears that kids have relatively little problem tracking it. I have been told that whereas LADs (language acquisition devices, aka: kids) omit morphology in the early stages of acquisition (e.g. ‘He go’), they don’t produce illicit “positive” combinations (e.g. ‘They leaves’). I have even been told that this holds true for languages with rich determiner systems and noun classes and fancy intricate verbal morphology: it seems that kids are very good at correctly classifying these and quickly master the relevant morphological paradigms.  So, LADs (and LASs; Language Acquisition Systems) are good at learning these horrifying (to an outsider or second language learner) details and at deploying them effectively as native speakers. So, again, why morphology?

Unburdened by any knowledge of the subject matter, I can think of four possible reasons for morphology’s ubiquity within NLs. I should add that what follows is entirely speculative and I hope that this post motivates others to speculate as well. I would love to have some ideas to chase down. So here goes.

The first possibility is that visible morphology is a surface manifestation of a deeper underlying morphology. This is a pretty standard Generative assumption going back to the heyday of comparative syntax in the early 80s. The first version of this was Jean-Roger Vergnaud’s (terrific) theory of abstract case. The key idea is that all languages have an underlying abstract case system that regulates the distribution of nominal expressions.  If we further assume that this abstract system can be phonetically externalized, then the seeds of visible morphology are inherent in the fundamental structure of FL. The general principle then is that abstract morphemes (provided by UG) are wont to find phonetic expression (are mapped to the sensory and motor systems (S&M)), at least some of the time.

This idea has been developed repeatedly. In fact, the following is still a not an unheard of move: We find property P in grammar G overtly, we then assume that something similar occurs in all Gs, at least covertly. This move is particularly reasonable in the context of “Greed” based grammars characteristic of early minimalism. If all operations are “forced” and the force reduces to checking abstract features, then using the logic of abstract case theory, we should not be surprised if a GL expresses these phonetically.

Note that if something like this is correct (observe the if), then the existence of overt morphology is not particularly surprising, though the question remains of why some Gs externalize these abstracta and some remain phonetically more mum.  Of late, however, this Greed based approach has dimmed somewhat (or at least that’s my impression) and generate and filter models of various kinds are again being actively pursued. So…

A second way to try and explain morphology piggy-backs on Chomsky’s recent claims that Gs are not pairings of sound and meaning but pairings of meanings with sound. His general idea is that whereas the mapping from lexical selection to CI is neat and pretty, the externalization to the S&M systems is less straightforward. This comports with the view that the first real payoff to the emergence of grammar was not an enhancement of communication but a conceptual boost expanding the range of cognitive computations in the individual, i.e. thinking and planning (see here). Thus externalization via S&M is a late add-on to an already developed system. This “extra” might have required some tinkering to allow it to hook onto the main lexicon-to-CI system and that tinkering is manifest as morphology. In effect then, morphology is family related to Chomsky and Halle’s old readjustment rules.  From where I sit, some of the work in Distributed Morphology might be understood in this way (it packages the syntax in ways palpable to S&M), though, I am really no expert in these matters so beware anything I say about the topic. At any rate, this could be a second source for morphology, a kluge to get Gs to “talk.”

I can think of a third reason for overt morphology that is at right angles to these sorts of more grammatically based considerations. There are really two big facts about human linguistic facility: (i) the presence of unbounded hierarchical Gs and (ii) the huge vocabulary speakers have.  Though it’s nice to be able to embed, it’s also nice to have lots of words.  Indeed, if travelling to a foreign venue where residents speak V and given the choice of 25 words of V plus all of GV or 25,000 words of V plus just the grammar of simple declaratives (and maybe some questions), I’d choose the second over the first hands down. You can get a lot of distance on a crappy grammar (even no grammar) and a large vocabulary.  So, here’s the thought: might morphology facilitate vocabulary development?  Building a lexicon is tough (and important) and we do it rapidly, very rapidly. Might overt morphology aid this process, especially if word order in a given language (and hence PLD of that language) is not all that rigid?  It could aid this process by providing stable landmarks near which content words could be found. If transitional probabilities are a tool for breaking into language (and the speech stream, as Chomsky proposed in LSLT and later rediscovered by Aislin, Saffran and Newport), then having morphological landmarks that probabilistically vary at different rates than the expressions that sit within these landmarks then it might serve to focus LADs and LASs on the stuff that needs learning; content words. On this story, morphology exists to make word learning easier by providing frames within a sentence for the all-important lexical content material.

There is a second version of this kind of story that I would like to end with. I should warn you that it is a little involved. Here goes. Chomsky has long identified two surprising properties of NLs. The first is unbounded hierarchical recursion, the second is our lexical profligacy. We not only can combine words but we have lots of words to combine. A typical vocabulary is in the 50,000 word range (depending on how one counts). How do we do this. Well, assume that at the very least, each new vocabulary item consists of some kind of tag (i.e. a sound or a hand gesture). In fact, for simplicity say that acquiring a word is simply tagging it (this is Quin’s “museum myth,” which like many myths may in fact be true). Now this sounds like it should be fairly easy, but is it?  Consider manufacturing 50,000 semantically arbitrary tags (remember, words don’t sound the way they do because they mean what the do, or vice versa).  This is hard. To do this effectively requires a combinatoric system, Indeed, something very like a phonology, which is able to combine atomic units into lexical complexes. So, assume that to have a large lexicon we need something like a combinatoric phonology and the products of this system are the atoms that the syntax combines into further hierarchically structured complexes. Here’s the idea: morphology mediates the interactions of these two very different combinatoric systems.  Meshing word structures and sentence structure is hard because the modes of combination of the two kinds of systems are different. Both kinds play crucial (and distinctive) roles in NL and when they combine morphology happens!  So, on this conception, morphology is not for lexical acquisition, but exists to allow words with their structures to combine into phrases with their structures.

The four speculations above are, to repeat, all very speculative and very inchoate. They don’t appear to be mutually inconsistent, but this may be because they are so lightly sketched. The stories are most likely naïve, especially so given my virtually complete ignorance of morphology and its intricacies. I invite those of you who know something about morphology to weigh in. I’d love to have even a cursory answer to the question.

Tuesday, July 16, 2013

The Evolution of Complexity

There seems to be some discussion in the bio literature about the sources of "complexity." A standard position is that Natural Selection is required for complexity to arise. This is roughly Dawkin's view.  This article discusses some other non-selection sources for complexity's evolution.  I suspect that this might be of interest for linguists interested in the question of how language might have evolved. Enjoy!

Sunday, July 14, 2013

Levels of Adequacy in Generative Grammar


REVISED July 15/2013

First off, thanks to Ben, David and Andrew for correcting me. The discussion is not in Aspects as stated below (or at least in the way I recalled: but see chapter 4 for some discussion), but in Current Issues in Linguistic Theory. So don't fret if in your (re)reading of Aspects chapter 1 you miss it.  Ben also asked that I expand the discussion a bit of the relation between observational and descriptive adequacy. I have added something at the end to expatiate a bit on the distinction. Thanks to Avery for adding to this. I encourage people to look at the comment section for further enlightenment.

****

In the last several posts I have defended the position that one of the benchmarks for theory evaluation ought to be how well a given proposal helps answer a central why question. Plato’s Problem lays out the central puzzle in GB, and it is worth reviewing how Generative Grammar (GG) incorporated concern for both low-level observational details and higher level problems in evaluating ongoing proposals. The locus classicus is, of course, in Aspects chapter 1.

BTW, before going on, I found to my dismay that students no longer read Chapter 1 as a matter of course. In teaching intro to Minimalism here at the LSA summer institute, I ran a random poll of the class and discovered that much less than a third had ever read it.  Given that this is one of the seminal documents in GG and lays out many of the basic foundational premises as clearly as has ever been done, this is just nuts. So, if you dear reader have never read it (in fact multiple times (aim for about 50)), then stop reading this silly blog now and go and read it!!

Ok, to resume: In Aspects Chomsky distinguished three levels of empirical adequacy:

(i)             Observational Adequacy
(ii)           Descriptive Adequacy
(iii)          Explanatory Adequacy

In the context of Plato’s Problem (PP) they were interpreted as follows: a proposed grammar G of a language L (i.e. GL) is observationally adequate if it covers the relevant data points (generally consisting of acceptability(-under-an-interpretation) judgments). A GL is descriptively adequate if it accurately describes the internal (steady) state of a native speaker of L. Last, a theory of UG is explanatorily adequate when it, in combination with the PLDL (the primary linguistic data of L that the child uses to attain GL) can derive the descriptively adequate GL. We can then extend the notion of explanatory adequacy to GL as well: the proposal is explanatory just in case it covers the relevant observational data and is derivable from UG given PLDL.

Note the important distinction between observational and descriptive adequacy. The evaluation of GLs faces in two directions: towards the sentential observations/judgments that derive from descriptively adequate Gs and towards FL from which these Gs descend.  Thus, descriptive adequacy commits hostages to the structure of FL for a proposal fails to be such if it fails to grow out of FL when appropriately watered with PLDL. In other words, covering the sample data was never considered in itself sufficient for attaining descriptive adequacy. This required attention to all three levels, i.e. to cover the data with the “correct” grammar derived from an adequate UG.

I always thought that one could add a forth level to this, doing for UG what (ii) does for particular Gs:

(iv)          Explanatory Descriptive Adequacy

This last would select out the actual UG we have, not merely one capable of deriving GL from PLDL + some UG.  However, given how hard the problem of finding any explanatorily adequate UG, it makes sense to be satisfied with theories of UG that can manage (iii) (to my knowledge, none has yet been produced, though there are sketches). I return to (iv) below.

Not surprisingly, within the context of PP all three levels play an important justificatory role for particular analyses. Thus, an account that fails to “cover” the relevant judgment data, is less well regarded than one that does. So too, if a particular story covers these data but does not describe our actual G, is justly discarded.  One relevant question is how one might know that an account that is observationally adequate is nonetheless possibly false.  One way is that it fails to comport with what we take to be features of UG (others include being “too complex and inelegant,” “missing a generalization,” etc.[1]).  So, for example, we might find a theory that covers the relevant language particular data but invokes operations that are UG inconsistent, something we discerned in studying the Gs of other Ls.  As we assume that any given Gs must be structurally consistent with other Gs (e.g. the G of English can constrain what we take to be a possible rule in the G of Japanese), we can argue that Gs that are observationally adequate (i.e. cover the relevant facts within any given L) are nonetheless inadequate as they invoke operations or principles inconsistent with what we find in other Gs.  As any child, we assume, can learn any language, the G that it actually (accidentally) acquires should look similar to the other Gs that it could have acquired. In this way, the larger explanatory goal, discovering the structure of FL/UG, impacts our evaluation of candidate proposals.

A similar three level scheme with a regulative role for Darwin’s Problem can be developed for the Minimalist Program (MP). Recall that the aim is to understand why we have the FL/UG we have and not some other. So a MP proposal will be observationally adequate if it replicates the right “laws of grammar.” (e.g. the principles within GB are “roughly” observationally adequate). An MP proposal will be descriptively adequate if they describe the actual FL/UG that is our biological endowment and a theory will be explanatorily adequate if combined with a plausible evolutionary scenario it can account for how the descriptively adequate FL/UG arose in the species. I have elsewhere discussed various intermediate projects that could push this explanatory project forward. Here I just want to note that as in the discussion above, it is useful to have the higher- level explanatory goal regulate the evaluation of specific proposals.

I hope it goes without saying that Plato’s and Darwin’s Problems are pursued in tandem and not seriatim. Both are looking for the correct FL/UG. It’s a factor in evaluating descriptive adequacy in the context of Plato’s Problem and it is the object of inquiry in pursuing Darwin’s.  Observational concerns affect explanatory ambitions and, hopefully, the reverse is true as well.  Thus, the projects are not separate, but go hand in hand.  Nonetheless, it is interesting to note that as Chomsky likes to say “from the earliest days of Generative Grammar” between observational work and higher level questions that this work was in the service of pursuing. Thus, the minimalist emphasis on explanation and the insistence that it play a central role in theory evaluation is hardly a GG novelty. It was always such, and now it is such in spades.

Addendum:

Here's how I have understand the distinction between observational and descriptive adequacy. Chomsky draws an implicit distinction between grammars that cover the relevant data of interest and grammars that accurately describe a speaker's mental grammatical state.  The former grammars are observationally adequate, the latter descriptively adequate.  Now, as Avery nicely observes, there are even levels of observational adequacy: does it cover the right range of acceptability judgments, get the right interpretations (theta roles, scope, binding), does it "capture" the apparent generalizations. Grammars that do so are observationally adequate. Note, this is not a trivial hurdle.  A good observationally adequate grammar is nothing to sneeze at.  Even more so, when as I suggested, that we extend the same courtesy to UGs in the context of MP. However, being observationally adequate is one virtue, being descriptively adequate another. A Grammar of L attains descriptive adequacy if it correctly describes the mental state achieved by a competent speaker of L (actually an ideal speaker-hearer, but let's put that aside).  What makes for the gap between observational and descriptive adequacy?  Well, being acquirable as a product of FL.  Descriptively adequate grammars are those that FL/UG generate on the basis of PLD of L.  So, among the possible observationally adequate grammars, the descriptively adequate ones are those that are products of FL. Thus descriptive adequacy adverts to the concerns of Plato's Problem, a sketch of the latter bearing (at least implicitly) on the step from observational to descriptive adequacy.

If this is correct, then from the earliest days of GG, Chomsky has insisted placing the evaluation of grammatical proposals in the widest possible context, with Plato's Problem (a later dubbing) playing a central role. This role became more prominent, in my view, as we made more progress. However, it was there from the start.  Minimalism has enriched this evaluative measure yet again by highlighting various other factors (e.g. Ockham, Darwin's Problem, etc.) and including them as relevant measures of theory evaluation. Of course, there is no algorithm for weighing these factors and trading them off against one another when they conflict. But that's par for the course.  That's a domain for scientific judgment, not mindless rules of method. However, it is easy to ignore the more abstract concerns on favor of covering he more easily available data points. The discussion of levels of adequacy is a useful conceptual reminder that the fish we are trying to fry is a very big one and that there are many things that go into frying it.




[1] For a classic argument like this see Chomsky’s discussion of the inadequacy of Phrase Structure Grammars as one’s that clearly are too cumbersome, redundant and miss obvious generalizations. In effect, they fail to be compact enough.

Saturday, July 13, 2013

Are Minds in Brains?

Take a look at this wild article (here). It discusses a recent macabre experiment that involves worm head decapitations. Worms are able to regenerate their heads, so this is not as dire as what happened to Marie Antoinette, but still it's no picnic I would think. Nonetheless, something really weird happened. When the head and brain grew back it grew back with all its former memories intact. Yes, cutting off the worm brain did NOT eliminate its prior knowledge. Not surprisingly, there are more guesses about how this might be the case than there are serious theories.  However, one seems quite interesting (at least to me) given prior (almost groundless, but not quite groundless) speculation about where computing happens in the brain (see here and here).  The idea gets mooted that the information might be stored "in cells that are outside the brain."  The exciting way of reading this is that memories are actually stored in cells/neurons rather than as in weighted networks of cells. I am sure that I am over-reading this, but it does seem interesting that cutting off something's brain seems to leave memories intact especially if memories are stored not in neurons but in the weighted connections between neurons. Of course, one weird fact does not a revolution make, but I think that this sort of thing serves to remind us how little we know about how mental phenomena are incarnated. Brains? Nervous Systems? Recall the ancients put in a plug for hearts and livers! Stay tuned.

Tuesday, July 9, 2013

A further note on falsificationism


This note is spurred by some points made by Alex Clark (here):

Every scientific theory aims for truth. This is not the exclusive goal of falsificationism. However, finite beings that we are we cannot gaze directly at a theory and see if it is true or not. Hence we look for marks or signature properties of truth and falsity. Falsificationism puts great weight on one of these features and lesser weight on others. The first virtue, and not first among equals, is "getting the facts right." Non-naive falsificationists (e.g. Lakatos) then gussy this all up to the point that the method turns out to be "do the best you can realizing that it's a complex mess." This, of course, is a nice version of Feyerabend's "anything goes" but said more politely. I personally like Percy Bridgeman's version: "use your noodle and no holds bared." The net effect of these methodological nostrums is to make them impossible to apply qua method for there is nothing methodical about them. Given the (apparently) insatiable desire for mechanistic answers to complex issues, people fall back onto "covering the data" as the best applicable proxy. And this is where the problem arises.

The problem is two fold: (i) what the relevant data are is not self evident, (ii) this leads to ignoring all the stuff that makes the non-naïve approaches non-naïve. So, in practice, what we get is a ham handed application of the sophisticated views that, in practice, is just our old friend naïve falsificationism.

Let me say a quick word about (i).  Naivety often begins with a robust sense of what the data is.  As the sophisticates know, this is hardly self-evident. So take an example close to the home of some: does Kobele’s work demonstrate once and for all that natural language grammars are not mildly context sensitive because they fail to display constant growth? I have not witnessed mass recantations of the MCS view from my computational colleagues. No, what I have seen is attempts (actually more often hopes that some attempts would be forthcoming) to reanalyze the morphological and syntactic data (pronounced copies) so that it is defanged. 

Let me add that I am entirely sympathetic to this effort in principle. It’s what we do all the time and it is one way that we try to protect our favored accounts.  We do this not (merely) out of a misplaced love of our own creations (but who doesn’t love her own mental products?) but because often our favorite theories have EO that would be lost were it rendered false. Moreover, the more that would be lost the more rational it is to ask for a high level of proof before we give up on the old theory, and even then we will demand of the replacement that it get the previous theories successes right (usually as limit cases). 

My main point has been that we should extend a similar set of considerations to theories with high EO/low BI.  But in order to do this we must try to keep firmly in mind what the central targets of explanation are. This is why Plato’s and Darwin’s problems are so important. Yes, they are vague and need precisfication. But they are not powerless and can be used to evaluate proposals. They, in other words, provide a second axis of evaluation in addition to “data coverage” for theory evaluation.

Two last points: First, there are other factors in addition to the two mooted above that serve as marks of truth: simplicity, elegance, Occamite stuff, etc.  These have served to advance thinking and are very hard to make precise, hence the injunction to “use your noodle.” Second, arguments gain in persuasiveness the more local they are.  These dicta, though very vague in the abstract are remarkably clear when applied in local circumstances, at least most of the time.  As Alex Drummond has repeatedly noted: it’s often all too easy to recognize a counter-example/bit of recalcitrant relevant data when it comes flying your way. Same with EO, elegance, etc.  In particular contexts, these high-flying notions get tied down and can become useful. The aim of lots of theoretical work is to figure out how to do this. However, not surprisingly, this is more art than technique, and hence the disagreements about what factor should weigh most heavily in deciding which way to turn. My view, is that in this hubbub we should not loose sight of the animating problems and discount them in favor of what is more easily at hand. Further, in my view, this is precisely what falsificationism encourages.

Monday, July 8, 2013

Two Cheat Sheets on Computation for Minimalist Syntacticians

NOTE: Apparently the first posting of this had a broken link making the second paper unavailable. I've fixed these so now both PDFs are downloadable. Sorry.


One of the pleasures of being at the LSA summer institute is the chance to interact with people that you've always wanted to know better. I had a couple of personal targets this summer, one of which being Rick Lewis, a computational psychologist here at U Mich.  I've met Rick once before when he came to UMD to give a talk. However, I really got to know about him through his work. He has become a minor celebrity among the UMD processing crowd (I am an informal syntax theory consultant to this group) and several theses have deployed the processing model that he developed (see here for papers. The 2005 with Vasisth is a minor classic where I come from). The proposal provides an account of the observation going back to Miller and Chomsky regarding the processing difficulty of self-embedded sentences (e.g. That that that Bill kissed is surprising is silly is evident). The same model appears to explain some fascinating data Ted Gibson discovered as well. At any rate, the students at UMD have had a field day exploring, criticizing and extending this proposal and I have had a lot of fun listening to their efforts.

This, however, does not exhaust Rick's interests. He is also one of that rare breed, a psychologist cum computer scientist that thinks that Chomsky should be taken seriously. Moreover, Rick has become interested in Minimalism and the idea that one might try to consider grammatical theories from the view point of computational efficiency. Indeed, he has written a couple of papers on the topic that I asked him if I could post. One is a very useful glossary of terms (here). The other (here) is a more extended disquisition on how one may go about thinking of optimal computation in a minimalist setting. I especially like the distinction between 'simplicity' and 'efficiency' and ways they may be tied together (if at all).  Questions of computational efficiency are both important and, sadly, obscure. These papers may help you get some bearings on these issues. They helped me.

Saturday, July 6, 2013

Why, How and When We Need Cowbell


The comments, especially by always level headed David Pesetsky, have provoked the following restatement of my beef with falsificationism.  Its main defect is that it presents a lopsided and hence counter-productive image of scientific practice. It thus encourages the wrong methodological ideals and this has a baleful effect on the research environment.  How so?

First, falsificationism tacitly assumes that the primary defect a proposal/story/theory can have is failure to cover the data.  In other words, in tacitly presents the view that the primary virtue a theory has is data coverage.  In sophisticated versions, other benchmarks are recognized, however, even non-naïve versions of falsificationism, in virtue of being falsificationist, place primary stress on getting the data points well organized. Everything else is second to this. I don’t buy this. In the sciences in general, and linguistics in particular, the enterprise is animated by larger ‘why’ questions and the goal of a proposal is to answer (or at least address) these.  This means that there are at least two dimensions in proposal evaluation (PE)[1]: (i) how they cover the “facts,” and  (ii) how they advance explanation. Moreover, neither is intrinsically more important than the other, though in some contexts one is weighted higher than the other.  A good project is to try to identify these contexts, at least roughly. However, the null position should be to recognize that both are equally important dimensions of evaluation ceteris paribus.

Second point: in practice, I believe, the two virtues are not genrally equally weighted.  For lots of linguistics (and many outside critics of the field), the second is only weakly attended to.  Many think it meet to devalue a proposal if it does not cover the “relevant” facts. The number of times this criticism has been levied with approval are numerous.  Far rarer are the times when suspicion is cast on a proposal because it entirely misses the explanatory boat.  However, and here is my main point, failure to advance the explanatory agenda is just as problematic as failure to cover some data points.  In fact, in practice, it is often easier to evaluate if a particular story has any Explanatory Oomph (EO) than to determine whether the data points not covered are actually relevant. We all agree (or should) that only data points worth covering are the relevant ones and that determining what’s relevant is no simple matter. However, we often act as if it is clear what the relevant data is whereas what the explanatory goal is we take to be irremediably obscure.  I don’t agree.

The goal of early syntactic theory was to explain why grammatically competent humans were able to use and understand sentences never before encountered. Answer: they had internalized generative grammars composed of recursive rules. The fine structure of possible grammars were investigated and some consensus was reached on what kinds of rules they deployed and constraints they obeyed.  This set up the next question (Plato’s Problems): what allows humans to develop generative grammars? And we answered this by attributing to human minds an FL and the project was to describe it.  We concluded that an FL with a principles and parameters architecture in which parameter values are set by PLD would explain why we are linguistically capable, i.e. a description of how this is done, answers the ‘why’ question.[2] The current Minimalist project, I have argued, rests on a similar ‘why’ question: why do we have the FL we have and not some other conceivable kind?  And we are looking for an answer here too. Roughly we are betting on the following kind of story being right: FL is a congery of domain general powers (aka operations and principles) with a small dollop of linguistic specificity thrown in.  If we can theoretically actualize this sort of picture we will have answered Darwin’s Problem, yet another ‘why’ question.  My modest proposal: part of proposal evaluation should involve seeing how a particular story helps us answer these questions.

Let me go further. There are times to emphasize one of the two criteria over the other.  A good time to have high regard for (ii) is when a program is starting out. The goal of the early stages of inquiry into a new question is to develop a body of doctrine (BOD) and this in practice requires putting some recalcitrant facts to the side. One develops a BOD by showing what a proposal buys you, in particular how, if correct, it can address an animating ‘why’ question. When there is some BOD with some EO in place, empirical coverage becomes crucial, as this is how we refine and choose between the many basic approaches all of which are of the right kind to answer the motivating ‘why’ question.  Of course, both activities go on at the same time. It’s not like for the first 10 years we value (ii) and ignore (i) and vice versa for the second ten.  Research is not so discrete.  However, there are times when answers to (ii) are hard to come by and at these times valuing PEs that potentially meet these kinds of demands is, I would argue, very very advisable. Again, the trouble with falsificationsim is that it encourages a set of attitudes that devalue the virtue of (ii).

Where does this leave us/me?  I agree with David Pesetsky that there are different levels of ‘why’ questions, and that they can be pursued in parallel. Addressing one does not preclude addressing another. I also agree that we pursue these higher-level questions by making proposals of how things work.  I also endorse the view that how and why are intimately intertwined. However, I suspect that there might be a disagreement of emphasis: I think that whereas we both value how “data” allows us to develop and judge our proposals, we don’t equally weight the impact of EO.  David has no problem relegating the big ‘why’ questions to “the grand scheme of things” making it sound like some far off fairyland (like Keyne’s long run, it’s where we are all dead?).  True, David mitigates this by adding the qualifier that we should not try and live there “all the time,” suggesting that occasional daydreaming is fine.  However, my point is that even in the “more humble scheme of things” when we work on detailed analyses of specific phenomena, indeed when we muddle along finding “semi-organized piles of semi-analyzed, often accidental discoveries,” even then we should try to keep our eyes on the explanatory prize and ask how what we are doing bears on these animating questions.  Why? Because they are important in evaluating what you are doing no less than seeing if the story covers some set of forms/sentences in some paradigm.  Both are critical, though to my eyes only one (i.e. (i) above) is uncontroversially valued and considered part of everyone’s every day research basket of values. This is what I want my rejection of falsificationsim to call into question as it is something that even sophisticated versions (which of course are correct if sophisticated in the right ways), tend to still operationally relegate to a secondary position.




[1] I would normally use ‘theory’ in place of ‘proposal’ but it sounds too grand. I think that non encompassing self-perceived smaller scale projects are subject to the same dual evaluation streams.
[2] Let me add before I am inundated by misplaced comments that I do not believe that we have “solved” Plato’s Problem.  I have written about this elsewhere.