Comments

Showing posts with label Pietroski. Show all posts
Showing posts with label Pietroski. Show all posts

Friday, March 21, 2014

Let's pour some oil on the flames: A tale of too simple a story

Olaf K asks in the comments section to this post why I am not impressed with ML accounts of Aux-to-C (AC) in English. Here’s the short answer: proposed “solutions” have misconstrued the problem (both the relevant data and its general shape) and so are largely irrelevant. As this judgment will no doubt seem harsh and “unhelpful” (and probably offend the sensibilities of many (I’m thinking of you GK and BB!!)) I would like to explain why I think that the work as conducted heretofore is not worth the considerable time and effort expended on it. IMO, there is nothing helpful to be said, except maybe STOP!!! Here is the longer story. Readers be warned: this is a long post. So if you want to read it, you might want to get comfortable first.[1]

It’s the best of tales and the worst of tales. What’s ‘it’? The AC story that Chomsky told to explicate the logic of the Poverty of Stimulus (POS) argument.[2] What makes it a great example is its simplicity. To be understood requires no great technical knowledge and so the AC version of the POS is accessible even to those with the barest of abilities to diagram a sentence (a skill no longer imparted in grade school with the demise of Latin).

BTW, I know this from personal experience for I have effectively used AC to illustrate to many undergrads and high school students, to family members and beer swilling companions how looking at the details of English can lead to non-obvious insights into the structure of FL. Thus, AC is a near perfect instrument for initiating curious tyros who into the mysteries of syntax.

Of course, the very simplicity of the argument has its down sides. Jerry Fodor is reputed to have said that all the grief that Chomsky has gotten from “empiricists” dedicated to overturning the POS argument has served him right. That’s what you get (and deserve) for demonstrating the logic of the POS with such a simple straightforward and easily comprehensible case. Of course, what’s a good illustration of the logic of the POS is, at most, the first, not last, word on the issue. And one might have expected professionals interested in the problem to have worked on more than the simple toy presentation. But, one would have been wrong. The toy case, perfectly suitable for illustration of the logic, seems to have completely enchanted the professionals and this is what critics have trained their powerful learning theories on. Moreover, treating this simple example as constituting the “hard” case (rather than a simple illustration), the professionals have repeatedly declared victory over the POS and have confidently concluded that (at most) “simple” learning biases are all we need to acquire Gs. In other words, the toy case that Chomsky used to illustrate the logic of the POS to the uninitiated has become the hard case whose solution would prove rationalist claims about the structure of FL intellectually groundless (if not senseless and bankrupt).

That seems to be the state of play today (as, for example, rehearsed in the comments section of this). This despite the fact that there have been repeated attempts (see here) to explicate the POS logic of the AC argument more fully. That said, let’s run the course one more time. Why? Because, surprisingly, though the AC case is the relatively simple tip of a really massive POS iceberg (c.f. Colin Phillips’ comments here March 19 at 3;47), even this toy case has NOT BEEN ADEQUATELY ADDRESSED BY ITS CRITICS! (see. In particular BPYC dhere for the inadequacies).  Let me elaborate by considering what makes the simple story simple and how we might want to round it out for professional consideration.

The AC story goes as follows. We note, first, that AC is a rule of English G. It does not hold in all Gs. Thus we cannot assume that the AC is part of FL/UG, i.e. it must be learned. Ok, how would AC be learned, viz: What is the relevant PLD? Here’s one obvious thing that comes to mind: kids learn the rule by considering its sentential products.[3] What are these? In the simplest case polar questions like those in (1) and their relation to appropriate answers like (2):

(1)  a. Can John run
b. Will Mary sing
c. Is Ruth going home

(2)  a. John can run
b. Mary will sing
c. Ruth is going home

From these the following rule comes to mind:

(3)  To form a polar question: Move the auxiliary to the front. The answer to a polar question is the declarative sentence that results from undoing this movement.[4]

The next step is to complicate matters a tad and ask how well (3) generalizes to other cases, say like those in (4):

(4)  John might say that Bill is leaving

The answer is “not that well.” Why? The pesky ‘the’ in (3). In (4), there is a pair of potentially moveable Auxs and so (3) is inoperative as written. The following fix is then considered:

            (3’) Move the Aux closest to the front to the front.

This serves to disambiguate which Aux to target in (4) and we can go on. As you all no doubt know, the next question is where the fun begins: what does “closest” mean? How do we measure distance? It can have a linear interpretation: the “leftmost” Aux and, with a little bit of grammatical analysis, we see that it can have a hierarchical interpretation: the “highest” Aux. And now the illustration of the POS logic begins: the data in (1), (2) and (4) cannot choose between these options. If this is representative of what there is in the PLD relevant to AC, then the data accessible to the child cannot choose between (3’) where ‘closest’ means ‘leftmost’ and (3’) where ‘closest’ means ‘highest.’ And this, of course, raises the question of whether there is any fact of the matter here. There is, as the data in (5) shows:

(5)  a. The man who is sleeping is happy
b. Is the man who is sleeping happy
c. *Is the man who sleeping is happy

The fact is that we cannot form a polar question like (5c) to which (5a) is the answer and we can form one like (5b) to which (5a) is the answer. This argues for ‘closest’ meaning ‘highest.’ And so, the rule of AC in English is “structure” dependent (as opposed to “linear” dependent) in the simple sense of ‘closest’ being stated in hierarchical, rather than linear, terms.

Furthermore, choice of the hierarchical conception of (3’) is not and cannot be based on the evidence if the examples above are characteristic of the PLD. More specifically, unless examples like (5) are part of the PLD it is unclear how we might distinguish the two options, and we have every reason to think (e.g. based on Childes searches) that sentences like (5b,c) are not part of the PLD. And, if this is all correct, then we have reason for thinking that: (i) that a rule like AC exists in English and whose properties are in part a product of the PLD we find in English (as opposed to Brazilian Portuguese, say) (ii) that AC in English is structure dependent, (iii) that English PLD includes examples like (1), (2) and maybe (4) (though not if we are a degree-0 learners) but not (5) and so we conclude (iv) if AC is structure dependent, then the fact that it is structure dependent is not itself a fact derivable from inspecting the PLD. That’s the simple POS argument.

Now some observations: First, the argument above supports the claim that the right rule is structure dependent. It does not strongly support the conclusion that the right rule is (3’) with ‘closest’ read as ‘highest.’ This is one structure dependent rule among many possible alternatives. All we did above is compare one structure dependent rule and one non-structure dependent rule and argue that the former is better than the latter given these PLD.  However, to repeat, there are many structure dependent alternatives.[5] For example, here’s another that bright undergrads often come up with:

            (3’’) Move the Aux that is next to the matrix subject to the front

There are many others. Here’s the one that I suspect is closest to the truth:

            (3’’) Move Aux

(3’’) moves the correct Aux to the right place using the very simple rule (3’’) in conjunction with general FL constraints. These constraints (e.g. minimality, the Complex NP constraint (viz. bounding/phase theory)) themselves exploit hierarchical rather than linear structural relations and so the broad structure dependence conclusion of the simple argument follows as a very special case.[6] Note, that if this is so, then AC effects are just a special case of Island and Minimality effects. But, if this is correct, it completely changes what an empiricist learning theory alternative to the standard rationalist story needs to “learn.” Specifically, the problem is now one of getting the ML to derive cyclicity and the minimality condition from the PLD, not just partition the class of acceptable and unacceptable AC outputs (i.e. distinguish (5b) from (5c)). I return to a little more discussion of this soon, but first one more observation.

Second, the simple case above uses data like (5) to make the case that the ‘leftmost’ aux cannot be the one that moves. Note that the application of (3’)-‘leftmost’ here yields the unacceptable string (5c). This makes it easy to judge that (3’)-‘leftmost’ cannot be right for the resulting string is clearly unacceptable regardless of what it is intended to mean. However, using this sort of data is just a convenience for we could have reached the exact same conclusion by considering sentences like (6):

(6)  a. Eagles that can fly swim
b. Eagles that fly can swim
c. Can eagles that fly swim

(6c) can be answered using (6b) not (6a). The relevant judgment here is not a simple one concerning a string property (i.e. it sounds funny) as it is with (5c). It is rather unacceptability under an interpretation (i.e. this can’t mean that, or, it sounds funny with this meaning). This does not change the logic of the example in any important way, it just uses different data, (viz. the kind of judgment relevant to reaching the conclusions is different).

Berwick, Pietroski, Yankama and Chomsky (BPYC) emphasize that data like (6), what they dub constrained homophony, best describes the kind of data linguists typically use and have exploited since, as Chomsky likes to say, “the earliest days of generative grammar.” Think: flying planes can be dangerous, or I saw the woman with the binoculars, and their disambiguating flying planes is/are dangerous and which binoculars did you see the woman with.  At any rate, this implies that the more general version of the AC phenomena is really independent of string acceptability and so any derivation of the phenomenon in learning terms should not obsess over cases like (5c). They are just not that interesting for the POS problem arises in the exact same form even in cases where string acceptability is not a factor.

Let’s return briefly to the first point and then wrap up. The simple discussion concerning how to interpret (3’) is good for illustrating the logic of POS. However, we know that there is something misleading about this way of framing the question. How do we know this? Well, because, the pattern of the data in (5) and (6) is not unique to AC movement. Analogous dependencies (i.e. where some X outside of the relative clause subject relates to some Y inside it) are banned quite generally. Indeed, the basic fact, one, moreover that we all have known about for a very long time, is that nothing can move out of a relative clause subject. For example: BPYC discuss sentences like (7):

(7)  Instinctively, eagles that fly swim

(7) is unambiguous, with instinctively necessarily modifying fly rather than swim. This is the same restriction illustrated in (6) with fronted can restricted in its interpretation to the matrix clause. The same facts carry over to examples like (8) and (9) involving Wh questions:

(8)  a. Eagles that like to eat like to eat fish
b. Eagles that like to eat fish like to eat
c. What do eagles that like to eat like to eat

(9)  a. Eagles that like to eat when they are hungry like to eat
b. Eagles that like to eat like to eat when they are hungry
c. When do eagles that like to eat like to eat

(8a) and (9a)  are appropriate answers to (8c) and (9c) but (8b) and (9b) are not. Once again this is the same restriction as in (7) and (6) and (5), though in a slightly different guise. If this is so, then the right answer as to why AC is structure dependent has nothing to do with the rule of AC per se (and so, plausibly, nothing to do with the pattern of AC data). It is part of a far more general motif, the AC data exemplifying a small sliver of a larger generalization. Thus, any account that narrowly concentrates on AC phenomena is simply looking at the wrong thing! To be within the ballpark of the plausible (more pointedly, to be worthy of serious consideration at all), a proffered account must extend to these other cases of as well. That’s the problem in a nutshell.[7]

Why is this important? Because criticisms of the POS have exclusively focused on the toy example that Chomsky originally put forward to illustrate the logic of POS.  As noted, Chomsky’s original simple discussion more than suffices to motivate the conclusion that G rules are structure dependent and that this structure dependence is very unlikely to be a fact traceable to patterns in the PLD. But the proposal put forward was not intended to be an analysis of ACs, but a demonstration of the logic of the POS using ACs as an accessible database. It’s very clear that the pattern attested in polar questions extends to many other constructions and a real account of what is going on in ACs needs to explain these other data as well. Suffice it to say, most critiques of the original Chomsky discussion completely miss this. Consequently, they are of almost no interest.

Let me state this more baldly: even were some proposed ML able to learn to distinguish (5c) from other sentences like it (which, btw, seems currently not to be the case), the problem is not just with (5c) but sentences very much like it that are string kosher (like (6)). And even were they able to accommodate (6) (which so far as I know, they currently cannot) there is still the far larger problem of generalizing to cases like (7)-(9). Structure dependence is pervasive, AC being just one illustration. What we want is clearly an account where these phenomena swing together; AC, Adjunct WH movement, Argument Wh Movement, Adverb fronting, and much much more.[8] Given this, the standard empiricist learning proposals for AC are trying (and failing) to solve the wrong problem, and this is why they are a waste of time. What’s the right problem? Here’s one: show how to “learn” the minimality principle or Subjacency/Barriers/Phase theory from PLD alone. Now, were that possible, that would be interesting. Good luck.

Many will find my conclusion (and tone) harsh and overheated. After all isn’t it worth trying to see if some ML account can learn to distinguish good from bad polar questions using string input? IMO, no. Or more precisely, even were this done, it would not shed any light on how humans acquire AC. The critics have simply misunderstood the problem; the relevant data, the general structure of the phenomenon and the kind of learning account that is required. If I were in a charitable mood, I might blame this on Chomsky. But really, it’s not his fault. Who would have thought that a simple illustrative example aimed at a general audience should have so captured the imagination of his professional critics! The most I am willing to say is that maybe Fodor is right and that Chomsky should never have given a simple illustration of the POS at all. Maybe he should in fact be banned from addressing the uninitiated altogether or only if proper warning labels are placed on his popular works.

So, to end: why am I not impressed by empiricist discussions of AC? Because I see no reason to think that this work has yielded or ever will yield any interesting insights to the problems that Chomsky’s original informal POS discussion was intended to highlight.[9] The empiricist efforts have focused on the wrong data to solve the wrong problem.  I have a general methodological principle, which I believe I have mentioned before: those things not worth doing are not worth doing well. What POS’s empiricist critics have done up to this point is not worth doing. Hence, I am, when in a good mood, not impressed. You shouldn’t be either.






[1] One point before getting down and dirty: what follows is not at all original with me (though feel free to credit me exclusively). I am repeating in a less polite way many of the things that have been said before. For my money, the best current careful discussion of these issues is in Berwick, Pietroski, Yankama and Chomsky (see link to this below). For an excellent sketch on the history of the debate with some discussion of some recent purported problems with the POS arguments, see this handout by Howard Lasnik and Juan Uriagereka.
[2] I believe (actually I know, thx Howard) that the case is first discussed in detail in Language and Mind (L&M) (1968:61-63). The argument form is briefly discussed in Aspects (55-56), but without attendant examples. The first discussion with some relevant examples is L&M. The argument gets further elaborated in Reflections on Language (RL) and Rules and Representations (RR) with the good and bad examples standardly discussed making their way prominently into view. I think that it is fair to say that the Chomsky “analysis” (btw, these are scare quotes) that has formed the basis of all of the subsequent technical discussion and criticism is first mooted in L&M and then elaborated in his other books aimed at popular audiences. Though the stuff in these popular books is wonderful, it is not LGB, Aspects, the Black Book, On Wh movement, or Conditions on transformations. The arguments presented in L&M, RL and RR are intended as sketches to elucidate central ideas. They are not fully developed analyses, nor, I believe, were they intended to be. Keep this in mind as we proceed.
[3] Of course, not sentences, but utterances thereof, but I abstract from this nicety here.
[4] Those who have gone through this know that the notion ‘Aux’ does not come tripping off the tongue of the uninitiated. Maybe ‘helping verb,’ but often not even this.  Also, ‘move’ can be replaced with ‘put’ ‘reorder’ etc.  If one has an inquisitive group, some smart ass will ask about sentences like ‘Did Bill eat lunch’ and ask questions about where the ‘did’ came from. At this point, you usually say (with an interior smile), to be patient and that all will be revealed anon.
[5] And many non-structure dependent alternatives, though I leave these aside here.
[6] Minimality suffices to block (4) where the embedded Aux moves to the matrix C. The CNPC suffices to block (5c). See below for much more discussion.
[7] BTW, none of this is original with me here. This is part of BPYC’s general critique.
[8] Indeed, every case of A’-movement will swing the same way. For example: in It’s fresh fish that eagles that like to eat like to eat, the focused fresh fish is complement of the matrix eat not the one inside the RC.
[9] Let me add one caveat: I am inclined to think that ML might be useful in studying language acquisition combined with a theory of FL/UG. Chomsky’s discussion in Chapter 1 of Aspects still looks to me very much like what a modern Bayesian theory with rich priors and a delimited hypothesis space might look like. Matching Gs to PLD even given this, does not look to me like a trivial task (and work by those like Yang, Fodor, Berwick) strike me as trying to address this problem. This, however, is very different from the kind of work criticized here, where the aim has been to bury UG not to use it. This has been a both a failure and, IMO, a waste of time.

Monday, February 24, 2014

DTC redux

Syntacticians have effectively used just one kind of probe to investigate the structure of FL, viz. acceptability judgments. These come in two varieties: (i) simple “sounds good/sounds bad” ratings, with possible gradations of each (effectively a 6ish point scale ok, ?, ??, ?*, *, **), and (ii) “sounds good/sounds bad under this interpretation” ratings (again with possible gradations). This rather crude empirical instrument has proven to be very effective as the non-trivial nature of our theoretical accounts indicates.[1] Nowadays, this method has been partially systematized under the name “experimental syntax.” But, IMO, with a few important conspicuous exceptions, these more refined rating methods have effectively endorsed what we knew before. In short, the precision has been useful, but not revolutionary.[2]

In the early heady days of Generative Grammar (GG), there was an attempt to find other ways of probing grammatical structure. Psychologists (following the lead that Chomsky and Miller (1963) (C&M) suggested) took grammatical models and tried to correlate them with measures involving things like parsing complexity or rate of acquisition. The idea was a simple and appealing one: more complex grammatical structures should be more difficult to use than less complex ones and so measures involving language use (e.g. how long it takes to parse/learn something) might tell us something about grammatical structure. C&M contains the simplest version of this suggestion, the now infamous Derivational Theory of Complexity (DTC). The idea was that there was a transparent (i.e. at least a homomorphic) relation between the rules required to generate a sentence and the rules used to parse it and so parsing complexity could be used to probe grammatical structure.

Though appealing, this simple picture can (and many believed did) go wrong in very many ways (see Berwick and Weinberg 1983 (BW) here for a discussion of several).[3] Most simply, even if it is correct that there is a tight relation between the competence grammar and the one used for parsing (which there need not be, though in practice there often is, e.g. the Marcus Parser) the effects of this algorithmic complexity need not show up in the usual temporal measures of complexity, e.g. how long it takes to parse a sentence. One important reason for this is that parsers need not apply their operations serially and so the supposition that every algorithmic step takes one time step is just one reasonable assumption among many. So, even if there is a strong transparency between competence Gs and the Gs parsers actually deploy, no straightforward measureable time prediction follows.

This said, there remains something very appealing about DTC reasoning (after all, it’s always nice to have different kinds of data converging on the same conclusion, i.e. Whewell’s consilience) and though it’s true that the DTC need not be true, it might be worth looking for places where the reasoning succeeds. In other words, though the failure of DTC style reasoning need not in and of itself imply defects in the competence theory used, a successful DTC style argument can tell us a lot about FL. And because there are many ways for a DTC style explanation to fail and only a few ways that it can succeed, successful stories if they exist can shed interesting light on the basic structure of FL.

I mention this for two reasons. First, I have been reading some reviews of the early DTC literature and have come to believe that its demonstrated empirical “failures” were likely oversold. And second, it seems that the simplicity of MP grammars has made it attractive to go back and look for more cases of DTC phenomena. Let me elaborate on each point a bit.

First, the apparent demise of the DTC. Chapter 5 of Colin Phillips’ thesis (here) reviews the classical arguments against the DTC.  Fodor, Bever and Garrett (in their 1974 text) served as the three horsemen of the DTC apocalypse. They interned the DTC by arguing that the evidence for it was inconclusive. There was also some experimental evidence against it (BW note the particular importance of Slobin (1966)). Colin’s review goes a very long way in challenging this pessimistic conclusion. He sums up his in depth review as follows (p.266):

…the received view that the initially corroborating experimental evidence for the DTC was subsequently discredited is far from an accurate summary of what happened. It is true that some of the experiments required reinterpretation, but this never amounted to a serious challenge to the DTC, and sometimes even lent stronger support to the DTC than the original authors claimed.

In sum, Colin’s review strongly implies that linguists should not have abandoned the DTC so quickly.[4] Why, after all, give up on an interesting hypothesis, just because of a few counter-examples, especially ones that when considered carefully seem on the weak side? In retrospect, it looks like the abandonment of the strong hypothesis was less a matter of reasonable retreat in the face of overwhelming evidence than a decision that disciplines occasionally make to leave one another alone for self-interested reasons. With the demise of the DTC, linguists could assure themselves that they could stick to their investigative methods and didn’t have to learn much psychology and psychologists could concentrate on their experimental methods and stay happily ignorant of any linguistics. The DTC directly threatened this comfortable “live and let live” world and perhaps this is why its demise was so quickly embraced
by all sides.

This state of comfortable isolation is now under threat, happily.  This is so for several reasons. First, some kind of DTC reasoning is really the only game in town in cog-neuro. Here’s Alec Marantz’s take:

...the “derivational theory of complexity” … is just the name for a standard methodology (perhaps the dominant methodology) in cognitive neuroscience (431).

Alec rightly concludes that given the standard view within GG that what linguists describe are real mental structures, there is no choice but to accept some version of the DTC as the null hypothesis. Why? Because, ceteris paribus:

…the more complex a representation- the longer and more complex the linguistic computations necessary to generate the representation- the longer it should take for a subject to perform any task involving the representation and the more activity should be observed in the subject’s brain in areas associated with creating or accessing the representation or performing the task (439).

This conclusion strikes me as both obviously true and salutary, with one caveat. As BW has shown us, the ceteris paribus clause can in practice be quite important.  Thus, the common indicators of complexity (e.g. time measures) may be only indirectly related to algorithmic complexity. This said, GG is (or should be) committed to the view that algorithmic complexity reflects generative complexity and that we should be able to find behavioral or neural correlates of this (e.g. Dehaene’s work (discussed here) in which BOLD responses were seen to track phrasal complexity in pretty much a linear fashion or Forster’s work finding temporal correlates mentioned in note 4).

Alec (439) makes an additional, IMO correct and important, observation. Minimalism in particular, “in denying multiple routes to linguistic representations,” is committed to some kind of DTC thinking.[5] Furthermore, by emphasizing the centrality of interface conditions to the investigation of FL, Minimalism has embraced the idea that how linguistic knowledge is used should reveal a great deal about what it is. In fact, as I’ve argued elsewhere, this is how I would like to understand the “strong minimalist thesis,” (SMT) at least in part. I have suggested that we interpret the SMT as committed to a strong “transparency hypothesis” (TH) (in the sense of Berwick & Weinberg), a proposal that can only be systematically elaborated by how linguistic knowledge is used.

Happily, IMO, paradigm examples of how to exploit “use” and TH to probe the representational format of FL are now emerging. I’ve already discussed how Pietroski, Hunter, Lidz and Halberda’s work relates to the SMT (e.g. here and here). But there is other stuff too of obvious relevance: e.g. BW’s early work on parsing and Subjacency (aka Phase Theory) and Colin’s work on how islands are evident in incremental sentence processing. This work is the tip of an increasingly impressive iceberg. For example, there is analogous work showing that that parsing exploits binding restrictions incrementally during processing (e.g. by Dillon, Sturt, Kush).

This latter work is interesting for two reasons. It validates results that syntacticians have independently arrived at using other methods (which, to re-emphasize, is always worth doing on methodological grounds). And, perhaps even more importantly, it has started raising serious questions for syntactic and semantic theory proper. This is not the place to discuss this in detail (I’m planning another post dedicated to this point), but it is worth noting that given certain reasonable assumptions about what memory is like in humans and how it functions in, among other areas, incremental parsing, the results on the online processing of binding noted above suggest that binding is not stated in terms of c-command but some other notion that mimics its effects.

Let me say a touch more about the argument form, as it is both subtle and interesting. It has the following structure: (i) we have evidence of c-command effects in the domain of incremental binding, (ii) we have evidence that the kind of memory we use in parsing cannot easily code a c-command restriction, thus (iii) what the parsing Grammar (G) employs is not c-command per se but another notion compatible with this sort of memory architecture (e.g. clausemate or phasemate). But, (iv) if we adopt a strong SMT/TH (as we should), (iii) implies that c-command is absent from the competence G as well as the parsing G. In short, the TH interpretation of SMT in this context argues in favor of a revamped version of Binding Theory in which FL eschews c-command as a basic relation. The interest of this kind of argument should be evident, biut let me spell it out. We S-types are starting to face the very interesting prospect that figuring out how grammatical information is used at the interfaces will help us choose among alternative competence theories by placing interface constraints on the admissible primitives. In other words, here we see a non-trivial consequence of Bare Output Conditions on the shape of the grammar. Yessss!!!

We live in exciting times. The SMT (in the guise of TH) conceptually moves DTC-like considerations to the center of theory evaluation. Additionally, we now have some useful parade cases in which this kind of reasoning has been insightfully deployed (and which, thereby, provide templates for further mimicking). If so, we should expect that these kinds of considerations and methods will soon become part of every good syntactician’s armamentarium.




[1] The fact that such crude data can be used so effectively is itself quite remarkable. This speaks to the robustness of the system being studied for such weak signals should not be expected to be so useful otherwise.
[2] Which is not to say that such more careful methods don’t have their place. There are some cases where being more careful has proven useful. I think that Jon Sprouse has given the most careful thought to these questions. Here is an example of some work where I think that the extra care has proven to be useful.
[3] I have not been able to find a public version of the paper.
[4] BW note that Forster provided evidence in favor of the DTC even as Fodor et. al. were in the process of burying it. Forster effectively found temporal measures of psychological complexity that tracked the grammatical complexity the DTC identified by switching the experimental task a little (viz. he used an RSVP presentation of the relevant data).
[5] I believe that what Alec intends here is that in a theory where the only real operation is merge then complexity is easy to measure and there are pretty clear predictions of how this should impact algorithms that use this information. It is worth noting that the heyday of the DTC was in a world where complexity was largely a matter of how many transformations applied to derive a surface form. We have returned to that world again, though with a vastly simpler transformational component.

Sunday, April 7, 2013

Operationalizing the Strong Minimalist Thesis


In the sciences, it takes a lot of work for a new idea to take hold.  Aside from a modicum of conceptual clarity, a conceptual innovation must be operationalized, and this requires offering canonical or paradigmatic models of its application. I mention this because for me one of the recurring difficulties with the Minimalist Program has been figuring out what makes any given proposal/analysis minimalist (i.e. the path from Minimalist Program to Minimalist Theory is often obscure). So, while being a minimalist is all very nice (e.g. it greatly (minimally?) enhances my self-esteem), what is more absent than it should be are clear examples of doing minimalism; examples of what makes a particular analysis/proposal minimalist or how the abstract leading ideas get concretized in everyday work. Fortunately, I have recently read some papers that, I believe, can serve as parade cases of minimalist thinking and provide clear examples of how the Strong Minimalist Thesis can be incarnated. They are the topic of today’s sermon.

First, what’s the Strong Minimalist Thesis (SMT)?  It is the claim that “language is an optimal solution” to interface conditions:

…the human faculty of language FL [is] an optimal solution to minimal design specifications, conditions that must be satisfied for language to be usable at all…for each language L (a state of FL), the expressions generated by L must be “legible” to systems that access these objects at the interface between FL and external systems – external to FL, internal to the person. (DbP 1).

The systems that L interfaces with use its generated objects. The SMT proposes that the objects of L are well designed for the cognitive interfaces that use them in doing what they do. Put another way, the generated objects can be used as is (i.e. without further alteration) to do what needs getting done (viz. the information they contain need not be further repackaged for the interfaces to use them for whatever tasks they set their “hands” to.)[1] What the hell does this mean?  Two papers by Pietroski, Lidz, Halberda and Hunter (PLHH) (here and here) provide a useful concrete model for interpreting these abstract claims. Their discussion centers on the correct representation of the meaning of most. Here’s what PLHH do.

The papers are interested in figuring out how most sentences affect visual perception (i.e. how the visual system uses grammatical information in making a visual judgment). Specifically, how does someone who hears (1) judge whether a certain presented array of dots verifies (1).

(1)  Most of the dots are blue

It’s a given that (1) is true iff the number of blue dots exceeds the number of non-blue dots. The question is how does one represent the italicized information and does it matter to what people do.  The problem becomes interesting in that there are several ways of representing the quantity information in (1) that are not intensionally equivalent (viz. they use different predicates and different relations) despite being truth functionally the same (in Frege speak: they involve different routes to the same truth value). Here are three possible representations for the meaning of most.

            (2)       a. OneToOnePlus*: [{x: D (x)}, [x: Y (x)}] iff some some set s, s Ì {X: D
    (x)} and OneToOne [s, {x: Y (x)}][2]
                        b. |{x: D (x) & Y (x)}| > {x: D (x) & - Y (x)}|
                        c. |{x: D (x) & Y (x)}| > |{ x: D (x)}| - |{x: D(x) & Y (x)}|

For D= ‘dot’ and Y= ‘blue,’ the (2a) representation carries out the evaluation of the dot scene by pairing the blue with the non-blue dots and seeing if there is at least one extra blue dot left over. The second, in (2b), sees if the size of the set of blue dots is greater than the size of the set of non-blue dots and the third, (2c) sees if the size of the set of blue dots is greater than the size of the set of all the dots minus the set of blue dots.  (2a) differs from the others in using a distinct predicate (viz. OnToOne) while (2b) and (2c) differ in the sets are compared, the former directly calculating the set of non-blue dots, the latter never directly numerically evaluating this set (i.e. it does so indirectly by directly subtracting the blue dots from the entire set of dots).

PLHH reason as follows: There are two possibilities when speakers are asked to evaluate dot scenes on hearing (1):  (i) the visual/counting system might find some of these representations more congenial than the others. (ii) Or, the visual/counting system might use any of these three truth functionally equivalent representations given the right circumstances.  If (i) holds then this visual-counting interface favors one representation over the other two. Why? Because “linguistic meanings are related the cognitive systems that are used to evaluate sentences for truth and falsity.” More specifically: “a declarative sentence S is semantically associated with a canonical procedure for determining whether S is true…[and] competent speakers are biased towards strategies that directly reflect canonical specifications of truth conditions.” They dub this thesis the Interface Transparency Thesis (ITT). Put in slightly more “minimalist” terms: in a well designed grammar, its products will supply the information the interfaces need in a way transparent to those needs. More specifically, the interfaces will “use” the kinds of information the linguistic structure directly encodes.  Or, the information that the grammatical representations encode and the information that the interface uses is one and the same.

Before going on, observe that the ITT provides a useful (and, as PLHH demonstrate, usable) interpretation of “optimal solution to interface conditions”: for a given interface (i.e. system that uses language) how transparent is the mapping between the information made available from L and the information that the interface uses to do what it does?  SMT amounts to the hypothesis that a strong transparency holds between the information as coded in L and the information these various interfaces exploit to do what they do.  If considerable transparency holds then SMT is vindicated. If not, not.[3] So among other things, one very useful contribution of PLHH’s papers is that they provide a substantive yet manageable interpretation of the SMT. But that is not all.

I would not be going through all of this were it not the case that PLHH demonstrate that not all representations of most are created equal.  They provide very good reasons to conclude that (2c) is the right semantic representation of most (or, is clearly superior to (2a,b)). Demonstrating this is conceptually simple but the argument is quite complicated and rich. Showing (2c) is superior to (2a/b) requires knowing a lot about properties of the interface, in this case knowing how people count (humans use two counting systems with different properties) and how people “count” what they see.  Luckily, this is a well-studied domain of visual perception (Halberda (one of the ‘H’s) has done a lot of basic work on how humans do this) and so it is possible to contrive visual dot scenes that would favor one or another of the representational formats in (2) and see what happens. The answer is that the information in (2c) is what humans compute, even when things are visually arranged so that (2a) or (2b) would be simple to apply.[4] The bottom line: humans have a bias for (2c) and the source of this bias is reasonably attributed to the fact that the visual-counting system likes the information as represented in (2c), as per the ITT.

Assume that this is correct. Can we go further and explain the properties of the representation (2c) in terms of the properties of this interface?  Let me be clear: PLHH show that one representation is preferred to others. We can attribute this to the meaning of most being (2c) coupled with the ITT.  Given this we can ask the next question: is the representational format of (2c) explicable in terms of the properties of this interface? Recall, the SMT suggests that FL (a late emerging system) is the “optimal solution” to interface requirements. This suggests that the properties of FL are what they are because of the properties of the interfaces that use them.  PLHH show that for some features of (2c) this explanatory chit can be cashed in. Here’s their very interesting argument.

First, note that in (2b,c) different sets are being selected for enumeration (recall (2c) says nothing direct about the non-blue dots).  This said, it’s a fact that humans are very good at selecting positive features in an array (e.g. blue dots or red dots or green dots) but not at negatively specified features (e.g. not-blue dots). This clearly argues against (2b) in that one is directed to select the non-blue dots.  Second, it has been shown that subjects (human adults) “always attend and enumerate the superset of all dots,” which is good news for (2c) as this is a required part of the specified computation. Third, it can be shown that when subjects use the Approximate Number System (ANS), the one used in this task, they can “estimate the cardinality of up to three sets in parallel,” which means that if there are blue dots, red dots, yellow dots, green dots and mauve dots that (2b) could not be used to evaluate (1) in such a scene (i.e. the requirements in (2b) do not scale up very well, whereas those in (2c) do).  In sum, as PLHH put it:

A meaning like [(2c)]…is straightforwardly verified with these resources, since the sets required for verification (one color plus the superset) are easily and automatically attended by the visual system. Moreover, this meaning does not become less plausible as the number of color subsets increases.

In other words, given the ITT in this domain (for which PLHH have provided evidence) and given the properties of the ANS and the visual system, representations like (2c) perfectly fit the structural capacities of the interface.  Thus, the meaning of most as specified in (2c) fits the noted interface specifications to a (SM)T!

The work is gorgeous. But aside from its stand-alone value, it’s really useful for minimalists to contemplate and absorb.  The Interface Transparency Thesis provides a useful concept for investigating how interfaces and grammars “fit.”  If such a fit can be established, it is possible (sometimes) to argue from properties of the interface to properties of the representations.  Minimalist should understand and absorb this two-step tango for it serves to operationalize the SMT, moving it from a frequently annoying slogan to a research problem, something every minimalist should welcome.

Let me end by noting that PLHH are not alone in deploying this argument. Berwick and Weinberg (BW) (here, where the notion ‘transparency’ was also mooted and discussed) develop an earlier version of this argument.[5] Their version of the ITT considers another interface, the parser (i.e. those interfaces that underlie parsing utterances in real time) and asks what grammatical properties would allow for optimal parsing (roughly parsing in linear time). BW showed that parsers with bounded left contexts would serve nicely and argued that grammars that respected some version of cyclicity+subjacency would perfectly fit the bill. So, if we assume that parsers use the structures generated by L to parse then a cyclic+subjacent compliant grammar would be the perfect fit.  The form of argument is exactly the same as that in PLHH, with the relevant interface this time being those that underlie parsing.

So, the upshot: there are now some paradigm cases out there of how to argue for the SMT.  Deploying these arguments requires knowing a lot about grammar and a lot about some interface property. However, as these two cases show, such arguments can be made. Moreover, they can be made convincingly.  It seems that the SMT is not merely a guiding regulative ideal but even one that can be empirically evaluated. Pretty damn good!

The take home message?  One effective way of investigating the SMT is to identify some interface system (the parser, the visual system, the ANS) and see how it uses the grammatical information provided by L.  The SMT leads to the expectation that it uses the information “transparently,” and that the details of how the interface works can explain why the representation looks like it does. This is hard to pull off, for it requires knowing a lot both about the grammar and the interface at issue.  It suggests that future syntacticians will need to have new skill sets and/or be very collaborative.  This will no doubt be demanding. But, hey, who every said that cognitive-biolinguistics would be easy. The most anyone promised was that it would be fun, and, if these cases are any indication, crammed with more than a touch of intellectual beauty as well.




[1] If I understand the notion “covering grammar” correctly, then one might say that the competence grammar is the covering grammar for the relevant interface.  In the best case it is the grammar that every interface uses.
[2] OneToOne is a function pairs individuals in D and Y one to one.
[3] Note the word ‘considerable.’  The relevant evaluation will revolve around some estimation of the degree of transparency and this may be a labile notion.  It may be possible to make these estimations on a case by case basis without having a general measure of transparency.
[4] As PLHH note: humans can in fact apply the predicates in (2a) and (2b) in non-quantificational tasks. Thus their failure to apply them in these “linguistic” contexts cannot be traced to some general human incapacity to deploy them.
[5] In addition Colin Phillips proposal that the grammar is identical to the parser can be interpreted as postulating a very strong transparency assumption for this interface.