Comments

Showing posts sorted by date for query On Wh movement. Sort by relevance Show all posts
Showing posts sorted by date for query On Wh movement. Sort by relevance Show all posts

Tuesday, January 15, 2019

Movement, islands and the ECP

Some papers reset the research agenda. This one by Lu and Yoshida (L&Y), I believe, is one of those (here is a slide conveying the basic point. The paper is under submission at LI and I assume it will be accepted and rapidly published (if not this is will tell us more about LI than it will about the quality of this paper)). The topic is the island status of Wh-in-situ (WIS) constructions in Chinese. The finding is that using judgment studies of the Sprouse experimental syntax (ES) variety provides evidence for two stunning conclusions: (i) that WISs respect islands and (ii) that there is no evidence for an argument/adjunct distinction wrt WISs. Both data points are theoretically pregnant and this post will largely concentrate on drawing out some of the implications. Many of these are mentioned in the paper (yup, I have a draft), so are not original with me. Let’s start.

L&Y is motivated by the premise that ES provides a useful tool for the refining linguistic judgments. The idea, as Sprouse has convincingly argued, is that grammatical complexity should induce a super-additivity effect in well-constructed judgment experiments (see, e.g. here and here for discussion and here for a nice review of the methodology). Importantly, super-additivity profiles arise in cases where less involved rating studies find nothing indicating un-grammaticality.

Before pressing on, let’s make an important and obvious point: all GGers distinguish (or should distinguish) acceptability from grammaticality. Acceptability is a probe for grammaticality. Acceptability is an observable property of utterances. Grammaticality is an abstract property of I-linguistic mental representations. Grammaticality is inferred from acceptability under the right conditions (given the right controls as realized by the appropriate minimal pairs). All of this is old hat, but a still very stylish and durable hat. 

Happily, for most of what we have done in GG, acceptability closely tracks grammaticality, but we also know that the two notions can and do diverge (see here for some discussion). ES is particularly useful for cases where this happens and the simple judgment elicitation procedure (e.g. ask a native speaker) indicates all is well. Diogo Almeida has dubbed cases of the latter “subliminal.” One of ES’s important contributions to syntax has been the discovery of such subliminal effects (SE), SEs being cases where ES procedures reveal super-additivity effects while more regular elicitation suggests grammaticality. So, for example, we now have many examples where standard elicitation has indicated that a certain dependency in a certain language shows no island sensitivity (i.e. the sentences are judged (highly) acceptable) while ES techniques indicate sensitivity to these same island effects (i.e. the relevant data display super-additivity effects).

We also find the converse: standard techniques indicating a profound difference in acceptability, while ES techniques showing nothing at all.[1]All in all then, ES has provided a useful additional kind of data, one that is often more sensitive to G structure than the quick and dirty (and largely accurate and hence very useful) standard judgment techniques, which sometimes fail to track these. 

So, back to the main point: L&Y is an ES study of WISs in Chinese and it has two important findings: that allWISs in Chinese exhibit relative clause island effects (henceforth RCI) (i.e. they alldisplay the super-additivity profile) and that there is no ES evidence that long “why” movement from an RCI is appreciably worse than long “why” movement absent an RCI (i.e. these cases when contrasted do not show a super-additvity profile). The first result argues that WISs are island sensitive and the second argues that there is no additional ECP effect distinguishing WISs like who/what from WISs like why. If correct, this is very big news, and, IMO, very welcome news. Let me say why.

First, as L&Y emphasizes, this result rules out most of the standard approaches to WIS constructions. In particular the result rules out two kinds of theories: (i) accounts that distinguish between overt movement vs covert movement (e.g. Huang’s) and treat island effects as effectively reflexes of overt movement (say, via a chain condition at SS) and (ii) theories that postulate two different kinds of operations (Movement vs Binding) to license WISs with movement subject to islands and binding exempt from them (as in, say, a Rizzi-Cinque approach to ECP effects). Both such kinds of theories will have problems with the apparent fact that WISs induce super-additivity effects.

It is worth noting, furthermore, that the sensitivity of WISs to islands is not the only example of apparent non-movement generated structures being island compliant. The same holds wrt resumptive pronoun (RP) constructions. These also appear to respect islands despite the absence of the main hallmark of movement (i.e. a gap in the “movement” site).[2]Both this RP data and now the WIS data point to the same conclusion: that island effects are notPF effects.[3]From my reading of the literature, this is the most popular current approach to islands and it has some terrifically interesting evidence in support (in particular the fact that some ellipsis (i.e. sluicing) obviates island violations). However, if L&Y are right, then we may have to rethink this assumption (see note 3 however).

Indeed, I would go further (and here it is NH speaking rather than L&Y). There have long been two general approaches to islands. 

First, we have Chomsky’s view of subjacency elaborated in ‘On wh movment’ that treats islands as reflecting bounds on the computational procedure. Island effects reflect the subjacency condition (aka PIC), which bounds the domain of computation (an idea motivated by the reasonable assumption that bounding a domain of computation makes doing computations more tractable).[4]

The second approach can be traced back to Ross’s thesis (islands restrict chopping rules) but has been developed as part of the linearization industry spurred by Kayne’s seminal work and mooted most explicitly by Uriagereka.[5]

The L&Y results argue pretty strongly, IMO, for Chomsky’s original conception precisely because they appear to hold whether or not the construction involves an obvious phonetic gap (gaps being problematic as they undo linearizations). If this is so, then it argues against linearization based approaches to the problem (leaving, of course a very big question: what to do about sluicing).[6]

We can go further still. The L&Y results also argue for a Merge only syntax. Here is what I mean. IMO, the central empirical thesis of the Minimalist Program (MP) is the Merge Hypothesis (MH). MH is the claim that the only specifically linguistic operation of FL is Merge. This entails that allG dependencies are merge mediated. The strong version of the thesis excludes operations like long distance Agree, which spans the same domains as I-merge but is a different operation. Note that it is natural to suppose that I-merge is movement and Agree is some kind of binding or feature sharing. At any rate, the classical conceptions gain empirical benefit from the “observation” that WISs do not display island effects. Why? Because, we might say, they are licensed by Agree not by I-merge and only the latter (being the MP analogue of movement) is subject to subjacency (or its current analogue). But as L&Y indicates this is precisely the wrong conclusion. WISs are subject to islands. A merge only syntax insists that all A’-dependencies are formed in the same way, via I-Merge, as this as the only way to establish any non-local grammatical dependency. So if WISs are G licensed, then they must be G licensed via I-merge and so will form a natural class with overt Wh movement. And this is what L&Y find. Both show super-additivity effects across islands. Thus, L&Y’s findings are what we should expect from a merge only syntax and it cautions against larding this best of MP theories with Agree/Probe-Goal titivations.[7]

We can milk a second important conclusion from L&Y. It solves a giant problem for MP. Which problem? The problem of unifying subjacency with the ECP. I have suggested elsewhere (see here) that the argument/adjunct asymmetries at the heart of the ECP are very MP problematic. This is so for a variety of reasons. The three that move me most are the fact that the ECP is a trace licensing condition and MP eschews traces, the huge theoretical redundancy between ECP and subjacency, and the “ugliness” of the basic technical machinery required to allow the ECP to track the argument/adjunct asymmetry. One of the nice implications of L&Y is that we need not worry about the problems that the ECP generates for MP because the theoretical apparatus is based on a mistaken description of the data. If L&Y is right, then there is no argument/adjunct asymmetry. Poof, the MP problem disappears and with it the ad-hoc theoretically unmotivated (within MP) technical apparatus required to track it.  

Of course, this overstates matters. It behooves us to go over the ECP data more carefully and see how to resolve the difference in acceptability that the standard literature identified. Why after all if all Whs are created equal do long distance adjuncts resist movement more fiercely than do long distance arguments?[8]L&Y offer a suggestion (no spoiler from me, read the squib when it comes out). But whatever the right answer is, it does not rest on making an invidiousgrammaticaldistinction between the two kinds of dependencies. And this is just what MP needs in order to start distinguishing ECP effects from the ECP theoretical apparatus in GB.

Let me hit this a bit harder. The ugliest parts of theories of A’-dependency within GB arise in response to argument/adjunct asymmetry effects. The technical machinery in Barriers (built on Lasnik and Saito foundations), though successful empirically (IMO, Lasnik and Saito’s theory was considerably more empirically effective than Barriers) had little of the virtual conceptual necessity MPers pine for. Nor were alternative theories (Generalized Binding, Connectedness) much prettier. Nonetheless, we put up with that stuff and developed it theoretically because it appeared to be empirically called for. The right aesthetic conclusion should have been (and actually was) that it was too contrived to be correct. L&Y provides courage for our aesthetic convictions. We should have judged these theories as suspect because ugly, though we would have been empirically premature in drawing that conclusion. Given L&Y, the facts are not what we took them to be despite reflecting very different acceptability profiles. 

There is a moral here, and you can all guess what it is but I cannot resist making it explicit anyhow. L&Y provide evidence for a methodological precept that we fail to respect enough: facts can change, no less than theory can. Or to put this another way: just as we can make theoretical wrong turns that we come to revise, we can adopt empirical generalizations that turn out to be misleading. The standard view is that data is hard and theory is fluffy and when the two clash it is best to revise the theory than rethink the data. L&Y provides a case where this is reversed. And I say hooray!

Let me make one more point and I will end this overly long post (long, and yet, filled with endlessly many loose ends). L&Y exemplifies something that I think is important. It is an empirical paper whose purpose is to directlytest a core theoretical assumption. This is not something that we generally see within syntax. Most papers are not out to test theoretical assumptions. Most papers use theory to explore more data. Theory might be tested but it is generally a by-product of better descriptive coverage. L&Y works differently. It starts from the theory and constructs an empirical intervention to probe it. Moreover, the question is quite precise and the assumptions required to answer it are clear within the confines of the project. This has all the look and smell of an honest to god experiment, a process whereby we query the theory using a relatively well-understood probe. Both empirical methods of exploration are worthwhile, but they are different and it is only relatively recently, I think, that we are seeing examples of the second experimental kind gaining traction. 

Curiously (perhaps), a feature of this second kind of paper is that the paper is short. L&Y is a squib. Empirical explorations in linguistics often read like novellas. L&Y is very definitely a very very short story. I would like to suggest that experimental papers like L&Y reflect the scientific health of linguistics. It is now possible to ask a sharp question, and give a sharp answer. We need more of these kinds of short pointed experimental forays into the data starting from well-formulated theoretical starting points.

That’s it. I have gone on far too long. L&Y is terrific. If correct, it is very important. I personally hope the results stand up. It would go a long way to cleaning up a particularly untidy part of syntactic theory and thereby further vindicating the promise of MP, indeed a particularly strong version of MP, one that endorses a Merge only conception of grammar.


[1]ES techniques have, for example, suggested that adjunct island effects might not be of a piece with other islands as they (often) fail to display super-additivity effects.
[2]To be slightly more careful and squinting at the ES results wrt resumptives the following is more accurate: resumptives uniformly ameliorate fixed subject constraint violations but note “mere” subjacency violations. Thus, resumptives inside islands seem to show the same super-additvity profiles as their moved counterparts.
[3]Perhaps a more felicitous way of putting matters is that RPs and WISs are also products of I-merge and so expected to be subject to islands. One reason for treating islands as PF effects is to capture the distinctionbetween these cases and more conventional examples of overt movement. If, however, they pattern the same then the motivation for treating islands as PF effects weakens. This said, I am pretty sure it is possible to model these cases as formed via movement/I-merge by, for example, treating them as cases of remnant movement, the moved Wh or Q morpheme starts as part of a doubled structure including the RP or WIS. 
[4]See Chomsky’s ‘On wh movement’ for discussion. See herefor more prose.
[5]Versions of this original idea were developed by many people including Hornstein, Lasnik and Uriagereka, and Fox and Pesetsky. The idea centers on the idea that the problem with movement is that it reorders elements and so can come into conflict with the ordering algorithm. In this sense, gaps are a big deal and what distinguish movement from other kinds of long distance dependencies like binding.
[6]Or again, it argues for treating the operation (e.g. movement) as in need of constraint rather than the output of the operation (a gap, or new linear order). What plausibly unifies cases of “overt” WH (as in English), “covert” WH (as in Chinese) and resumptive WH (as in Hebrew) is that they all involve relating an A’ element to a non-local syntactic position that can be arbitrarily far away. IT’s the span that seems to matter, not what sits at the tail of the chain.
[7]RP constructions must also be formed via I-merge and so too all forms of binding. This requirement fits well with the observation that RPs obey islands. Binding, especially pronominal binding, is likely to be more problematic. As they say at this point in a journal paper in reply to referee 2; these are topics for further research.
[8]As I have noted in other places, this description of the ECP contrast is not quite correct, as we have known since Rizzi’s work on minimality. The distinction seems less a matter of argument vs adjunct than object centered vs non object centered quantification. But this is a topic for another time.

Wednesday, October 31, 2018

FL and the envelope of variability

I read a nice little paper that I would like to bring to your attention. The article(by Alison Henry) is ostensibly about Q-float in varieties of Irish English and it elaborates a point made in earlier work by Jim McCloskey (2000). Jim’s paper used the distribution of quantifiers floated off of WH elements to provide evidence for successive cyclicity (the Qs could be stranded in what are effectively phase edges (aka “escape hatches”)). You are all likely more familiar with this work than I am so I won’t review it here. Henry’s interesting paper makes three additional interesting points:

1.     The paper observes that the Q stranding provides evidence for vP as a phase edge as it seems possible to strand material in this position in several dialects.
2.     The parallel between the stranding facts in A’ movement configurations with that of Q float in A-movement configurations suggests that these should be treated uniformly andthis implies that Q-float in A-movement configurations piggy backs on movement as Sportiche originally suggested rather than these Qs being based generated like adverbials in vP edge positions for semantic reasons.
3.     It suggests an interesting way of parsing these Q stranding effects to shed light on what wrt the phenomenon reflects universal features of FL/UG and what is more G specific.

Like I said, this is a nice little paper and a quick and illuminating read. Before ending, let me say a word or two about the points above.

First, if correct, it provides as the paper notes, interesting evidence for the claim that vP is a phase edge. There is also counter evidence for this claim coming from Keine and Bhatt’s work on non-local agreement in Hindi (thx to Omer for bringing this to my attention). However the Henry data seems compelling, at least to me, and clearly points to something like a landing site under CP for A’-movement. At the very least this now sets up a nice research question for some ambitious syntax grad student: how to reconcile the Irish English data with the Hindi data. Good luck.

The paper’s second point also seems to me quite solid. If the stranding data under A-movement is a proper subset of that under A’-movement then it is hard to see what could motivate treating them as generated by entirely different mechanisms. In fact, theoretically, this would seem to me to be a disaster, invoking the worst kind of constructionism. Linguists have a habit of theoretically reifying surface data in their generative procedures. This leads to multiplication of G operations that have similar effects, which requires enriching the structure of FL/UG. This is a habit to be resisted, IMO. In fact, as a working hypothesis, I believe that we should standardly assume that FL can only do things in one way. There are never two roads to Rome! Henry’s paper shows how productive this kind of assumption can be empirically.

Third, and this IMO is the paper’s most interesting feature, it proposes a very reasonable view of one rich source of variation. Henry’s paper notes that we can find Qs stranded in anyphase edge (and base position) when we take the unionof all the dialects. No singledialect appears to allow Qs to surface in every position. Thus, it might seem as if each dialect has a different Q float mechanism. And, of course, in one sense this is correct. Each G must have somedifference or there would be no dialectal differences. However, as Henry’s paper argues, we can see this another way. The data points to the conclusion that FL/UG actually permits stranding in anyposition but specific LADs acquire Gs with further restrictions. In other words, FL/UG provides an envelope of possibilities that particular Gs further restrict. How? Via learning from the input. The paper makes the plausible point that PLD could fix the specific landing spots allowing the dialect specific G to use the FL/UG provided options as templates for where Qs could appear. This seems to me like a very reasonable idea and allows us to use the full range of variation as a window into the properties of FL/UG.

Two points: First, I have no idea how robust the Q float data are in the PLD and whether there is enough there to fix the various dialects.[1]However, Henry’s speculation can be tested. We are talking about data that should be easy to spot in a CHILDES data base for Irish English (if there is one).  One nice feature of this data: it will all fall into the domain of 0+learning (discussed here) and so be the right rain size to be acquirable via PLD. 

Second, the idea that Henry’s paper illustrates with Q float is one that others (e.g. Idsardi and Lidz and yours truly) have suggested for other syntactic phenomena.[2]We know that I-Merge generates copies in many places and which copy is pronounced should have an impact on surface order given standard linearization procedures. We can put these things together in Henryish fashion and note that what FL/UG provides via Merge is an envelope of possibilities that PLD then winnows out to provide some basic word order templates. On this view, FL/UG provides representations for the class of possible dependencies and PLD provides evidence for selecting among these possibilities wrt linearization. If this is correct, then specificlinearizations in specific Gs are not going to reflect much on the structure of FL/UG though the full range of typological options attested might well do so. At any rate, Henry provides a nice case study of the logic that Idsardi and Lidz were proposing more generally.

Enough said. Like I said, Henry’s paper is interesting and very well written and reasonably compact. Wish I had written it. 



[1]In fact, I have a sneaking suspicion that the range of variation might be more idiolectal than dialectal, but I really do not know enough to ground this suspicion.
[2]Eric Raimy and Lidz have suggested something like this for phonological phenomena as well. They argue that phonological structures are graphs, not strings, and so linearization is as much an issue in phonology as it is in syntax. If you haven’t read this stuff, you should take a look. It’s quite cool.

Friday, March 30, 2018

Three varieties of theoretical research

As FoLers know, I do not believe that linguists (or at least syntacticians) highly prize theoretical work. Just the opposite, in fact. This is why, IMO, the field has tolerated (rather than embraced) the minimalist project (MP) and why so many professionals believe MP to have largely been a failure despite, what (again IMO) is its evident overall successes. As I’ve argued this at length before, I will not do so again here. Rather I would like to report on an interesting paper that I have just re-read that tries to elucidate three distinct kinds of theoretical work. The paper is an old one (published in 2000). It’s called “Thinking about Mechanisms” and the authors, three philosophers, are Peter Machamer, Lindley Darden and Carl Craven (MDC). Here is a link. The paper concentrates on elucidating the notion of a mechanism and argues that it is the key explanatory notion within neurobiology and molecular genetics. The discussion is interesting and I recommend it. In what follows, I would like to pick out some points at random that MDC makes and relate it to linguistic theorizing. This, I hope, will encourage others to look more kindly on theoretical work.

MDC defines the notion of a as follows:

Mechanisms are entities and activities organized such that they are productive of regular changes from start or set-up to finish or termination conditions… To give a description of a mechanism for a phenomenon is to explain that phenomenon, i.e. to explain how it was produced. (3)

So, mechanisms are theoretical constructs whose features (the “entities,” their “properties,” and the “activities” they partake in) explain how phenomena of interest arise. So mechanisms produce phenomena (in biology, in real time) in virtue of the properties of their parts and the activities they engender. [1]

MDC divides a mechanistic description into three parts: (i) Set-up Conditions, (ii) Termination Conditions and (iii) Intermediate activities.

The first, set-up conditions, are “idealized descriptions” of the beginning of the mechanism. Termination conditions are “idealized states or parameters describing a privileged endpoint.” The intermediate steps provide an account of how one gets from the initial set-up to the termination conditions, which describe the phenomenon of interest. (11-12)

This should all sound vaguely familiar. To me it sounds very much like what linguists do in providing a grammatical derivation of a sentence of interest. We start with an initial structure (e.g. a D(eep) S(tructure) representation and explain some feature of a sentence (e.g why the syntactic subject is interpreted as a thematic object) by showing how various operations (i.e. transformations) lead from the initial to the termination state. Doing this explains why the sentence of interest has the properties to be explained. Indeed, it has them in virtue of being the endpoint of the licit derivation provided.

Note too that both mechanisms and GPs focus on idealized situations. GPs describes the linguistic competence of an ideal speaker-hearer or the FL of an idealized LAD. So too with biological mechanisms. They describe idealized hearts or kidneys or electrical conduction at a synapse. Actual instances are not identical to these, though they function in the same ways (it is hoped). No two hearts are the same, yet every idealized heart is identical to any other.

This said, derivations are not actually mechanisms in MDC’s sense for they do not operate in real time (unlike the one’s biologists are typically describing (e.g. synaptic transmission or protein synthesis)). However, generative procedures (GP) are the “mechanisms” of interest within GG for it is (at least in large part) in virtue of the properties of GPs that we explain why native speakers judge the linguistic objects in their native languages as they do and why Gs have the properties they have. Furthermore, as in biology, the aim of linguistics is to elucidate the basic properties of GPs and try to explain why they have the properties they have and not others. So, GPs in linguistics are analogous to mechanisms in other parts of biology. Phenomena are interesting exactly to the degree that they serve to shine light on the fine structures of mechanisms in biology. Ditto with GPs in linguistics.

MDC notes that a decent way to write a history of biology is to trace out the history of its mechanisms. I cannot say whether this is so for the rest of biology, but as regards linguistics, there are many worse ways of tracing the history of modern GG than by outlining how the notion of GP has evolved over the last 60 years.  There is a reasonable argument to be made (and I have tried to make it (see here and following four posts) that the core understanding of GP has become simpler and more general in this period, and that the Minimalist Program is conservative extension of prior work describing the core properties of a human linguistic GP. Not surprisingly, this has analogues in the other kinds of biological theorizing MDC discusses.

So, the core explanatory construct in biology according to MDS is the mechanism. As MDC puts it: “…a mechanistic explanation…renders a phenomenon intelligible…Intelligibility arises not from an explanation’s correctness, but rather from an elucidative relation between the explanans (the set-up conditions and the intermediate entities and activities) and the explanadum (the termination condition or the phenomenon to be explained)” (21).

MDC is at pains to point out that this “elucidative relation” holds regardless of the accuracy of the description. So explanatory potential is independent of truth, and what theory aims at are theories with such potential. Explanatory potential relies on elucidating how something could work, not how it does. The gap between possibility and actuality is critical for the theoretical enterprise. It’s what allows it a certain degree of autonomy.  

For such autonomy to be possible it is critical to appreciate that explanatory potential (what I have elsewhere called “oomph”) is not reducible to regularity of behavior. Again MDC (21-22):

We should not be tempted to follow Hume and later logical empiricists into thinking that the intelligibility of activities (or mechanisms) is reducible to their regularity. Description or mechanisms render the end stage intelligible by showing how it is produced by bottom out entities and activities. To explain is not merely to redescribe one regularity as a series of several. Rather, explanation involves revealing the productive relation. It is the unwinding, bonding, and breaking that explain protein synthesis; it is the binding, bending, and opening that explain the activity of Na+ channels. It is not the regularities that explain the activities but the activities that sustain the regularities.

In other words, mechanisms are not (statistical) summaries of what something regularly does. Regularities/summaries do not (and cannot) explain, and as mechanisms aim to explain they must be more than such summaries no matter how regular. Mechanisms outline how a phenomenon has (or could have) arisen, and this requires outlining the structures and principles that mechanisms deploy to “generate” the phenomenon of interest.[2]

Importantly, it is the relative independence of explanatory potential from truth that allows theory to have an independent existence. MDC suggest three different grades of theoretical involvement summed up in three related but different questions: How possibly? How plausibly? How actually? Let me elaborate.

Explanations are hard. They are hard precisely because they must go beyond recapitulating the phenomenon of interest. Finding the right concepts and putting them together in the right way can be demanding. Here is an example of what I mean (see here for an earlier discussion).

The Minimalist Program (MP) has largely ignored ECP effects of the argument/adjunct asymmetry variety.  Why so? I would contend it is because it is quite unclear how to understand these effects in MP terms. In this respect ECP effects contrast with island effects. There are MP compatible versions of the latter, largely recapitulating earlier versions of Subjacency Theory. IMO, such accounts are not particularly elegant, nor particularly insightful. However it is possible to pretty directly trade bounding nodes for phases, escape hatches for phase edges and Subjacency Principles for Phase Impenetrability Conditions in largely a one for one swap and thereby end up with a theory no worse than the older GB stories but cast in an acceptable MP idiom. This does not constitute a great theoretical leap forward (and so, if this is correct, for these phenomena thinking a la MP does not deepen our understanding), but at least it is clear how island effects could hold within an MP style conception of G. They reduce to Subjacency Effects albeit with all the parts suitably renamed. In other words, the theoretical and conceptual resources of MP are adequate to recapitulate (if not much illuminate) the theoretical and conceptual resources of earlier GB.

This is not so for ECP effects. Why not? Well for several reasons, but the two big ones are that the ECP is a trace licensing condition and the technology behind it appears to run afoul of inclusiveness. Let’s discuss each point in turn.

The big idea behind the ECP is that traces are grammatically toxic unless tamed. They can be tamed by being marked (gamma-marked) by a local antecedent throughout the course of the derivation. The distinction between arguments and adjuncts arises from the assumption that argument A’-chains can be reduced, thereby eliminating their -gamma marked carriers and thereby not cancelling the derivation at LF (recall, -gamma marked expressions kill a derivation). So, traces are toxic, +gamma marking tames them, and deletion acts differently for adjuncts and arguments which is why the former are more restricted than the latter. This, plus a kind of uniformity principle on chains (not a great or intuitive principle IMO, but maybe this is just me) which invidiously distinguishes adjunct from argument chains,[3] yields the desired empirical payoff.

Given the complexity of the ECP data, this is an achievement. Whether it constitutes much of an explanation is something people can disagree about. However, whatever it’s value, it runs afoul of what appear to be basic MP assumptions. For example, MP eschews traces, hence there is little conceptual place for a module of the grammar whose job it is to license them. Second, MP derivations reject adding little diacritics to expressions in the course of a derivation. If indices are technicalia non grata, what to make of +/-gamma marks. Last, MP derivations are taken to be monotonic (No Tampering), hence frowning upon operations that delete information on the “LF” side of a derivation. But deleting –gamma marked traces is what “explains” the argument/adjunct difference. So, the standard GB story doesn’t really fit with basic MP assumptions and this makes it fruitful to ask how ECP effects could possibly be modeled in MP style accounts.  And this is a job for theorists: to come up with a story that could fit, to find the right combination of MP compatible concepts that would yield roughly the right empirical outcomes.[4] The theoretical challenge, then, in the first instance, is to explain how to possibly fit ECP effects into an MP setting, given the elimination of traces and a commitment to derivational monotonicity.

There are additional why questions out there begging for how possible scenarios: e.g. Why case? Why are phrase markers organized so that theta domains are within case/agreement domains that are within A’ (information structure) domains? Why are reflexivization and pronominalization in complementary distribution? Why is selection and subcategorization so local? I could go on and on. These why questions are hard not because we have tons of possible explanatory options but cannot figure out which one to run with, but because we have few candidate theories to run with at all. And that is a theoretical challenge, not just an empirical one. It’s in situations like these that how-possibly becomes a pressing and interesting issue. Sadly, it is also something that many working syntacticians barely attend to.

MDC notes a second level of theoretical involvement: how plausible is a certain possible story. Clearly to ask this requires having a how possible scenario or two sketched out. It is tempting to think that plausibility is largely a matter of empirical coverage. But I would like to suggest otherwise. IMO, plausibility is evaluated along two dimensions: how well the novel theory covers the older (gross) empirical terrain and how many novel lines of inquiry it prompts. A theory is plausible to the degree that it largely conserves the results and empirical coverage of prior theory (what one might call the “stylized facts”) and the degree to which it successfully explains things that earlier theory left stipulative. Again, let me illustrate.

Clearly, plausibility is more demanding than possibility. Plausible theories not only explain, but have verisimilitude (we think that they have a decent chance of being correct). What are the marks of a plausible account? Well, they cover roughly the same empirical territory of the theory they are replacing and they explain what earlier theory stipulated. Here are a couple of examples.

I believe that movement theories of binding and control are plausible precisely because they are able to explain why Obligatory Control (OC) and reflexivization have many of the properties they do. For example, we typically find neither in the subject position of finite clauses (e.g. John expects PRO to/*will win, John expects him(he)self to/*will win). Why not? Well if the movement theory is right, then they are parts of A-chains and so should pattern like what we find in analogous raising constructions (e.g. John was expected t to/*would win), and they do. So the movement theory derives what is largely stipulated in earlier accounts and exposes as systematic relations that earlier theory treated as coincidental (that finite subject positions don’t allow PRO, reflexives or A-traces).  Does his make such accounts true? Nope. But it does enhance their plausibility. Thus, being able to unify these disparate phenomena and provide principled explanations for the distribution of OC PRO and for the relative paucity of nominative reflexives enhances their claims on truth.

Note that here plausibility hinges on (1) accepting that prior accounts are roughly descriptively accurate (i.e. doing what decent science always does; building on past work and insight) and (2) explaining their stipulated features in a principled way. When a story has these two features it moves from possible to plausible. Of course, demonstrating plausibility is not trivial, and what some consider plausibility enhancing others will find wanting. But that is as it should be. The point is not that theorizing is dispositive (nothing is) but that it strives for goals different from empirical coverage (and this is not intended to disparage the latter).

Let me out this another way. When one has a possible explanation in hand it is time to start looking for evidence in its favor. In other words, rather than looking for ways to reject the account one looks for reasons to accept it as a serious one. Trying to falsify (i.e. rigorously test) a proposal has its place, but so does looking for support. However, trying to falsify a possible theory is premature. What one should test are the plausible ones, and that means finding ways to elevate the possible to a higher epistemological plain; the territory of the plausible. That’s what how-plausibly theory aims to do; find the fit between something that is possible and what has come before and showing that the new possible story is a fecund extension of the old. It is an extension in that it covers much of the same territory. It is fecund in that it improves on what came before. This kind of theorizing is also hard to pull off, but like how-possible theory, it relies heavily on theoretical imagination.

Which brings us to how-actually investigations. This is where the theory and the data really meet and where something that family resembles falsification comes into play. Say we have a plausible theory, the next step is to tease out ways of testing its central assumptions. This, no doubt, sounds obvious. But I would beg to differ. Much of what goes on in my little area of linguistics fails to test central postulates and largely concentrates on seeing how to fit current theoretical conceptions to available data (e.g. how to apply a Probe-Goal account to some configuration of agreement/case data). There is nothing wrong with this, of course. But it is not quite “testing” the theory in the sense of isolating its central premises/concepts and seeing how they fly. Let me give you an example.

I personally know of very few critical tests driven by thoughtful theorizing. But I do know of one: the Aoun/Choueiri account of reconstruction effects (RE). The reasoning is as follows: If REs are reflections of the copy theory of movement (as every good Minimalist believes) then where there is no movement, there should be no reconstruction (notice movement is a necessary, not sufficient condition for RE). There is no movement from islands, therefore there should be no RE within islands. Aoun/Choueiri then goes onto argue that resumption in Lebanese Arabic is a movement dependency (Demirdache argued this first I believe) and argues that whereas REs are available when an antecedent binds a resumptive outside an island, they fail systematically to arise with resumptives inside islands. This argues for two central conclusions: (i) that REs are indeed parasitic on movement and (ii) that resumption is a movement dependency. This vindicates the copy theory, and with it a central precept of MP.

For now forget about whether Aoun/Choueiri is right about the facts.[5] The important point here is the logic. The test is interesting because it very clearly implicates key features of current theory: the copy theory of movement, islands as restrictions on movements and REs as piggybacking on copies.  These are three central features and the argument if correct tests them. And this is interesting precisely because they are central ideas in any MP style account. Moreover, it is very clear how the premises bear on the testable conclusion.[6] They can be laid out (that’s where theory comes in BTW, in laying out the premises and showing how together they have certain testable consequences) and a prediction squeezed from them. Moreover, the premises, as noted, are theoretically robust. The Copy Theory of Movement is a core feature of MP architectures, locality as islands are central parts of any reasonable GG theory of syntax. Hence if these pulled apart it would indicate something seriously amiss with how we conceptualize the fundamentals of FL/UG. And that is what makes the Aoun/Choueiri argument impressive.

Like I said, I personally know of only a couple of cases like this. What makes it useful here is that it illustrates how to successfully do how-actually theory (i.e. it is a paradigm case of how-actually theoretical practice). Find consequences of core conceptions and use them to test the core ideas.  We all know that most of what we believe today is likely wrong at least in detail. Knowing this however, does not mean that we cannot test the core features of our accounts. But this requires determining what is central (which requires theoretical evaluation and judicious imagination) and figuring out how to tease consequences from them (which requires analytical acumen). In testing a proposal to see how-actual we need to lead from theory to data, and this means thinking theoretically by respecting the deductive structure that makes a theory the theory it is.

How possible, how plausible, how actual; three grades of theoretical involvement. All are useful. All require attention to the deductive structure of the core ideas that constitute theory. All start with these ideas and move outwards towards the phenomena that, correctly used, can help us refine and improve them. Right now, theoretical work is largely absent from the discipline, at least the how possibly and how plausibly variety. Even the how actually kind is far less common than commonly supposed.


[1] MDC distinguishes “substantivilists” and “process” ontologists” wrt their different understanding of mechanism. The difference appears to reside in whether mechanisms comprise both “entities” and “activities” or whether activities alone suffice. MDC takes it as obvious that reducing entities to activities is hopeless (“As far as we know, there are no activities in …biology…that are not activities of entitites” (5)). I mention this for the discussion is redolent of the current discussion on FoL (Idsardi and Raimy discussing Hale and Reiss) concerning substance free phonology. It is curious that the same kind of discussion takes place in a very different venue and so it is worth taking a look at it in this domain to gain leverage on the one in ours.
[2] As an old friend (Louise Antony) once remarked: in answer to a question like “why did this book drop to the floor when I let go?” it is not helpful to answer that “it always drops whenever anyone lets go.”
[3] I am sure it is not news that the distinction is not accurately described in terms of arguments and adjuncts. But for the record, the absence of pair-list readings of WHs extracted from weak islands seems to show the same acceptability profile as adjuncts even though the WH moves from the complement position. The difference seems to be less argument/adjunct and more individual variable vs higher level variable interpretation.
[4] FWIW, I think that this is where maybe a minimality style explanation of the Rizzi-Cinque variety might be a better fit than the Lasnik-Saito/Barriers approach. But even this story need some detailed reworking.
[5] There is some evidence from Jordanian Arabic contradicting it, though I am not sure whether I believe it yet. Of course, you can take what I believe and still need the full fare of about $2 to get a metro ride in DC.
[6] I often find it surprising how few papers of a purported theoretical nature actually set out their premises clearly and deduce the conclusion of interest. More often, we sue theory like putty and smear it on our favorite empirical findings to see if copiously applied it can be used to hold the data together. Though this method can yield interesting results it is not theory driven and generally fails to address an identifiable theoretical question.
 

Tuesday, February 27, 2018

Universals; structural and substantive

Linguistic theory has a curious asymmetry, at least in syntax.  Let me explain.

Aspects distinguished two kinds of universals, structural vs substantive.  Examples of the former are commonplace: the Subjacency Principle, Principles of Binding, Cross Over effects, X’ theory with its heads, complements and specifiers; these are all structural notions that describe (and delimit) how Gs function. We have discovered a whole bunch of structural universals (and their attendant “effects”) over the last 60 years, and they form part of the very rich legacy of the GG research program. 

In contrast to all that we have learned about the structural requirements of G dependencies, we have, IMO, learned a lot less about the syntactic substances: What is a possible feature? What is a possible category? In the early days of GG it was taken for granted that syntax, like phonology, would choose its primitives (atomic elements) from a finite set of options. Binary feature theories based on the V/N distinction allowed for the familiar four basic substantive primitive categories A, N, V, and P. Functional categories were more recalcitrant to systematization, but if asked, I think it is fair to say that many a GGer could be found assuming that functional categories form a compact set from which different languages choose different options. Moreover, if one buys into the Borer-Chomsky thesis (viz. that variation lives in differences in the (functional) lexicon) and one adds a dash of GB thinking (where it is assumed that there is only a finite range of possible variation) one arrives at the conclusion that there are a finite number of functional categories that Gs choose from and that determine the (finite) range of possible variation witnessed across Gs. This, if I understand things (which I probably don’t (recall I got into syntax from philosophy not linguistics and so never took a phonology or morphology course)), is a pretty standard assumption within phonology tracing back (at least) to Sound Patterns. And it is also a pretty conventional assumption within syntax, though the number of substantive universals we find pale in comparison to the structural universals we have discovered. Indeed, were I incline to be provocative (not something I am inclined to be as you all know), I would say, that we have very few echt substantive universals (theories of possible/impossible categories/features) when compared to the many many plausible structural universals we have discovered. 

Actually one could go further, so I will. One of the major ambitions (IMO, achievements) of theoretical syntax has been the elimination of constructions as fundamental primitives. This, not surprisingly, has devalued the UG relevance of particular features (e.g. A’ features like topic, WH, or focus), the idea being that dependencies have the properties they do not in virtue of the expressions that head the constructions but because of the dependencies that they instantiate. Criterial agreement is useful descriptively but pretty idle in explanatory terms. Structure rather than substance is grammatically key. In other words, the general picture that emerged from GB and more recent minimalist theory is that G dependencies have the properties they have because of the dependencies they realize rather than the elements that enter into these dependencies.[1]

Why do I mention this? Because of a recent blog post by Martin Haspelmath (here, henceforth MH) that Terje Lohndal sent me. The post argues that to date linguists have failed to provide a convincing set of atomic “building blocks” on the basis of which Gs work their magic. MH disputes the following claim: “categories and features are natural kinds, i.e. aspects of the innate language faculty” and they form “a “toolbox” of categories that languages may use” (2-3). MH claims that there are few substantive proposals in syntax (as opposed to phonology) for such a comprehensive inventory of primitives. Moreover, MH suggests that this is not the main problem with the idea. What is? Here is MP (3-4):

To my mind, a more serious problem than the lack of comprehensive proposals is that linguistics has no clear criteria for assessing whether a feature should be assumed to be a natural kind (=part of the innate language faculty).

The typical linguistics paper considers a narrow range of phenomena from a small number of languages (often just a single language) and provides an elegant account of the phenomena, making use of some previously proposed general architectures, mechanisms and categories. It could be hoped that this method will eventually lead to convergent results…but I do not see much evidence for this over the last 50 years. 

And this failure is principled MH argues relying that it does on claims “that cannot be falsified.”

Despite the invocation of that bugbear “falsification,”[2] I found the whole discussion to be disconcertingly convincing and believe me when I tell you that I did not expect this.  MH and I do not share a common vision of what linguistics is all about. I am a big fan of the idea that FL is richly structured and contains at least some linguistically proprietary information. MP leans towards the idea that there is no FL and that whatever generalizations there might be across Gs are of the Greenberg variety.

Need I also add that whereas I love and prize Chomsky Universals, MH has little time for them and considers the cataloguing and explanation of Greenberg Universals to be the major problem on the linguist’s research agenda, universals that are best seen as tendencies and contrasts explicable “though functional adaptation.” For MH these can be traced to cognitively general biases of the Greenberg/Zipf variety. In sum, MH denies that natural languages have joints that a theory is supposed to cut or that there are “innate “natural kinds”” that give us “language-particular categories” (8-9).

So you can see my dilemma. Or maybe you don’t so let me elaborate.

I think that MH is entirely incorrect in his view of universals, but the arguments that I would present would rely on examples that are best bundled under the heading “structural universals.” The arguments that I generally present for something like a domain specific UG involve structural conditions on well-formedness like those found in the theories of Subjacency, the ECP, Binding theory, etc. The arguments I favor (which I think are strongest) involve PoS reasoning and insist that the only way to bridge the gap between PLD and the competence attained by speakers of a given G that examples in these domains illustrate requires domain specific knowledge of a certain kind.[3]
And all of these forms of argument loose traction when the issue involves features, categories and their innate status. How so?

First, unlike with the standard structural universals, I find it hard to identify the gap between impoverished input and expansive competence that is characteristic of arguments illustrated by standard structural universals. PLD is not chock full of “corrected” subjacency violations (aka, island effects) to guide the LAD in distinguishing long kosher movements from trayf ones. Thus the fact that native speakers respect islands cannot be traced to the informative nature of the PLD but rather to the structure of FL. As noted in the previous post (here), this kind of gap is where PoS reasoning lives and it is what licenses (IMO, the strongest) claims to innate knowledge. However, so far as I can tell, this gap does not obviously exist (or is not as easy to demonstrate) when it comes to supposing that such and such a feature or category is part of the basic atomic inventory of a G. Features are (often) too specific and variable combining various properties under a common logo that seem to have little to do with one another. This is most obvious for phi-features like gender and number, but it even extends to categories like V and A and N where what belongs where is often both squishy within a G and especially so across them. This is not to suggest that within a given G the categories might not make useful distinctions. However, it is not clear how well these distinctions travel among Gs. What makes for a V or N in one G might not be very useful in identifying these categories in another. Like I said at the outset, I am not expert in these matters, but the impression I have come away with after hearing these matters discussed is that the criteria for identifying features within and across languages is not particularly sharp and there is quite a bit of cross G variation. If this is so, then the particular properties that coagulate around a given feature within a given G must be acquired via experience with that that particular feature in that particular G. And if this is so, then these features differ quite a bit in their epistemological status from the structural universals that PoS arguments most effectively deploy. Thus, not only does the learner have to learn which features his G exploits, but s/he even has to learn which particular properties these features make reference to, and this makes them poor fodder for the PoS mill.

Second, our theoretical understanding of features and categories is much poorer than our understanding of structural universals. So for example, islands are no longer basic “things” in modern theory. They are the visible byproducts of deeper principles (e.g. Subjacency). From the little I can tell, this is less so for features/categories. I mentioned the feature theory underlying the substantive N,V,A,P categories (though I believe that this theory is not that well regarded anymore). However, this theory, even if correct, is very marginal nowadays within syntax. The atoms that do the syntactic heavy lifting are the functional ones, and for this we have no good theoretical unification (at least so far as I am aware). Currently, we have the functional features we have, and there is no obvious theoretical restraint to postulating more whenever the urge arises.  Indeed, so far as I can tell, there is no theoretical (and often, practical) upper bound on the number of possible primitive features and from where I sit many are postulated in an ad hoc fashion to grab a recalcitrant data point. In other words, unlike what we find with the standard bevy of structural universals, there is no obvious explanatory cost to expanding the descriptive range of the primitives, and this is too bad for it bleaches featural accounts of their potential explanatory oomph.

This, I take it, is largely what MH is criticizing, and if it is, I think I am in agreement (or more precisely, his survey of things matches my own). Where we part company is what this means. For me this means that these issues will tell us relatively little about FL and so fall outside the main object of linguistic study. For MH, this means that linguistics will shed little light on FL as there is nothing FLish about what linguistics studies. Given what I said above, we can, of course, both be right given that we are largely agreeing: if MH’s description of the study of substantive universals is correct, then the best we might be able to do is Greenberg, and Greenberg will tell us relatively little about the structure of FL. If that is the argument, I can tag along quite a long way towards MH’s conclusion. Of course, this leaves me secure in my conclusion that what we know about structural universals argues the opposite (viz. a need for linguistically specific innate structures able to bridge the easily detectable PoS gaps).

That said, let me add three caveats.

First, there is at least one apparent substantive universal that I think creates serious PoS problems; the Universal Base Hypothesis (UBH). Cinque’s work falls under this rubric as well, but the one I am thinking about is the following. All Gs are organized into three onion like layers, what Kleanthes Grohmann has elegantly dubbed “prolific domains” (see his thesis). Thus we find a thematic layer embedded into an agreement/case layer embedded into an A’/left periphery layer.  I know of no decent argument arguing against this kind of G organization. And if this is true, it raises the question of why it is true. I do not see that the class of dependencies that we find would significantly change if the onion were inversely layered (see here for some discussion). So why is it layered as it is? Note that this is a more abstract than your typical Greenberg universal as it is not a fact about the surface form of the string but the underlying hierarchical structure of the “base” phrase marker. In modern parlance, it is a fact about the selection features of the relevant functional heads (i.e. about the features (aka substance) of the primitive atoms). It does not correspond to any fact about surface order, yet it seems to be true. If it is, and I have described it correctly, then we have an interesting PoS puzzle on our hands, one that deals with the organization of Gs which likely traces back to the structure of FL/UG. I mention this because unlike many of the Greenberg universals, there is no obvious way of establishing this fact about Gs from their surface properties and hence explaining why this onion like structure exists is likely to tell us a lot about FL.

Second, it is quite possible that many Greenberg universals rest on innate foundations. This is the message I take away from the work by Culbertson & Adger (see here for some discussion). They show how some order within nominals relating Demonstratvies, Adjectives, Numerals and head Nouns are very hard to acquire within an artificial G setting. They use this to argue that their absence as Greenberg options has a basis in how such structures are learned.  It is not entirely clear that this learning bias is FL internal (it regards relating linear and hierarchical order) but it might be. At any rate, I don’t want anything I said above to preclude the possibility that some surface universals might reflect features of FL (i.e. be based on Chomsky Universals), and if they do it suggests that explaining (some) Greenberg universals might shed some light on the structure of FL.

Third, though we don’t have many good theories of features or functional heads, a lazy perusal of the facts suggest that not just anything can be a G feature or a G head. We find phi features all over the place. Among the phi features we find that person, number and gender are ubiquitous. But if anything goes why don’t we find more obviously communicatively and biologically useful features (e.g. the +/- edible feature, or the +/- predator feature, or the +/- ready for sex feature or…). We could imagine all sorts of biologically or communicatively useful features that it would be nice for language to express structurally that we just do not find. And the ones that we do find, seem from a communicative or biological point of view to often be idle (gender (and, IMO, case) being the poster child for this). This suggests that whatever underlies the selection of features we tend to see (again and again) and those that we never see is more principled than anything goes. And if that is correct, then what basis could there be for this other than some linguistically innate proclivity to press these features as opposed to those into linguistic service.  Confession: I do not take this argument to be very strong, but it seems obvious that the range of features we find in Gs that do grammatical service is pretty small, and it is fair to ask why this is so and why many other conceivable features that we could imagine would be useful are nonetheless absent.

Let me reiterate a point about my shortcomings I made at the outset. I really don’t know much about features/categories and their uniform and variable properties. It is entirely possible that I have underestimated what GG currently knows about these matters. If so, I trust the comments section will set things straight. Until that happens, however, from where I sit I think that MH has a point concerning how features and categories operate theoretically and that this is worrisome. That we draw opposite conclusions from these observations is of less moment than that we evaluate the current state of play in roughly the same way.



[1] This is the main theme of On Wh Movement and I believe what drives the unification behind Merge based accounts of FL.
[2] Falsification is not a particularly good criterion of scientific adequacy, as I’ve argued many times before. It is usually used to cudgel positions one dislikes rather than push understanding forward. That said, in MH, invoking the F word does not really play much more than an ornamental role. There are serious criticisms that come into play.
[3] I abstract here from minimalist considerations which tries to delimit the domain specificity of the requisite assumptions. As you all know, I tend to think that we can reduce much of GB to minimalist principles. The degree to which this hope is not in vain, to that degree the domain specificity can be circumscribed to whatever it is that minimalism needs to unify the apparently very different principles of GB and the generalizations that follow from them.