Comments

Friday, October 16, 2015

A new model for ling journals

Johan Rooryck sent me the following very interesting piece on open access journals for linguistics. It addresses an issue that has been discussed on this blog regarding the practicalities of such a venture. Many of you have weighed in on the topic. This should be useful in assessing the options. Thx Johan.

****

Open Access publishing is often said to be the future of academic journals, but the actual move from a subscription model to an Open Access model is not easily achieved. Several international linguistics journals are currently transitioning from their traditional publisher to a new Open Access publisher, transferring their entire editorial staff, authors, and peer reviewers from the traditional subscription model to Fair Open Access.
Such a transfer is made possible thanks to a new organization called Linguistics in Open Access (Ling-OA) (www.lingoa.eu). Ling-OA is a non-profit foundation representing linguistics journal editors who wish to publish under the conditions of Fair Open Access outlined below:
•    The editorial board or a learned society owns the title of the journal.
•    The author owns the copyright of his articles, and a CC-BY license applies.
•    All articles are published in Full Open Access (no subscriptions, no ‘double dipping’).
•    Article processing charges (APCs) are low (around 400 euros), transparent, and in proportion to the work carried out by the publisher.
Ling-OA facilitates the move to Fair Open Access by paying for the Article Processing charges of the articles published in these journals for the next five years. Although Ling-OA will consider any publisher subscribing to the principles of Fair Open Access, the journals associated with Ling-OA will in principle be published by Ubiquity Press with the Open Library of Humanities as a long-term sustainability partner. OLH, whose platform is also provided by Ubiquity Press, will guarantee the continued publication of the journals associated with Ling-OA after the first five years through its consortial library funding model. OLH is a charitable organization dedicated to publishing Open Access scholarship with no author-facing APCs (https://www.openlibhums.org). In short, Ling-OA pays for APCs during the first 5 years, then OLH takes over these payments. This will provide long-term sustainability for Fair Open Access journals, ensuring that no researcher will ever have to pay for APCs out of their own pocket. The idea of this pay model is that university libraries are key:they largely pay for the expensive subscriptions of scientific journals now, and can be convinced to pay for the less expensive APCs financing the journals in the future.
The first journals to be published by Ubiquity Press, from January 2016, as part of this initiative are LabPhon and the Journal of Portuguese Linguistics. The journals Lingua and Journal of Greek Linguistics are currently renegotiating their collaboration with their publishers. We hope several other journals will transition to the new platform shortly.
Ling-OA has obtained considerable financial guarantees to cover Article Processing Charges for the first 5 years, provided by the Association of Dutch Universities (VSNU) and the Netherlands Organization for Scientific Research (NWO). It enjoys further support from the Royal Netherlands Academy of Arts and Sciences (KNAW), the Consortium of Dutch University Libraries (UKB), and in particular the Radboud University Library, which played a major role in initiating LingOA. The impact of the journals in transition will be monitored by CWTS Leiden (http://www.cwts.nl).

After the successful transition of these journals, Ling-OA hopes to convince the editors of many other linguistics journals to join them. We believe that a successful transition to fair Open Access can be achieved within a relatively small and close-knit discipline such as linguistics. As such, Ling-OA hopes to become a model for the transition to fair Open Access in other disciplines as well.

Monday, October 12, 2015

What's in UG (part 2)

In the previous post (here) I provided some relevant background for discussing a forthcoming paper in Cognition by Cole, Hermon and Yanti (CHY) that argues against a UG interpretation of the Binding Theory (BT), and, by extension, against most PoS forms of reasoning to UG. Here we get into some details. There is one more post to follow. Told you this was going to be involved.

The paper argues that there are languages where the relatively clean morphological distinctions found in English between anaphors and pronominals is considerably less clear. Thus, there are some languages where some “semantically dependent expressions” (SDE, I return in a moment to explicate this notion) are completely exempt from BT and that there are many where there is no clear morphological distinction between reflexives and pronouns (as their appears to be in English). None of this is news. In fact, as CHY notes, and has been known for quite a while even to typological illiterates like me, there are even languages where the morphological distinction between pronominals and anaphors is hard to discern.[1] In other words, there are languages where the functional surface vocabulary (FSV) relevant for binding does not cleanly reflect the underlying categories relevant for BT. Or, to put this another way, it’s easier to study BT in some languages than in others precisely because in some languages the overt FSV more clearly marks the underlying grammatical distinctions than in other languages. Isn’t this one of the reasons for doing comparative syntax? To find those languages where the system of interest is easiest to study? Does it matter that in some languages the system of interest is muddied by a lousy functional inventory? Not that I can see.[2]

However, CHY’s main argument is that this is a problem for BT and that it shows that UG is not key to understanding binding phenomena and that PoS problems do not support nativist conclusions in the domain of binding. Here’s how CHY puts it (p.3):

UG constitutes a general solution to the problem of the poverty of the stimulus, and would be expected to provide the solution to this dilemma in the realm of Binding, as well as other areas of syntax. If, however, Binding is not determined by UG, and, hence, must be learned by the language learner (as we claim below is in fact the case) then Binding must be learnable pace a priori claims to the contrary. For this reason, Binding constitutes a good test case for the claim that the facts of syntax generally cannot be accounted for without making the assumption that much of our knowledge of syntax is innate (i.e. determined by UG).

So the aim of CHY is to argue that BT facts can do without a UGish BT. However, what CHY actually shows is that it is often hard to fix which overt morphemes/expressions (if any) fall under which parts of BT, i.e. to fix the FSVs of a given language’s G. Thus, the fact that BT does not tell us how to categorize overt morphemes wrt binding categories is taken by CHY to show that a UG BT cannot play a part in solving PoS problems in the domain of binding. In other words, CHY argues that because BT is not by itself a full theory of acquisition, (in particular does not provide a general account of fixing FSV) then it does nothing at all. And this is a very poor argument against UG.

Let me be clear exactly what my criticism of CHY is: nowhere does CHY outline just what the PoS problem in the domain of binding is (as I did in the previous post). It doesn’t discuss how facts like (3) in the post here, could be learned purely on the basis of PLD. Nor does CHY argue that the binding principles would not be useful in categorizing morphemes. The CHY argument is entirely based on the observation that the BT categories are often opaque (i.e. that the mapping from overt FSVs to BT relevant categories is often unclear).[3] However, as noted above, this cannot be an argument against the UG nature of BT because the claim that BT is part of UG never assumed that it could solve this problem (all by itself, though it almost surely contributes towards solving it). And a good thing too for we know that whatever is in UG it had better not be a specification of which language particular morphemes are subject to BT-A, B or C. If anything is learned, it’s this.  So, the CHY conclusions do not, as a matter of logic, argue against the claim that BT is part of FL/UG. Let’s consider some further details.

The CHY argument comes in three parts. Part one is the observation that though many languages have a BT functional vocabulary that clearly separates BT anaphors from BT pronominals (see CHY section 2), this is not universally true.[4] In section 3 CHY notes that there are languages (a version of Javanese) with BT exempt anaphors, expressions that are anaphoric (in one sense) but do not fall under BT at all. CHY takes this to be problematic. In particular, CHY claims that the existence of BT exempt anaphors (p. 7):

…constitutes a serious challenge for UG-based approaches to Binding. The presence in a language of a form that is used anaphorically but which is exempt from the Binding requirements of UG would impose a considerable burden on the child acquiring the language…The existence of UG sanctioned categories for anaphora simplifies learning only if all anaphoric elements (in the non technical sense of “anaphoric” that includes both pronouns and reflexives) are subject to UG principles. Thus, in the context of a system containing a mixture of UG compliant and UG exempt elements, UG principles do not provide a solution to the poverty of stimulus problem. At the least, the distinction between exempt and non- exempt forms must be learned from experience…and it is doubtful that the data is sufficiently structured to make this distinction learnable on the basis of the distributional data that are available to the child.

What’s the problem? Not all forms that are “anaphoric” in the “non-technical sense” (i.e. they are semantic but not syntactic dependents) are anaphoric in the BT sense (i.e. subject to BT-A) and, CHY claims, there is no PLD that can help the PLD figure out if an expression belongs in one category or another. Let’s consider these claims much more closely, for they come close to inviting a form of magical thinking.

First, what is the “non-technical sense” of anaphoric? Linguists have two senses of “anaphoric.” The first is “semantic.” An expression is an anaphoric dependent of another expression if its interpretation presupposes/requires/relies on the interpretation of another. This is what I meant by SDEs above. English contains expressions like this. For example, ‘the others’ is an expression that has no interpretation sans a semantic antecedent (i.e. it must have a semantic antecedent). However, it is not subject to BT-A as it can have a non sentence internal antecedent. In fact, all non-deictic pronouns (e.g. cataphoric pronouns) are another example.

(4)  a. Three of the men are wearing tuxedos. The others are sporting jeans.
b. John ate an ice cream. Mary then kissed him.

In (4a), the others depends for its interpretation on three of the men in the previous sentence, as does him on John in (4b). These, then are semantically anaphoric, but BT-A exempt. Are these then a problem for a UG account of BT-A? Only if categorizing them as exempt from BT-A is problematic. So is it? Is there no available positive PLD that might inform the LAD that despite appearances, these expressions are not subject to BT-A (i.e. that despite being semantic anaphors they are not syntactic anaphors)? Well, how about the data in (4). Are these too exotic to be considered part of the PLD? Not from where I sit. Of course, I could be wrong, but it is certainly not “doubtful” that such data might be available to the LAD. In fact, I would go further: cross sentential dependencies of the sort in (4) are almost certainly part of the PLD. [5]

Is Javanese any different? Maybe, but as CHY notes (p. 6) the BT exempt anaphor it discusses can have discourse antecedents, in contrast to the BT-A forms that cannot. But this then is exactly the kind of evidence that would advise the LAD against categorizing them as syntactically anaphoric (i.e. subject to BT-A). If so, there is no PoS problem as regards the categorization of these overt expressions into more abstract syntactic categories (i.e. treating them (or not) as FSVs for the abstract grammatical category BT-anaphor).

Conclusion: CHY’s argument here fails. It is certainly right in concluding that the world would be a better place for the LAD if semantic anaphora were an infallible indicator of syntactic anaphora. A world where visible diagnostics are perfect indicators of underlying structure is a nice place (of course, this is no less true for measles and quarks than for BT-A anaphors). But, this does not mean that a world where the two pull apart (i.e. our own) implies that the distinction cannot be acquired on the basis of PLD. It all depends on the PLD and the learner, and CHY does not consider its own cited data in concluding that the PLD is sadly lacking. Is cross-sentential “anaphora” that exotic in the PLD that the LAD would never have access to it? Maybe, but I would like a lot more than simple assertion to that effect before concluding that it is so.

In fact, we can go further. CHY must assume that this categorization can be fixed on the basis of PLD (contrary to their apparent claim to the contrary). Why so? CHY argues that distinguishing BT-A from BT-A-exempt FSVs is not based on innate features of the LAD. That, in fact, is its main point (which, btw, must be correct given the obvious variation in surface forms). But if it’s not fixed via FL/UG then it must be learned. But to be learned there needs to be evidence that the LAD can use to learn it. Thus, there must be evidence in the PLD relevant to fixing the categorization. I mention this as CHY appears to deny this. Specifically CHY asserts (p. 7):

At the least, the distinction between exempt and non- exempt forms must be learned from experience…and it is doubtful that the data is sufficiently structured to make this distinction learnable on the basis of the distributional data that are available to the child

But as a matter of logic, this pair of claims is close to contradictory (That’s being weaselish. IMO it is a contradiction). If the categorization is learned, then the PLD must be “sufficiently structured” to allow it to be learned. And if the PLD is not “sufficiently structured” to allow it to be learned, then it must be innately determined. There is no third alternative.  There is no learning fairy that fits neatly between these possibilities.[6] Luckily, relevant data appears to be plausibly accessible, and so there is no clear PoS problem as regards binning potential “anaphoric” FSVs into different BT relevant categories.

Last point: recall, that even where it very hard to fix BT exemption on the basis of PLD it would not argue against the classical BT being part of UG. Why? Because CHY does not address how BT-A compliant FSVs acquire all of their distributional properties? In specific, CHY does not address how native speakers acquire competence wrt the data in (3) (see previous post). Recall that this is the PoS problem that BT was developed to address. And in this case, as we outlined in the previous post, there is a clear PoS problem. Thus, the LAD cannot inductively acquire knowledge of facts like (3) (in earlier post) as there is no relevant data in the PLD for fixing this knowledge. This argument stands regardless of how the problems with BT-A exempt expressions is resolves.  In short, it is a very odd criticism of any theory that because it (BT in this case) does not solve a problem that it was never intended to address that it also does not solve the problem that it was constructed to address. And, to further conclude, that because some other theory (not yet provided I may add) solves a problem BT was not constructed to address that the same solution obviously solves the problem that BT did address.

Conclusion: this argument comes nowhere close to showing anything about the UG status of BT. And, for the record, I don’t believe that CHY has shown that categorizing FSVs as BT exempt cannot be inferred from a judicious use of the PLD (and, as I noted above, CHY can’t really assume this either).[7]

The next and last post addresses CHY’s second argument.


[1] The most well-known cases involve languages with so called “long distance anaphors.” Are these really anaphors or just bound pronouns? What’s the difference? Are they both or neither? Tough questions. But does linguistic theory need to provide strict answers? Or does it only need to provide ways of classifying things to the degree that they can be so classified? Can you spot the rhetorical question?
[2] Nor is looking for the easiest window into the mechanism limited to linguists. Ask your favorite neuroscientist about squid axons sometime and see why they were a research favorite for such a long time.
[3] For example, on p 9 CHY states:

…the division of anaphora into reflexives and pronouns cannot be simply a matter of compliance with UG principles…[this] suggest strongly that the pattern modeled in Binding Theory is not primarily due to the interaction among principles of UG…

This suggests that CHY (wrongly) understands BT as aiming to explain how FSV gets mapped into BT categories. And CHY is right. BT does not do this. But it does other things, like explain the pattern in (3), which is not attested in the PLD.

[4] Oddly, IMO, CHY group English and Chinese together as well-behaved languages wrt BT. However, if English is the BT poster child (which it really shouldn’t be, see below) then Chinese makes things less transparent. So, for many speakers, ‘ziji’ can serve both as the overt form of the reflexive and of the long distance anaphor. True, ‘ta ziji’ distributes much like ‘himself,’ but Chinese has a simple ‘self’ form that English does not have and this might well make the categorization problem in Chinese harder than it is in English, at least for the canonical data. Of course, not even English is that well behaved as the well-known facts about picture NPs indicates.
[5] JL, my friendly local acquisitionist, tells me that I am correct in thinking that such data are available in the PLD.
[6] There is a third option, but it is irrelevant to CHY’s concerns. It could be that some feature is not specified by PLD or intrinsically fixed by LAD. In such a case, the “parameter” could be random in the population (see Han, Lidz and Musolino LI (2007) for a worked out example in Korean “dialects”). However, even in this case, there is an innate learning principle at work: if no data relevant to P then flip a coin and set P among UG available options. As I noted, this possibility is not relevant for the CHY argument.
[7] CHY argue against the hypothesis that the BT exempt anaphor is ambiguously either a local anaphor or a pronoun depending on context. I sort of liked this proposal. CHY argues against it. It argues instead that the exempt anaphor is under-specified wrt being an anaphor and a pronominal. CHY proposes that this under-specification allows such anaphors to be simultaneously interpreted anaphors and pronouns (as opposed to being one or the other in any given context). The argument against this involves strict vs sloppy readings under ellipsis. CHY assumes that BT-A anaphors only license sloppy readings under ellipsis. It further notes that the BT exempt anaphor in Javanese it discusses allows strict readings even when the anaphor in the ellipsis licensing antecedent would be locally bound (i.e. even when the ambiguity thesis should treat the relevant structure as one of unambiguous BT-A binding).

The form of this argument is fine (or sorta fine: it relies on an ad hoc stipulation that expressions that are underdetermined wrt anaphor and pronominal features is simultaneously both anaphoric and pronominal semantically. So far as I know, this finely-tuned assumption does not follow from anything else about binding or under-specification, and so is ad hoc.). However the premise seems to me empirically ill founded. Assume, as CHY does, that English is well behaved wrt BT-A then we predict that reflexive binding should never license strict readings under ellipsis in English. However this is false (or there is some challenging counter-evidence). One standard case where the strict reading is acceptable is in (i):
(i)             John1 defended himself1 more competently than his lawyer did (defend him1)
In (i) the “defend John” reading is fine. For me, ditto for sentences like (ii) and (iii).  
(ii)           John1 defended himself1 ably in court. His1 lawyers did not (defend him1 ably)
(iii)          John1 admires himself1 greatly. Nobody else does (admire him1).
Thus, the diagnostic tool CHY uses to argue against the ambiguity thesis is not correct in general as paradigmatic reflexive antecedents can license strict readings in ellipsis sites.

Why?  Well this requires some theory of ellipsis (which CHY does not provide). The technical issue of relevance will be what licenses ellipsis. One standard theory sees ellipsis as deletion under identity. What are the relevant identity parameters? Say that morpho-phonological identity is one such, then the fact that the same FSV in Javanese is ambiguous between a reflexive and a pronoun could allow the strict reading to be readily available (the “reflexive” antecedent would be morpho-phonologically identical to the “pronominal” in the ellipsis site). The expression in the ellipsis site interpreted as a pronoun would be morphologically identical to the expression in the antecedent clause with the expression interpreted as a reflexive. It is well known that such conditions are grammatically relevant (e.g. ATB and Parasitic Gap constructions often require morphological case identity to be licit).
            Last point: it is worth observing that the strict/sloppy dichotomy CHY uses as a diagnostic is not part of BT. In fact, BT is mum concerning these matters. It is a strictly empirical question, which CHY argues one way. My point is not that CHY is wrong, but that the premise is dubious and this leaves the ambiguity hypothesis alive and kicking.

Wednesday, October 7, 2015

What's in UG (part 1)?

This is the first of three posts on a forthcoming Cognition paper arguing against UG. The specific argument is against the Binding Theory. But the form is intended to generalize. The paper is written by excellent linguists, which is precisely why I spend three posts exposing its weaknesses. The paper, because it will appear in Cognition, is likely to be influential. It shouldn’t be. Here’s the first of three posts explaining why.

Let’s start with some truisms: not every property of a language particular G is innate. Here’s another one: some features of G reflect innate properties of the language acquisition device (LAD). Let’s end with a truth (that should be a truism by now but is still contested by some for reasons that are barely comprehensible): some of the innate LAD structure key to acquiring a G is linguistically dedicated (i.e. not cognitively general (i.e. due to UG)). These three claims should be obvious. True truisms. Sadly, they are not everywhere and always recognized as such. Not even by extremely talented linguists. I don’t know why this is so (though I will speculate towards the end of this note), but it is. Recent evidence comes from a forthcoming paper in Cognition (here) by Cole, Hermon and Yanti (CHY) on the UG status of the Binding Theory (BT).[1] The CHY argument is that BT cannot explain certain facts in a certain set of Javanese and Malay dialects. It concludes that binding cannot be innate. The very strong implication is that UG contains nothing like BT, and that even if it did it would not help explain how languages differ and how kids acquire their Gs. IMO, this implication is what got the paper into Cognition (anything that ends with the statement or implication that there is nothing special about language (i.e. Chomsky is wrong!!!) has a special preferential HOV lane in the new Cognition’s review process). Boy do I miss Jacques Mehler. Come back Jacques. Please.

Before getting into the details of CHY, let’s consider what the classical BT says.[2] It is divided into three principles and a definition of binding:

A.   An anaphor must be bound in its domain
B.    A pronominal cannot be bound in its domain
C.    An R-expression cannot be bound

(1)  An expression E binds an expression E’ iff E c-commands E’ and E is co-indexed with E’.

We also need a definition of ‘domain’ but I leave it to the reader to pick her/his favorite one. That’s the classical BT.

What does it say? It outlines a set of relations that must hold between classes of grammatical expressions. BT-A states that if some expression is in the grammatical category ‘anaphor’ then it must have a local c-commanding binder. BT-B states that if some expression is in the category ‘pronominal’ then it cannot have a local c-commanding binder. And BT-C states, well you know what it states, if

Now what does BT not say? It says nothing about which phonetically visible expressions fall into which class. It does not say that every overt expression must fall into at least one of these classes. It does not say that every G must contain expressions that fall into these classes. In fact, BT by itself says nothing at all about how a given “visible” morphologically/phonetically visible expression distributes or what licensing conditions it must enter into. In other words, by itself BT does not tell us, for example, that (2) is ungrammatical. All it says is that if ‘herself’ is an anaphor then it needs a binder. That’s it.

            (2) John likes herself

How then does BT gain empirical traction? It does so via the further assumption that reflexives in English are BT anaphors (and, additionally, that binding triggers morphologically overt agreement in English reflexives). Assuming this, ‘herself’ is subject to principle BT-A and assuming that John is masculine, herself has no binder in its domain, and so violates BT-A above. This means that the structure underlying (2) is ungrammatical and this is signaled by (2)’s unacceptability.

As stated, there is a considerable distance between a linguistic object’s surface form and its underlying grammatical one. So what’s the empirical advantage of assuming something as abstract as the classical BT? The most important reason, IMO, is that it helps resolve a critical Poverty of Stimulus (PoS) problem. Let me explain (and I will do this slowly for CHY never actually explains what the specific PoS problem in the domain of binding is (though they allude to the problem as an important feature of their investigation), and this, IMO, allows the paper to end in intellectually unfortunate places).

As BT connoisseurs know, the distribution of overt reflexives and pronouns is quite restricted. Here is the standard data:[3]

(3) a. John1 likes herself1/*2
b. John1 likes himself1/*2
c. John1 talked to Bill2 about himself1/2/*3
d. John1 expects Mary2 to like himself*1/*2/*3
e. John1 expects Mary2 to like herself*1/2/*3
f. John1 expects himself1/*2/*3 to like Mary2
g. John1 expects (that) he/himself*1/*2/*3 will like Mary2

If we assume that reflexives are BT-A-anaphors then we can explain all of this data. Where’s the PoS problem? Well, lots of these data concern what cannot happen. On the assumption that the ungrammatical cases in (3) are not attested in the PLD, then the fact that a typical English Language Acquisition Device (LAD, aka, kid) converges on the grammatical profile outlined in (3) must mean that this profile in part reflects intrinsic features of the LAD. For example, the fact that kids do not generalize from the acceptability of (3f) to conclude that (3g) should also be acceptable needs to be explained and it is implausible that the LAD infers that that this is an incorrect inference by inspecting unacceptable sentences like (3g), for being unacceptable they will not appear in the PLD.[4] Thus, how LADs come to converge to Gs that allow the good sentences and prevent the bad ones looks like (because it is) a standard PoS puzzle.

How does assuming that BT is part of UG solve the problem? Well, it doesn’t, not all by itself (and nobody ever thought that it could all by itself). But it radically changes it. Here’s what I mean.

If BT is part of UG then the acquisition problem facing the LAD boils down to identifying those expressions in your language that are anaphors, pronominals and R-expressions. This is not an easy task, but it is easier than figuring this out plus figuring out the data distribution in (3). In fact, as I doubt that there is any PLD able to fix the data in (3) (this is after all what the PoS problem in the binding domain consists in) and as it is obvious that any theory of binding will need to have the LAD figure out (i.e. learn) using the PLD which overt morphemes (if any) are BT anaphors/pronominals (after all, ‘himself’ is a reflexive in English but not in French and I assume that this fact must be acquired on the basis of PLD) then the best story wrt Plato’s Problem in the domain of binding is where what must obviously be learned is all that must be learned. Why? Because once I know that reflexives in English are BT anaphors subject to BT-A then I get the knowledge illustrated by the data in (3) as a UG bonus.  That’s how PoS problems are solved.[5] So, to repeat: all the LAD needs do to become binding competent is figure out which overt expressions fall into which binding categories. Do this and the rest is an epistemic freebie.

Furthermore, it’s virtually certain that the UG BT principles act as useful guides for the categorization of morphemes into the abstract categories BT trucks in (i.e. anaphor, pronominal, and R-expression).  Take anaphors. If BT is part of UG it provides the LAD with some diagnostics for anaphoricity. Anaphors must have antecedents. They must be local and high enough. This means that if the LAD hears a sentence like John scratched himself in a situation where John is indeed scratching himself then he has prima facie evidence that ‘himself’ is a reflexive (as it fits A constraints). Of course, the LAD may be wrong (hence the ‘prima facie’ above). For example, say that the LAD also hears pairs of sentences like John loves Mary. She loves himself too and ‘himself’ here is anaphoric to John, then the LAD has evidence that reflexives are not just subject to BT-A (i.e. they are at best ambiguous morphemes and at worst not subject to BT-A at all). So, I can see how PLD of the right sort in conjunction with an innate UG provided BT-A would help with the classification of morphemes to the more abstract categories using simple PLD in the.[6]  That’s another nice feature of an articulate UG.

Please observe: on this view of things UG is an important part of a theory of language learning. It is not itself a theory of learning. This point was made in Aspects, and is as true today as it was then. In fact, you might say that in the current climate of Bayesian excess that it is the obvious conclusion to draw: UG limns the hyporthesis space that the learning procedure explores. There are many current models of how UG knowledge might be incorporated in more explicit learning accounts of various flavors (see Charles Yang’s work or Jeff Lidz’s stuff for some recent general proposals and worked out examples).

Does any of this suppose that the LAD uses only attested BT patterns in learning to classify expressions? Of course not. For example, the LAD might conclude that ‘itself’ is a BT-A anaphor in English on first encountering it. Why? By generalizing from forms it has encountered before (e.g. ‘herself’, ‘themselves’). Here the generalization is guided not by UG binding properties but by the details of English morphology.  It is easy to imagine other useful learning strategies (see note 6). However, it seems likely that one way the LAD will distinguish BT-A from BT-B morphemes will be in terms of their cataphoric possibilities positively evidenced in the PLD.

So, BT as part of UG can indeed help solve a PoS problem (by simplifying what needs to be acquired) and plausibly provides guide-posts towards that classification. However, BT does not suffice to fix knowledge of binding all by itself nor did anyone ever think that it would.  Moreover, even the most rabid linguistic nativist (I know because I am one of these) is not committed to any particular pattern of surface data. To repeat, BT does not imply anything about how morphemes fall into any of the relevant categories or even if any of them do or even if there are any relevant surface categories to fall into.

With this as background, we are now ready to discuss CHY. I will do this in the next post.


[1] I have been a great admirer of both Cole and Hermon’s work for a long time. They are extremely good linguists, much better than I could ever hope to be. This paper, however, is not good at all. It’s the paper, not the people, that this post discusses.
[2] I will discuss the GB version for this is what CHY discusses. I personally believe that this version of BT is reducible to the theory of movement (A-chain dependencies actually). The story I favor looks more like the old Lees & Klima account. I hope to blog about the differences in the very near future.
[3] As GGers also know, the judgments effectively reverse if we replace the reflexive with a bound pronoun. This reflects the fact that in languages like English, reflexives and bound pronouns are (roughly) in complementary distribution. This fact results from the opposite requirements stated in BT-A and BT-B. The same effect was achieved in earlier theories of binding (e.g. Lees and Klima) by other means.
[4] From what I know, sentences like (3g) are unattested in CHILDES. Indeed, though I don’t know this, I suspect that sentences with reflexives in ECM subject position are not a dime a dozen either.
[5] I assume that I need not say that once one figures out which (if any) of the morphemes are pronominals then BT-B effects (the opposite of those in (3) with pronouns replacing reflexives) follow apace. As I need not say this, I won’t.
[6] Please note that this is simply an illustration, not a full proposal. There are many wrinkles one could add. Here’s another potential learning principle: LADs are predisposed to analyze dependencies in BT terms if this is possible. Thus the default analysis is to treat a dependency as a BT dependency. But this principle, again, is not an assumption properly part of BT. It is part of the learning theory that incorporates a UG BT.

Wednesday, September 30, 2015

What are theta roles for?

So here’s my question: What’s the point of theta theory? What does a theta role do? Here’s my impression: we want theta theory to do two different kinds of things and it is not clear to me that any theory can (or should) do both. What are these two things? They are an integral part of the semantic interpretation of a sentence and they are the means by which arguments are linked to syntactic positions in “D-structure.” I should point out that noting these dual desiderata is not original to me, but arises from what I recall were earlier important discussions of these matters by Dowty, Grimshaw, and others.  Nonetheless, I feel that these issues have become more obscure over time and I would like to engage in a rambling re-think. This is all in the way of excusing the shambolic nature of what follows. Hard as it is for me to present clear arguments in general, in this case I am not even going to try. I just want to sorta kinda survey the options and try to clear up my own confusion. Needless to say, I am relying on the kindness of others to clear up the mess. Here goes.

The literature seems to have two different (though possible related, we shall see) desiderata for theta roles:

(1)  Theta roles are required for semantic interpretation
(2)  Theta roles are required to get the LAD from primary linguistic data to a G.

Let’s discuss each of these a little bit. The first view of theta roles treats them as essential semantic notions. Without theta roles, arguments would not have a semantic interpretation and given that Gs map meanings and (“with,” if you are a thoroughly modern minimalist (TMM)) sounds then we need some conception of meaning which is the target of the mapping and theta roles are taken to be one component of a well-formed meaning.

The second view treats theta roles as levers for getting a language acquisition device (aka: a child or LAD) from primary linguistic data (PLD) to a G, most particularly, from PLD to a “D”-structure. The ‘D’ here is in scare quotes here for as any TMM knows we have dispensed with D-structure in the GB sense, yet, so far as I know, every theory of GG has some analogue thereof, including current minimalist accounts. By ‘D-structure’ I just mean the G structure in which arguments are grammatically linked up to their predicates (or vice versa).[1] A G establishes thematic links before establishing any further dependencies that expressions grammatically enter into (e.g. agreement, case, binding, and especially, movement). Fixing this first relation, the one where arguments join predicates, is very important because it is very hard to study all the other dependencies, especially movement, if you have no idea where expressions begin their grammatical lives. 

Is there a necessary relation between these two desiderata? Perhaps, perhaps not (though I suspect not). Here’s what I mean. It might be that the conception of theta role required for semantic interpretation is identical to the one that used to get LADs from PLD to Gs.  However, there is no obvious reason why this need be the case. In particular, the conception of theta role required for semantic interpretation seems to be at a different grain than the one useful from priming the G pump. Let me explain.

One conception of theta role is simply as a place-holder for the notion “argument.” For example, all we mean when we say that some DP has the “agent” theta role is that it is the external argument of some predicate.  The designation “agent” does not mean much, save indicating which of the ordered arguments of a predicate some DP is related to.[2]

The problem with this conception is that it is not clear how it helps with (2). In particular, it is quite unlikely that LADs know the meanings of the predicates they are being exposed to and so it is not clear how they could use this thin sense of theta role to acquire their G. Rather, what we would like is some good coarse rule of thumb that the LAD can use to vault into the G given some PLD. This is where notions like ‘agent’ and ‘patient’ gain their value. Being an agent or patient (a doer or done-to) is plausibly an observational feature of an event participant. In other words, the substantive interpretation of notions like agent and patient plausibly have what Chomsky called “epistemological priority” (EP). They are observable non­-linguistic predicates that can be used to map PLD to (non-observable) grammatical dependencies. An example of such a useful mapping rule would be “Agents are always external arguments, patients always internal arguments.” If every D-structure corresponded to a set of (observable) theta roles with the right linking rules, then we could solve the problem of how an LAD gets from PLD to abstract Gish structures.

Now the problem is that it turns out to be hard come up with such substantive thematic notions that are also plausibly semantically general. Another way of saying this is that though there are plausibly some clear cases of “agenthood,” it is not clear that the same notion extends usefully to all (or even most) verbal subjects. Thus, though kickers may be prototypical agents, lovers may not be. At any rate, one of the well-known problems is that such substantive theta roles have a problem extending to all predicates.

One important and influential solution to this is Dowty’s work (here). It defines super categories of theta roles, collapsing them into two “proto”-flavors; Proto-Agent (P-A) and Proto-Patient (P-P). Proto roles are defined over the full semantics of a predicate indexed to a particular argument position. Thus, P-As are “verbal entailments about the argument in question,” (i.e. those DPs that have more of the “agent” properties than any other DP in that argument structure).[3] On this conception, an argument is P-A if it has a preponderance of the following properties in a given proposition: it is volitional, sentient, a causer of events or changes of states in another participant, a mover, exists independently of event named by verb ((27): 572). Indices of P-P are undergoing a change of state, being an incremental theme, being casually affected by another participant, being stationary relative to movement of another participant, not existing independently of the named event ((28: 572). These are among the contributing factors that Dowty suggests for classifying arguments into one of the proto categories (he is quite clear that these may not exhaust the relevant entailments). Note that on this conception, proto roles are defined in terms of the more articulated semantics of the sentence. In other words, given the meaning of a sentence we can compute a coarser grained classification of arguments into super categories that “average” over the differences. On this conception, proto-roles “loose” information that the actual meaning of the sentence contains.

Not surprisingly, for Dowty proto-roles do not determine semantic interpretations for they presuppose them (i.e. proto-roles are defined in terms of the entailments of the argument in question in the specific proposition). Thus, on this view, proto-roles are not important for (1) above. Their special function (if they are important at all, which is something that Dowty often questions) is to provide an account of how arguments map to syntactic positions given that we know the verbal implications of that argument (i.e. what the proposition means).

IMO, the most interesting version of proto-role theory is Baker’s UTAH version (see here).[4] UTAH directly addresses the problem of how to get from pre-linguistic information into the syntax. The idea is that proto-roles mediate the mapping from PLD to “D-structure,” (e.g. P-As map to underlying subjects and P-Ps to underlying objects). Thus, proto-roles are understood to enjoy epistemological priority and are thus able to mediate a mapping to the linguistic system. What is less clear is that Baker’s understanding of proto-roles is really the same as Dowty’s. Why?

Well first, it seems unlikely, at least to me, that LADs compute proto-roles for a given predicate to see how they map onto the syntax. This presupposes that LADs have a rather rich understanding of the meaning of each predicate prior to having any linguistic analysis of the sentence. Some features of the “scene” may be evident (e.g. on hearing “Fido is biting the ball” it is evident that Fido is an “agent” and the ball a ‘patient’) but it seem to me unlikely that this is a consequence of a computation over the meaning of “bite” indexed to the subject and object positions. Rather, here the two notions are simple primitives applying more or less (im)perfectly to the scene at hand. To get from PLD to G, this kind of sloppy information may suffice (at least for a sufficient number of verbs) but it is unlikely to be based on a prior full understanding of the predicates involved. Rather the opposite. Of course, once the G is engaged, then there is more than theta theory available to guide the LAD. So, for the linking problem, all that UTAH must do is get the LAD into the G, then the G can offer other kinds of linguistic information useful for acquiring the G of interest.

Second, Baker also assumes that the theta roles that solve the linking problem are also inputs to the semantic interpretation of the sentence. Note that this is very different from Dowty. For Dowty, proto-roles are too coarse to provide a semantic interpretation. Baker’s suggestion that theta roles are critical to meaning (rather than notions derived from the meaning) assumes a different conception of linguistic meaning than Dowty’s conception. It is unclear to me whether this conception has been fully articulated.

There is a second influential view of theta roles, one that aims to tie it more tightly to a natural semantics. This eschews proto-roles and develops a more articulated inventory of thematic functions.  So, here we get not just two or three roles but a myriad of these. Agents, causers, experiencers, instruments, goals, sources, beneficiaries, targets of emotion, etc.  This richer conception allows theta structure to explicate argument structure. Theta roles don’t just reflect meaning. They determine it. Here, theta roles are cut thinly enough so that they can support intuitive differences in the meanings of different predicates. Not all agents/causers/experiencers are the same. We need hyphenated versions of these to get the full range of mappings that all the different predicates in a language manifest.

There are two main problems with this conception, I believe. First, as Dowty argues quite persuasively, we really don’t have an even approximately decent theory of what these richer roles are or how to specify them.  In particular, there are many many verbs where it is quite unclear what the theta roles of the relevant arguments is let alone how they differ. The most obvious cases involve symmetrical predicates like ‘face’ (e.g. “Carnegie Hall faces the Carnegie Deli”) or ‘resemble’ (“Bill resembles Sam”). In such cases it is quite difficult to see what thematic difference might distinguish one argument from the other. And this problem generalizes. Why? Because there are many different ways of being an agent and it is not at all clear that a hugger is an agent in the exact same way that a lover is. But if these differences are semantically relevant, then it appears that we will need about as many theta roles as we have predicates. This is effectively Dowty’s point, but in the other direction.  You can’t get from agents directly to huggers as the concept is intended to abstract away from what makes huggers different from lovers.  But if you want to get all the way to the actual semantic role that subjects of these particular predicates play, then you will need a lot of hyphenated theta roles.

Second, it is not clear whether this conception will get you any purchase on (2). Again, as Dowty notes, for this end we want a coarser notion, one that will allow us to map arguments to syntactic structure in some general way. Cutting roles too finely will not yield a simple mapping from roles to structure.

It is worth considering for a minute how these two conceptions interact with the theta criterion. As Grimshaw, among others, noted a long time ago, the theta criterion can make do with a very thin conception of theta role. All it requires is that whatever a theta role is a DP must get one and no more than one of them. It does not matter how we distinguish roles, only that we have some way of tying roles to syntactic positions. The prohibition amounts to the claim that an argument must saturate some position and cannot saturate more than one. So far as the theta criterion goes, we don’t really need a general conception of theta role, only of something like “argument position.” The theta criterion restricts arguments to one and only one of these.

A substantive theory of theta roles, one where the kinds of theta roles we have matter, then only really arises with the linking problem. Here we need theta roles that enjoy EP because grammatical notions are not observables, and so to prime FL, to get us to Gs, we need some notions that can bridge the G non-G divide (i.e. some observables that are (at least weakly) correlated to Gish concepts).

Let me be a little clearer. Subject-hood and object-hood are not observable except via an FL lens. Agent-hood and patient-hood likely are. My (ex) dog Sampson could parse many scenes into agents and patients (doers and done-tos), at least some of the time. If this is so (and I am certain that it is) then these sorts of notions have EP status (they are not parasitic on FL for their viability), and these notions can be used to prime FL via something like UTAH (i.e. agents are subjects, patients are objects). UTAH uses the EP thematic notions to access FL given some PLD.  But, if this is what one needs thematic notions for, then it is not at all clear that every argument in every sentence need have a theta role. All that is required is that enough PLD can be parsed in this way to get the G system off the ground. Once the LAD has accessed FL and started developing a G then these Gish notions can take over/supplement the analysis of the PLD. In other words, theta roles as EPs need not be very general (i.e. cover every conceivable predicate and argument), they just need to be general enough to cover enough PLD predicates to prime FL and get it going. Once FL is engaged then its resources are available for further linguistic analysis. And for this purpose, these notions can be (actually should be) quite coarse as their aim is not to provide an interpretation for the sentence but to just crack open the FL module and make it usable by the LAD, which, when on-line, is then able to provide (more) grammatical ways of analyzing the incoming PLD (e.g. this agrees with that so this is a subject, this is adjacent to the verb so this is the object, etc.).

One might go a step further here, I think. To solve the linking problem you want coarse roles that are not determined by calculating the verbal inferences of an argument. Why? Because this is just too fancy a procedure. You want very coarse indicators, those that Sampson could (and did) use. The problem with proto-roles as understood by Dowty is that they don’t seem to be EPish. They are not so much observables as inferables. What I mean is that to get proto-roles the LAD would need to compute inferences off of pretty sophisticated semantic representations. And these need not be very accessible. Better to have limited coarse-grained properties that fit a small number of available predicates than to have a sophisticated system that generalizes across all predicates. You just don’t need the latter if what you want to do is solve the linking problem.

I’ve rambled on long enough and repeated myself way too much (as if repetition and clarity go hand in hand!). Here is what appears to be the main conclusion: we seem to have been asking theta roles to do two things that don’t obviously pull in the same direction. We want them to provide an interpretation for the sentence and to solve the linking problem. However, the kind of roles we want for the first appear to be different from the kinds of roles we need for the second. IMO, the linking problem is the important one for GG. But if this is right, then having a theory of roles that applies to every DP in every sentence is unnecessary (or at least not obviously required). We need a few gross observational roles that apply to enough PLD predicates to get a G up and running. Once engaged, an LAD gets immediate access to a whole slew of linguistic features that the LAD can effectively use to continue acquiring its G. One this conception, we just don’t need a general theory of theta roles (i.e.one that assigns each argument an interpretive role). Which seems like a good thing given that one does not appear to be currently available or likely to be forthcoming.



[1] Everyone (including advocates of the movement theory of control, e.g. me) assumes that at least the following is accurate: every (contentful (i.e. non pleonastic)) DP enters the derivation through a thematic door. Thus the first relation that any such DP grammatically enters into is a thematic relation. This is also true of every version of minimalism that I am aware of.
[2] Even for neo-Davidsonians like Scheine and Pietroski where theta roles serve an important type-lifting role (they are relational predicates that tie a DP to an event variable), all that is generally required is a distinction between internal vs external argument. What flavor these are (whether they are agents or experiencers or causes or…) does not really matter much. The same is even truer for standard conceptions where arguments are effectively related to their predicates by saturating a variable position of the predicate via lambda conversion.
[3] Arguments actually, for it not defined syntactically but over the propositional structure. We can say that a DP has the proto-role in virtue of representing the relevant argument. I will leave such niceties aside here.
[4] This is an online version of the paper that appeared in Haegeman’s edited volume Elements of Grammar. It is a great paper. One of those that I wish that I had written.