Comments

Monday, September 12, 2016

The Generative Death March, part 2.

I’m sitting here in my rocking chair, half dozing (it’s hard for me to stay awake these days) and I come across this passage from the Scientific American piece by Ibbotson and Tomasello (henceforth IT):

“And so the linking problem—which should be the central problem in applying universal grammar to language learning—has never been solved or even seriously confronted.”

Now I’m awake. To their credit, IT correctly identifies the central problem for generative approaches to language acquisition. The problem is this: if the innate structures that shape the ways languages can and cannot vary are highly abstract, then it stands to reason that it is hard to identify them in the sentences that serve as the input to language learners. Sentences are merely the products of the abstract recursive function that defines them, so how can one use the products to identify the function? As Steve Pinker noted in 1989 “syntactic representations are odorless, colorless and tasteless.” Abstractness comes with a cost and so we are obliged to say how the concrete relates to the abstract in a way that is transparent to learners.

And IT correctly notes that Pinker, in his beautifully argued 1984 book Language Learnability and Language Development, proposed one kind of solution to this problem. Pinker’s idea was based on the idea that there are systematic correspondences between syntactic representations and semantic representations. So, if learners could identify the meaning of an expression from the context of its use, then they could use these correspondences to infer the syntactic reprentations. But, of course, such inferences would only be possible if the syntax-semantics correspondences were antecedently known. So, for example, if a learner knew innately that objects were labeled by Noun Phrases, then hearing an expression (e.g., “the cat”) used to label an object (CAT) would license the inference that that expression was a Noun Phrase. The learner could then try to determine which part of that expression was the determiner and which part the noun. Moreover, having identified the formal properties of NPs, certain other inferences would be licensed for free. For example, it is possible to extract a wh-phrase out of the sentential complement of a verb, but not out of the sentential complement of a noun:

(1) a. Who did you [VP claim [S that Bill saw __]]?
b.   * Who did you make [NP the claim [S that Bill saw __]]?

Again, if human children knew this property of extraction rules innately, then there would be no need to “figure out” (i.e., by general rules of categorization, analogy, etc) that such extractions were impossible. Instead, it would follow simply from identifying the formal properties that identified an expression as an NP, which would be possible given the innate correspondences between semantics and syntax. This is what I would call a very good idea. 

Now, IT seems to think that Pinker’s project is widely considered to have failed [1]. I’m not sure that is the case. It certainly took some bruises when Lila Gleitman and colleagues showed that in many cases, even adults can’t tell from a context what other people are likely to be talking about. And without that semantic seed, even a learner armed with Pinker’s innate correspondence rules wouldn’t be able to grow a grammar. But then again, maybe there are a few “epiphany contexts” where learners do know what the sentence is about and they use these to break into the grammar, as Lila Gleitman and John Trueswell have suggested in more recent work. But the correctness of Pinker’s proposals is not my main concern here. Rather, what concerns me is the 2nd part of the quotation above, the part that says the linking problem has not been seriously confronted since Pinker’s alleged failure [2]. That’s just plain false.

Indeed, the problem has been addressed quite widely and with a variety of experimental and computational tools and across diverse languages. For example, Anne Christophe and her colleagues have demonstrated that infants are sensitive to the regular correlations between prosodic structure and syntactic structure and can use those correlations to build an initial parse that supports word recognition and syntactic categorization. Jean-Remy Hochmann, Ansgar Endress and Jacques Mehler demonstrated that infants use relative frequency as a cue to whether a novel word is likely to be a function word or a content word. William Snyder has demonstrated that children can use frequent constructions like verb-particle constructions as a cue to setting an abstract parameter that controls the syntax of a wide range of complex predicate constructions that may be harder to detect in the environment. Charles Yang has demonstrated that the frequency of unambiguous evidence in favor of a particular grammatical analysis predicts the age of acquisition of constructions exhibiting that analysis; and he built a computational model that predicts that effect. Elisa Sneed showed that children can use information structural cues to identify a novel determiner as definite or indefinite and in turn use that information to unlock the grammar of genericity. Misha Becker has argued that the relative frequency of animate and inanimate subjects provides a cue to whether a novel verb taking an infinitival complement is treated as a raising or control predicate, despite their identical surface word orders. In my work with Josh Viau, I showed that the relative frequency of animate and inanimate indirect objects provides a cue to whether a given ditransitive construction treats the goal as asymmetrically c-commanding the theme or vice versa, overcoming highly variable surface cues both within and across languages. Janet Fodor and William Sakas have built a large scale computational simulation of the parameter setting problem in order to illustrate how parameters could be set, making important predictions for how they are set. I could go on [3].

None of this work establishes the innateness of any piece of the correspondences. Rather it shows that it is possible to use the correlations across domains of grammar in order to make inferences on the basis of observable phenomena in one domain to the abstract representations of another.  The Linking Problem is not solved, but there is a large number of very smart people working hard to chip away at it. 

The work I am referring to is all easily accessible to all members of the field, having been published in the major journals of linguistics and cognitive science. I have sometimes been told, by exponents of the Usage Based approach and their empiricist cousins, that this literature is too technical, that, “you have to know so much to understand it.” But abbreviation and argot are inevitable in any science, and a responsible critic will simply have to tackle it. What we have in IT is an irresponsible cop out from those too lazy to get out of their armchairs.

I think it’s time for my nap. Wake me up when something interesting happens.

____________________________________

[1] IT also thinks that something about the phenomenon of ergativity sank Pinker’s ship, but since Pinker spent considerable time in both his 1984 and 1989 books discussing that phenomenon, I think these concerns may be overstated.

[2] You can sign me up to fail like Pinker in a heartbeat.

[3] A reasonable review of some of this literature, if I do say so myself, can be found in Lidz and Gagliardi (2015) How Nature Meets Nurture: Statistical Learning and Universal Grammar. Annual Reviews of Linguistics 1. Also, the new Oxford Handbook of Developmental Linguistics (edited by Lidz, Snyder and Pater) is also full of interesting probes into the linking problem and other important concerns.

Sunday, September 11, 2016

Universals: a consideration of Everett's full argument

I have consistently criticized Everett’s Piraha based argument against Chomsky’s conception about Universal Grammar (UG) by noting that the conclusions only follow if one understands ‘universal’ in Greenberg rather than Chomsky terms (e.g. see here).  I have recently discovered that this is correct as far as it goes, but it does not go far enough. I have just read this Everett post, which indicates that my diagnosis was too hasty. There is a second part to the argument and the form is actually one of a dilemma: either you understand Chomsky’s claims about recursion as a design feature of UG in Greenbergian terms OR Chomsky’s position is effectively unfalsifiable (aka: vacuous). That’s the full argument.  I (very) critically discuss it in what follows. The conclusion is that not only does it fail to understand the logic of a CU conception of universal, it also presupposes a rather shallow Empiricist conception of science, one in which theoretical postulates are only legitimate if directly reflected in surface diagnostics. Thus, Everett’s argument gains traction only if one mistakes Chomsky Universals (CU) for Greenberg Universals (GUs), misunderstands what kind of evidence is relevant for testing CUs and/or tacitly assumes that only GUs are theoretically legit conceptions in the context of linguistic research. In short, the argument still fails, even more completely than I thought.

Let’s give ourselves a little running room by reviewing some basic material. GUs are very different from CUs. How so?

GUs concern the surface distributional properties of the linguistic objects (LO) that are the outputs of Gs. They largely focus on the string properties of these LOs.  Thus, one looks for GUs by, for example, looking at surface distributions cross linguistically. For example, one looks to see if languages are consistent in their directionality parameters (If ‘OP’ then ‘OV’). Or if there are patterns in the order in nominals of demonstratives, modifiers and numerals wrt the heads they modify (e.g. see here for some discussion).

CUs specify properties of the Faculty of Language (FL). FL is the name given to the mental machinery (whatever its fine structure) that outputs a grammar for L (GL) given Primary Linguistic Data from language L (PLDL). FL has two kinds of design features. The linguistically proprietary ones (which we now call UG principles) versus the domain general ones, which are part of FL but not specific to it. GGers  investigate the properties of FL by, first, investigating the properties of language particular Gs and second, via the Poverty of the Stimulus argument (POS). POS aims to fix the properties of FL by seeing what is needed to fill the gap between information provided about the structure of Gs in the PLD and the actual properties that Gs have. FL has whatever structure is required to get the Language Acquisition Device (LAD) from PLDL to GL for any L. Why any L? Because any kid can acquire any G when confronted with the appropriate PLD. 

Now on the face of it, GUs and CUs are very different kinds of things. GUs refer to the surface properties of G outputs. CUs refer to the properties of FL, which outputs Gs. CUs are ontologically more basic than GUs[1] but GUs are less abstract than CUs and hence epistemologically more available.

Despite the difference between GUs and CUs, GGers sometimes use string properties of the outputs of Gs to infer properties of the Gs that generate these LOs. So, for example, in Syntactic Structures Chomsky argues that human Gs are not restricted to simple finite state grammars because of the existence of sentences that allow non-local dependencies of the sort seen in sentences of the form ‘If S1 then S2’. Examples like this are diagnostic of the fact that the recursive Gs native speakers can acquire must be more powerful than simple FSGs and therefore that FL cannot be limited to Gs with just FSG rules.[2]  Nonetheless, though GUs might be useful in telling you something about CUs, the two universals are conceptually very different, and only confusion arises when they are run together.

All of this is old hat, and I am sorry for boring you. However, it is worth being clear about this when considering the hot topic of the week, recursion, and what it means in the context of GUs and CUs. Let’s recall Everett’s dilemma. Here is part 1:

1.     Chomsky claims that Merge is “a component of the faculty of language,” (i.e. that it is a Universal).[3]
2.     But if it is a universal then it should be part of the G of every language.
3.     Piraha does not contain Merge.
4.     Therefore Chomsky is wrong that Merge is a Universal.

This argument has quite a few weak spots. Let’s review them.

First, as regards the premises (1) and (2), the argument requires assuming that if something is part of FL, a CU, then it appears in every G that is a product of FL. For unless we assume this, it cannot be that the absence of Merge in Piraha is inconsistent with the conclusion that it is part of FL. But, the assumption that Merge is a CU does not imply that it is a GU. It simply implies that FL can construct Gs with that embody (recursive) Merge.[4] Recall, CUs describe the capacities of the LAD not its Gish outputs. FL can have the capacity to construct Merge containing Gs even if it can also construct Gs that aren’t Merge containing Gs. Having the capacity to do something does not entail that the capacity is always (or even ever) used. This is why a claim like Everett’s that argues from (3), the absence of Merge in the G of Piraha, does not argue against Merge as part of FL.

Second, what is the evidence that Merge is not part of Piraha’s G? Everett points to the absence of “recursive structures” in Piraha LOs (2). What are recursive structures? I am not sure, but I can hazard a guess. But before I do so, let me note that recursion is not properly a predicate of structures but of rules. It refers to rules that can take their outputs as inputs. The recursive nature of Merge can be seen from the inductive definition in (5):

5.   a. If a is a lexical item then a is a Syntactic Object (SO)
b. If a is an SO and b is an SO then Merge(a,b) is an SO

With (5) we can build bigger and bigger SOs, the recursive “trick” residing in the inductive step (5b). So rules can be recursive. Structures, however, not so much. Why? Well they don’t get bigger and bigger. They are what they are. However, GGers standardly illustrate the fact of recursion in an L by pointing to certain kinds of structures and these kinds have come to be fairly faithful diagnostics of a recursive operation underlying the illustrative structures. Here are two examples both of which have a phrase of type A embedded in another one of type A.

6.     S within an S: e.g. John thinks that Bill left in which the sentence Bill left is contained within the larger sentence John thinks that Bill left.
7.     A nominal within a nominal: e.g.  John saw a picture of a picture where the nominal a picture is contained within the larger nominal a picture of a picture.

Everett follows convention and assumes that structures of this sort are diagnostic of recursive rules. We might call them reliable witnesses (RW) for recursive rules. Thus (6)/(7) are RWs for the claim that the rule for S/nominal “expansion” can apply repeatedly (without bound) to their outputs.

Let’s say this is right. It does not imply that the absence of RWs implies the absence of recursive rules. As it is often said: the absence of evidence is not evidence of absence. Merge may be applying even though we can find no RWs diagnostic of this fact in Piraha.

Moreover, Everett’s post notes this. As it says: “…the superficial appearance of lacking recursion does not mean that the language culd not be derived from a recursive process like Merge. And this is correct” (2-3). Yes it is. Merge is sufficient to generate the structures of Piraha. So, given this, how can we know that Piraha does not employ a Merge like operation?

So far as I can tell, the argument that it doesn’t is based on the assumption that unless one has RWs for some property X one cannot assume that X is a characteristic of G. So absent visible “recursive structures” we cannot assume that Merge obtains in the G that generates these structures. Why?  Because Merge is capable of generating unboundedly big (long and deep) structures, and we have no RWs indicating that the rule is being recursively applied. But, and this I really don’t get; the fact that Merge could be used to generate “recursive structures” does not imply that in any given G it must so apply. So how exactly does the absence of RWs for recursive rule application in Piraha (note, I am here tentatively conceding that Everett’s factual claims might be right (which is likely incorrect, for they are likely wrong)) show that Merge is not part of a Piraha G? Maybe Piraha Gs can generate unboundedly large phrases but then applies some filters to the outputs to limit what surfaces overtly (this in fact appears to be Everett’s analysis).[5] In this sort of scenario, the Merge rule is recursive and can generate unboundedly large SOs but the interfaces (to use minimalist jargon) filters these out preventing the generated structures from converging. On this scenario, Piraha Gs are like English or Braizilain Portuguese or … Gs, but for the filters.

Now, I am not saying that this is correct. I really don’t know and I leave the relevant discussions to those that do.[6] But, it seems reasonable and if this is indeed what the right G analysis for Piraha is, then it too contains a recursive rule (aka, Merge) though because of the filters it does not generate RWs (i.e. the whole Piraha G does not output “recursive structures”).

Everett’s post rejects this kind of retort. Why? Because such “universals cannot be seen, except by the appropriate theoretician” (4). In other words, they are not surface visible, in contrast to GUs, which are. So, the claim is that unless you have a RW for a rule/operation/process you cannot postulate that rule/operation/process exists within that G. So, absence of positive evidence for (recursive) Merge within Piraha is evidence against (recursive) Merge being part of Piraha G. That’s the argument. The question is why anyone should accept this principle?

More exactly, why in the case of studying Gs should we assume that absence of evidence is evidence of absence, rather than, for example, evidence that more than Merge is involved is yielding the surface patterns attested. This is what we do in any other domain of inquiry. The fact that planes fly does not mean that we throw out gravity. The fact that balls stop rolling does not mean that we dumb inertia, the fact that there are complex living systems does not mean that entropy doesn’t exist. So why should the absence of RWs for recursive Merge in Piraha imply that Piraha does not contain Merge as an operation?

In fact, there is a good argument that it does. It is that many many other Gs have  RWs for recursive Merge (a point that Everett accepts). So, why not assume that Piraha Gs do too? This is surely the simplest conclusion (viz. that Piraha Gs are just like other Gs fundamentally) if it is possible to make this assumption and still “capture” the data that Everett notes.  The argument must be that this kind of reply is somehow illicit. What could license the conclusion that it is?

I can only think of one: that universals just are summaries of surface patterns. If so, then  without  surface patterns that are RWs for recursion in a given G means that there are no recursive rules in that G for there is nothing to “summarize.” All G generalizations must be surface “true.” The assumption is that it is scientifically illicit to postulate some operation/principle/process whose surface effects are “hidden” by other processes. The problem then with considering a theory according to which Merge applies in Piraha but its effects are blocked so that there are no RW-like “recursive structures” to reliably diagnose that it is there is that such an assumption is unscientific!

Note that this methodological principle applied in the real sciences would be considered laughable. Most of 19th century astronomy was dedicated to showing that gravitation regulates planetary motion despite the fact that planets do not appear to move in accord with the inverse square law. The assumption was made that some other mass was present and it was responsible for the deviant appearances. That’s how we discovered Neptune (see here for a great discussion). So unless one is a methodological dualist, there is little reason to accept Everett’s presupposed methodological principle.

It is worth noting that an Empiricist is likely to endorse this kind of methodological principle and be attracted to a Greenberg conception of universals. If universals are just summaries of surface patterns then absent the pattern there is no universal at play.

Importantly, adopting this principle runs against all of modern GG, not just the most recent minimalist bit. Scratch any linguist and s/he will note that you can learn a lot about language A by studying language B. In particular, modern comparative linguistics within GG assumes this as a basic operating principle. It is based on the idea that Gs are largely similar and that what is hard to see within the G of language A might be pretty easy to observe in that of language B. This, for example, is why we often conclude that Irish complementizer agreement tells us something about how WH movement operates in English, despite there being very little (no?) overt evidence for C to C movement in English. Everett’s arguments presuppose that all of this reasoning is fallacious. His is not merely an argument against Merge, but a broadside on virtually all of the cross-linguistic work within GG for the last 30+ years.

Thankfully, the argument is really bad. It either rests on a confusion between GUs and CUs or rests on bogus (dualist) methodological principles. Before ending, however, one more point.

Everett likes to say that a CU conception of universals is unfalsifiable. In particular, that a CU view of universals robs these universals of any “predictive power” (5). But this too is false.  Let’s go back to Piraha.

Say you take recursion to be a property of FL then what would you conclude if you ran into speakers that spoke a language without RWs for that universal? You would conclude that they could learn a language where that universal has clear overt RWs. So, assume (again only for the sake of discussion) that you find a Piraha speaker sans recursive G. Assuming that recursion is part of FL you predict that such speakers could acquire Gs that are clearly recursive. In other words, you would predict that Piraha kids would acquire, to take an example at random, Brazilian Portuguese just like non-Piraha kids do. And, as we know, they do! So, taking recursion to be a property of FL makes a prediction about the kinds of Gs LADs can/do acquire. And these predictions seem to be correct. So, postulating CUs does have empirical consequences and it does make predictions, it’s just that it does not make predictions about whether CUs will be surface visible in every L (i.e. provide RWs in every L) and there is no good reason that they should.

Everett complains in this post that people reject his arguments because they confuse GUs and CUs and that this is incorrect (i.e. they don’t make this confusion). However, it is clear that there is lots of other confusion and lots of methodological dualism and lots of failure to recognize the kinds of “predictions” a CU based understanding of universals does make. Both prongs of the dilemma that the argument against CUs rests on collapse on pretty cursory inspection. There is no there there.

Let me end with one more observation, one I have made before. That recursion is part of FL is not reasonably debatable. What kind of recursion there is, and how it operates is very debatable. The fact is not debatable because it is easy to see its effects all around you. It’s what Chomsky called linguistic productivity (LP). LP, as Chomsky has noted repeatedly, requires that linguistic competence involve knowledge of a G with recursive rules. Moreover, that any child can acquire any language implies that every child comes equipped to the language acquisition task with the capacity to acquire a recursive G. This means that the capacity to acquire a recursive G (i.e. to have an operation like Merge) must be part of every human FL.  This is a near truism and, as Chomsky (and many others, including moi) have endlessly repeated, it is not really contestable. But there is a lot that is contestable. What kind of rules/operations do Gs contain (e.g. FSGs, PSGs, TGs, MPs?)? Are these rules/operations linguistically proprietary (i.e. part of UG or not?)? How do Gs interact with other cognitive systems, etc.? These are all very hard and interesting empirical questions which are and should be vigorously debated (and believe me, they are). The real problem with Everett’s criticism is that it has wasted a lot of time by confusing the trivial issues with the substantive ones. That’s the real problem with the Piraha “debate.” It’s been a complete waste of time.



[1] By this I mean that whereas you are contingently a speaker of English and English is contingently SVO it is biologically necessary that you are equipped with an FL. So, Norbert is only accidentally a speaker of English (and so has a GEnglish) and it is only contingently the case that English is SVO (it could have been SOV as it once was). But it is biologically necessary that I have an FL. In this sense it is more basic.
[2] Actually, they are diagnostic on the assumption that they depict one instance of an unbounded number of sentences of the same type. If one only allows finite substitutions in the S positions then a more modest FSG can do the required work.
[3] Quote from 202 Science paper with Fitch and Hauser.
[4] I henceforth drop the bracketed modifier.
[5] Thanks to Alec Marantz for bringing this to my attention.
[6] This point has already been made By Nevins, Pesestsky and Rodrigues in their excellent paper.

A review of Wolfe worth the time to read

David Pesetsky posted a comment linking to this review of Wolfe's book by Joshua Leach. I bring it to your attention here for it is readable, incisive and informed.

There are two large questions that the Wolfe book invites. First, what could motivate someone like Wolfe to write the drivel that he wrote? The second: why does this junk gets positive reviews in our high brow press (HBP)?

The answer to the first question is not that interesting, I believe. Grab your favorite psychological theory of personality and have at it. Wolfe is not the first person who wants to be the smartest person in the room, nor will he be the last.

The second question is more interesting, for it tells us something about our current institutions and the "intellectuals" who aspire to lead them. Oddly, the recent coverage in the HBP does quite a bit to confirm Chomsky's poignant discussion in "The Responsibility of Intellectuals." The HBP feels threatened and something about Chomsky threatens it/them. That's the story. How can such junk get so much non-critical coverage? As Leach points out, Trumpism has moved up stream and has ensconced its intellectual values in the smart New Yorker, NYT reading set. That's the big story. Read the piece. It's very good.

Saturday, September 10, 2016

The Generative Death March, Part 1

It’s been a bad few weeks to be a Chomskyan. It seems like everywhere you turn, someone is claiming that the core ideas of generative linguistics are relics of a bygone era, that the original lessons of the cognitive revolution can be discarded and that those of us who study the language faculty from the standpoint of classical cognitive science are standing around throwing hissy-fits while the real scientists repeatedly show us how the facts on the ground disprove every idea we ever had. 

In Scientific American, Ibbotson and Tomasello (henceforth IT), here, take issue with two key features of the generative enterprise. First, IT tells us that languages do not have “an underlying computational structure”. Then IT tells us that there is no special biological foundation for language, no innate structure that determines how languages can and cannot be structured. IT tells us that “instead, grammar is the product of history (the processes that shape how languages are passed from one generation to the next) and human psychology (the set of social and cognitive capacities that allow generations to learn a language in the first place).”

As I inch along to my death alongside the ideas of generative linguistics, IT looks not like a revolutionary new idea, but the revival of some very old and discredited ideas with no answer to the fundamental questions that bankrupted them to begin with.

IT characterize their revolution thus: “They [children, JL] inherit the mental equivalent of a Swiss Army knife: a set of general-purpose tools—such as categorization, the reading of communicative intentions, and analogy making, with which children build grammatical categories and rules from the language they hear around them.”

The generative linguist is more than happy to grant the child these general purpose tools. Surely the child must be able to identify regularities in the environment, to read communicative intentions and to draw analogies. The central idea of Universal Grammar is that these things are not sufficient to explain the character of the languages we come to acquire. And the goal of the generative linguist is to add as little as possible to UG so that these abilities have increased potency. UG provides the dimensions along which analogies can be drawn and constrains the character of the representations that learners construct when faced with language data. It makes no claims about the trajectory of development, about the contributions of social cognition, memory architectures, cognitive control or statistical inference mechanisms. Rather it says, that given these endowments, we also need one more thing. Thus, any evidence that any nonlinguistic cognitive systems are involved in language acquisition are entirely silent about the existence of UG, unless it can be shown that they do work that we otherwise thought UG responsible for. Note also that the generativist would generally be delighted to learn that something they thought fell into their purview is better explained by something extralinguistic for it allows UG to be smaller, which everyone agrees is the strongest scientific position. 

So, what kinds of things is UG responsible for? In an earlier post here, I worked through one example concerning the treatment of subjects and the interpretive asymmetries between sentences like these:

(1) a. Norbert knows how proud of himself Alexander was after the conference.
b. Norbert knows which picture of himself Alexander posted after the conference.

(1b) is ambiguous in a way that (1a) is not, a fact not exhibited in speech to children and not obviously explained by factors external to grammar.

Here is another relevant case, dating back to several papers from the 1970s by Chomsky and Bresnan:

(2) a.  Valentine is a good value-ball player and Alexander is too
b. Valentine is a better value-ball player than Alexander is

In both of these examples, there is no pronounced predicate in the second clause, but we fill in this predicate in our minds as equivalent to the predicate in the first clause (i.e., a good value-ball player). Is this unpronounced predicate represented in the same way in the two sentences? Evidence suggests not. For example, in some contexts, they behave differently.

(3)  a. Valentine is a good value-ball player and I think Alexander is too
b. Valentine is a better value-ball player than I think Alexander is 
c.  Valentine is a good value-ball player and I heard a rumor that Alexander is too
d. * Valentine is a better value-ball player than I heard a rumor that Alexander is

The fact to be explained here is why the child learner when building representations for (2) doesn’t treat the silent predicate in the same way in the two cases. Both can be interpreted as identical to the main clause predicate in (3a/b); however, this dependency can only hold across the expression “hear a rumor that…” in the coordinate (3c) but not the comparative (3d). It is an analogy that could be drawn but apparently isn’t. Moreover, it seems that (2b) has a structure analogous to the structure of interrogatives.

4) a.  What do Valentine and Alexander like to play together?
b. What do you think that Valentine and Alexander like to play together?
c. * What did you hear a rumor that Valentine and Alexander like to play together?
In (4) there is a dependency between the wh-phrase “what” and the verb play. In (4b), we see that this dependency can be established across multiple clauses (just like the coordinate and comparative ellipses), and in (4c) we see that it cannot be established across “hear a rumor that” (like the comparative ellipsis and unlike the coordinate ellipsis).

Evidently the analogy that the child learner draws when acquiring English is that comparatives have the same kind of structure as wh-questions. Why do they draw this analogy and not the analogy between the comparative and the coordinate ellipsis, which shares more obvious surface features? These patterns, both the analogies that our grammars make and the ones that are tempting but not taken, have been at the center of the generative enterprise since the 1960s. They hold this privileged place because they invite grammar-internal explanations in the form of computational/representational mechanisms out of which sentences are built. To my knowledge, nobody in the field of usage-based linguistics has even attempted to show how such facts follow from “categorization, the reading of communicative intentions and analogy making.” Their silence suggests one of two things: (a) that their swiss-army knife doesn’t have the right tool, or (b) that they have dismissed such cases as irrelevant because they haven't seen how to integrate them with things they do understand. I actually think the answer is a combination of these two, a point I will elaborate on in a second post.

In the meantime, when the usage-based theorists have something to say about the range of grammatical phenomena, and the deep similarities found among widely diverse languages that animate discussion in generative syntax, we will be ready to engage. Until then, my friends and I will continue our long slow march to scientific obsolescence.

Thursday, September 8, 2016

Another constituency heard from

For those misguided souls like me who can't get enough of the awful coverage of the Wolfe book, here is a breath of fresh air.  Steven Poole (I think that this is him) has a short reasonable review in the Guardian (here) (Thx to David Hall for bringing it to my attention). The tone strikes me as just right, as is the main judgment. He notes that Wolfe's book is "a sad example of the interface of literary celebrity with publishing." Yup. Poole, however, is more optimistic than I am about the current state of publishing and critical thought among the high brow, for he continues by saying: "An author less famous and bankable than Wolfe would surely have been saved from such embarrassment by more critical editorial attention." If only. We have a natural experiment indicating otherwise, Knight's book which is by an unknown, has been published by Yale and is also deeply ignorant of the linguistic subject matter. It seems that anything goes when going after Chomsky. Why? Dunno. Maybe the threat his ideas presents are too great to take what he actually says as the subject matter for criticism. Maybe the aim is to make sure that they don't get discussed at all for fear that they will prove too persuasive. And what better way to prevent this than attacking straw men, and poorly fabricated ones at that.

Wednesday, September 7, 2016

Some filler

The papers are having a field day with Chomsky, getting hits by insisting that he is wrong about this and about that. The coverage is evidence that he casts a very long intellectual shadow and that his ideas are very attractive. You don't spend pages dumping on a nobody.  So, in one sense, all of the coverage is flattering. It is also deeply ignorant. I have spent some time rehearsing how Wolfe's views, based on Everett's misguided reasoning about Piraha and UG, is an intellectual (and moral) scandal (here). I have also examined in some detail how the high brow press spreads ignorance (see here for a discussion of Tom Bartlett and his agnotological efforts in the Chronicle).

I have also linked to Coyne's more cogent review of Wolfe's book in the Washington Post (here). I add another for your interest. It reviews Caitlin Flanagan's review in the NYT of the Wolfe book (here). It is by Nathan Robinson in Current Affairs (here). Robinson's review of Flanagan says does not deal much with linguistics, but then neither does Flanagan's review. In fact, Flanagan notes, quite rightly, that Wolfe's discussion of Chomsky's linguistics is entirely parasitic on the New Yorker piece on Everett in 2007. The only problem with Flanagan's discussion show it flags that the New Yorker piece. The review says that the New Yorker piece by John Colapinto "sums up the relevant Chomsky theories more clearly than anything in "the Kingdom of Speech." There is a sense in which this is correct, but a more important sense in which it is not. I take the phrasing to implicate that the New Yorker  piece does a credible job of summing up Chomsky's views. This is false. The New Yorker article not only fails to identify the relevant issues, it also manages to obfuscate them by missing the fact that the Everett claims about Piraha are logically irrelevant to Chomsky's claims about UG and FL. Sadly, no doubt given the prestige of the New Yorker as a high brow thinking person's mag, Colapinto's framing of the issues has surfaced repeatedly in all articles that have made the "debate" a central focus of Chomsky coverage. And given that Colapinto's framing distorts the relevant lay of the intellectual land so completely, it has been a baleful influence on all further popular writing on the matter.

At any rate, this is old news. I bring the Robinson reply to Falangan's review to your attention because it notes a vey critical aspect of all of this Chomsky bashing. It is based on carefully not reading what Chomsky has actually written. Robinson focuses on how most everything Flanagan says about Chomsky's "politics" in her price is actually the opposite of what he has written. As you know, this is also true of his linguistic views. It seems that critics consider Chomsky's views on matters linguistic and political so dangerous that actually presenting them accurately is potentially toxic. Better to attack views he does not have than to attack views that he holds. Of course, this is shoddy and dishonest, but hey, Chomsky's views must be discredited, or at least he must be.

Flanagan's NYT piece does serve an important function, something that Bartlett's piece mentioned but only sotto voce: the aim of the Wolfe book/Harper's article was to "fillet" the "New Left figure" Noam Chomsky. Note, not to "fillet" the ideas, nor the evidence, nor the argumentation, nor anything else, but to "fillet" the person. This is exactly right. And this is Flanagan's aim as well. Robinson demonstrates the intellectual shoddiness of Flanagan's asides. Like Wolfe, she too has not read those pieces that she feels comfortable dismissing. Like Wolfe, her misunderstandings are not profound, but rest on a simple refusal to read what Chomsky has written, as Robinson demonstrates. talk abut dumb!


Saturday, September 3, 2016

Brains and syntax: part 2

This is the second part of William Matchin's paper. Thanks again to William for putting this down on paper and provoking discussion.


One (very rough) sketch of a possible linking theory between a minimalist grammar and online sentence processing

I am going to try and sketch out what I think is a somewhat reasonable picture of the language faculty given the insights of syntactic theory, psycholinguistics, and cognitive neuroscience. My sketch here takes some inspiration from TAG-based psycholinguistic research (e.g., Demberg & Keller, 2008) and the TAG-based syntactic theory developed by Frank (2002) (thanks to Nick Huang for drawing this work to my attention).

Figure from Frank (2002). The dissociation between the inputs and operations of basic structure building and online processing/manipulation of treelets is clearly exemplified in the grammatical framework of Frank (2002).

The essential qualities of this picture of the language faculty are as follows. Minimalism is essentially a theory of the objects of language, the syntactic representations that people have. These objects are TAG treelets. TAG is a theory of what people do with these objects during sentence processing. TAG-type operations (e.g., unification, substitution, adjunction, verification) may be somehow identifiable with memory retrieval operations, opening up a potentially general cognitive basis for the online processing component of the language faculty, leaving the language-specific component to Merge. This proposal severs any inherent connection between Merge and online processing – although nothing in the proposal precludes the online implementation of Merge during sentence processing, much of sentence processing might proceed without having to implement Merge, but rather TAG operations operating over stored treelets.

I start with what I take to be the essential components of a Minimalist grammar – the lexicon and the computational system (i.e., Merge). Things work essentially as a Minimalist grammar says – you have some lexical atoms, Merge combines these elements (bottom-up) to build structures that are interpreted by the semantic and phonological systems, and there are some principles – some of them part of cognitive endowment, some of them “third factors” or general laws of nature or computation – that constrain the system (Chomsky, 1995; 2005).

The key difference that I propose is that complex derived structures can be stored in long-term memory. Currently, Minimalism states that the core feature of language, recursion, is the ability to treat derived objects as atoms. In other words, structures are treated as words, and as such are equally good inputs to Merge. However, the theory attributes the property of long-term storage only to atoms, and denies long-term storage to structures. Why not make structures fully equivalent to the atoms in their properties, including both Merge-ability AND long-term store-ability?

These stored structures or treelets can either be fully-elaborated structures with the leaves attached, or they might be more abstract nodes, allowing different lexical items to be inserted. It seems important from the psycholinguistic literature to have abstract structural nodes (e.g. NP, VP), so this theory would have to provide some means of taking a complex structure created by Merge and modifying it appropriately to eliminate the leaves (and perhaps many of the structural nodes) of the structure through some kind of deletion operation.

Treelets are the point of interaction between the syntactic system (essentially a Minimalist grammar) and the memory system. It may be the top-down activation of memory retrieval operations that “save” structures as treelets. Memory operations do much of the work of sentence processing – retrieving structures and unifying/substituting them appropriately to efficiently parse sentences (see Demberg & Keller, 2008 for an illustration). Much of language acquisition amounts to refining the attention/retrieval operations as well as the set of treelets and the prominence/availability of such treelets) that the person has available to them.

I think that there are good reasons to think that the retrieval mechanisms and the stored structures/lexical items live in language cortex. Namely, retrieval operations live in the pars triangularis of Broca’s area and stored structures/lexical items live in posterior temporal lobe (somewhere around the superior temporal sulcus/middle temporal gyrus).

This approach pretty much combines the Minimalist generative grammar and the lexicalist/TAG approaches. Note also that retrieving a stored treelet includes the fact that the treelet was created through applications of Merge. So when you look at structure that is finally said by a person, it is both true that the syntactic derivation of this structure is generated bottom-up in accordance with the operations and principles of a minimalist grammar, AND that the person used the thing by retrieving a stored treelet. We can (hopefully) preserve both insights – bottom-up derivation with stored treelets that can be targeted by working memory operations.

One remaining issue is how treelets are combined and lexical items inserted into them – this could be a substitution or unification operation from TAG, but Merge itself might also work for some cases (suggesting some role for Merge in actual online processing).

I think this proposal starts to provide potential insights into language acquisition. Say you’re a person walking around with this kind of system – you’ll want to start directing your attentional/working memory system to all these objects being generated by Merge and creating thoughts. You’ll also (implicitly) realize that other people are saying stuff that connects to your own system of thought, and you’ll start to align your set of stored structures and retrieval operations to match the patterns of what you’re seeing in the external world. This process is language acquisition, and it creates a convergence on the set of features, stored structures, and retrieval operations that are used within a language.


This addresses some of the central questions I posited earlier:

When processing a sentence, do I expect Merge to be active? Or not?

- Not necessarily, maybe minimally or not at all for most sentences.

What happens when people process things less than full sentences (like a little NP – “the dog”)? What is our theory of such situations?

- A little treelet corresponding to that sub-sentence structure is retrieved and interpreted.

Do derivations really proceed from the bottom up, or can they satisfactorily be switched to go top-down/left-right using something like Merge right (Phillips 1996)?

- Syntactic derivations are bottom-up in terms of Merge, but sentence processing occurs left-to-right roughly along the lines of TAG-based parsing frameworks (Demberg & Keller, 2008).

What happens mechanistically when people have to revise structure (e.g., after garden-pathing)?

- De-activate the current structure, retrieve new treelets/lexical items that fit better with what was presented. Lots of activity associated with processing lexical items/structures and memory retrievals, but there may not be an actual activation/implementation of Merge.

Are there only minimal lexical elements and Merge? Or are there stored complex objects, like “treelets”, constructions or phrase-structure rules?

- Yes, there are treelets, but we have an explanation for why there are treelets – they were created through applications of Merge at some point in the person’s life, but not necessarily online during sentence processing.

How does the syntactic system interact with working memory, a system that is critical for online sentence processing?

- The point of interaction between syntax and memory is the treelet. Somehow certain features encoded on treelets have to be available to the memory system.


Now that I have these answers, I can proceed to do my neuroimaging and neuropsychology experiments with testable predictions regarding how language is effected in the brain:

What’s the function of Broca’s area?

- Retrieval operations that are specialized to operate over syntactic representations.
- Which is why when you destroy Broca’s area you are still left with a bunch of treelets that can be activated in comprehension/production that you can use pretty effectively, although you have less strategic control over them.
- We expect patients with damage to Broca’s area to be able to basically comprehend sentences, but really have trouble in cases requiring recovery/revision, long-distance dependencies, prediction, and perhaps second language acquisition

What’s the function of posterior temporal areas?

- Lexical storage, including treelets.
- We expect activation for basic sentence processing, more activation for ambiguity/garden-path sentences when more structural templates are activated.
- We expect patients with damage to posterior temporal damage to have some real problems with sentence comprehension/production).

Where are fundamental structure building operations in the brain, e.g. Merge?

- Merge is a subtle neurobiological property of some kind.
- It might be in the connections between cortical areas, perhaps involving subcortical structures, or some property of individual neurons, but regardless, there isn’t a “syntax area” to be found.

What are the ramifications of this proposal for the standard contrast of sentences > lists that is commonly used to probe sentence processing in the brain?

- This contrast will highlight all sorts of things, likely including the activation of treelets, memory retrieval operations, semantic processing, but it might not be expected to drive activation for basic syntactic operations, i.e. Merge


Here I have tried to preserve Merge as the defining and simple feature of language – it’s the thing that allows people to grow structures. It also clearly separates Merge from the issue of “what happens during sentence processing”, and really highlights the core of language as something not directly tied to communication. Essentially, the theory of syntax becomes the theory of structures and dependencies, not producing and understanding sentences. On this conception of language, there is this Merge machinery creating structures, perhaps new in evolution that can be harnessed by an (evolutionarily older) attentional/memory system for the purposes of producing and comprehending sentences through storing treelets in long term memory. Merge is clearly separate from this communication/memory system, and an engine of thought. Learning a language then becomes a matter of refining the retrieval operations and what kinds of stored treelets you have that are optimized for communicating with others over time.

If this is a reasonable picture of the language faculty, thinking along these lines might start to help resolve some conundrums in the traditional domain of syntax. For example, there is often the intuition that syntactic islands are somehow related to processing difficulty (Kluender & Kutas 1993; Berwick & Weinberg, 1984), but there is good evidence that islands cannot be reduced to online processing difficulty or memory resource demands (Phillips, 2006; Sprouse et al., 2012). One approach might be to attribute islands to a processing constraint that somehow becomes grammaticalized (Berwick & Weinberg, 1984). The present framework provides a way for thinking about this issue, because the interaction between syntax and the online processing/memory system is specified. I have some more specific thoughts on this issue that might take the form of a future post.

At any rate, I would love any feedback on this type of proposal. Do we think this is a sensible idea of what the language faculty looks like? What are some serious objections to this kind of proposal? If this is on the right track, then I think we can start to make some more serious hypotheses about how language is implemented in the human brain beyond Broca’s area = Merge.

Many thanks to Nick Huang (particularly for pointing out relevant pieces of literature), Marta Ruda, Shota Momma, Gesoel Mendes, and of course Norbert Hornstein for reading this and giving me their thoughts. Thanks to Ellen Lau, Alexander Williams, Colin Phillips and Jeff Lidz for helpful discussion on these topics. Any failings are mine, not theirs.

References

Berwick, R. C., & Weinberg, A. S. (1983). The role of grammars in models of language use. Cognition, 13(1), 1-61.

Berwick, R., and Weinberg, A.S. (1984). The grammatical basis of linguistic performance. Cambridge, MA: MIT Press.

Bresnan, J. (2001). Lexical-Functional Syntax Blackwell.

Chomsky, N. (2005). Three factors in language design. Linguistic inquiry, 36(1), 1-22.

Chomsky, N. (1965). Aspects of the Theory of Syntax. MIT press.

Culicover, P. W., & Jackendoff, R. (2005). Simpler syntax. Oxford University Press on Demand.

Culicover, P. W., & Jackendoff, R. (2006). The simpler syntax hypothesis. Trends in cognitive sciences, 10(9), 413-418.

Demberg, V., & Keller, F. (2008, June). A psycholinguistically motivated version of TAG. In Proceedings of the 9th International Workshop on Tree Adjoining Grammars and Related Formalisms. Tübingen (pp. 25-32).

Embick, D., Marantz, A., Miyashita, Y., O'Neil, W., & Sakai, K. L. (2000). A syntactic specialization for Broca's area. Proceedings of the National Academy of Sciences, 97(11), 6150-6154.

Fedorenko, E., Behr, M. K., & Kanwisher, N. (2011). Functional specificity for high-level linguistic processing in the human brain. Proceedings of the National Academy of Sciences, 108(39), 16428-16433.

Fodor, J., Bever, A., & Garrett, T. G. (1974). The psychology of language: An introduction to psycholinguistics and generative grammar.

Frank, R. 2002. Phrase Structure Composition and Syntactic Dependencies. Cambridge, Mass: MIT Press.

Grodzinsky, Y. (2000). The neurology of syntax: Language use without Broca's area. Behavioral and brain sciences, 23(01), 1-21.

Grodzinsky, Y., & Friederici, A. D. (2006). Neuroimaging of syntax and syntactic processing. Current opinion in neurobiology, 16(2), 240-246.

Grodzinsky, Y. (2006). A blueprint for a brain map of syntax. Broca’s region, 83-107.

Jackendoff, R. (2003). Précis of foundations of language: brain, meaning, grammar, evolution. Behavioral and Brain Sciences, 26(06), 651-665.

Joshi, A. K., & Schabes, Y. (1997). Tree-adjoining grammars. In Handbook of formal languages (pp. 69-123). Springer Berlin Heidelberg.

Kluender, R., & Kutas, M. (1993). Subjacency as a processing phenomenon. Language and cognitive processes, 8(4), 573-633.

Lewis, S., & Phillips, C. (2015). Aligning grammatical theories and language processing models. Journal of Psycholinguistic Research, 44(1), 27-46.

Lewis, R. L., & Vasishth, S. (2005). An activation‐based model of sentence processing as skilled memory retrieval. Cognitive science, 29(3), 375-419.

Lewis, R. L., Vasishth, S., & Van Dyke, J. A. (2006). Computational principles of working memory in sentence comprehension. Trends in cognitive sciences, 10(10), 447-454.

Linebarger, M. C., Schwartz, M. F., & Saffran, E. M. (1983). Sensitivity to grammatical structure in so-called agrammatic aphasics. Cognition, 13(3), 361-392.

Lukyanenko, C., Conroy, A., & Lidz, J. (2014). Is she patting Katie? Constraints on pronominal reference in 30-month-olds. Language Learning and Development, 10(4), 328-344.

Matchin, W., Sprouse, J., & Hickok, G. (2014). A structural distance effect for backward anaphora in Broca’s area: An fMRI study. Brain and language, 138, 1-11.

Miller, G. A., & Chomsky, N. (1963). Finitary models of language users.

Mohr, J. P., Pessin, M. S., Finkelstein, S., Funkenstein, H. H., Duncan, G. W., & Davis, K. R. (1978). Broca aphasia Pathologic and clinical. Neurology, 28(4), 311-311.

Momma, 2016 (doctoral dissertation, University of Maryland, department of Linguistics)

Musso, M., Moro, A., Glauche, V., Rijntjes, M., Reichenbach, J., Büchel, C., & Weiller, C. (2003). Broca's area and the language instinct. Nature neuroscience, 6(7), 774-781.

Omaki, A., Lau, E. F., Davidson White, I., Dakan, M. L., Apple, A., & Phillips, C. (2015). Hyper-active gap filling. Frontiers in psychology, 6, 384.

Pallier, C., Devauchelle, A. D., & Dehaene, S. (2011). Cortical representation of the constituent structure of sentences. Proceedings of the National Academy of Sciences, 108(6), 2522-2527.

Phillips, C. (1996). Order and structure (Doctoral dissertation, Massachusetts Institute of Technology).

Phillips, C. (2006). The real-time status of island phenomena. Language, 795-823.

Rogalsky, C., & Hickok, G. (2011). The role of Broca's area in sentence comprehension. Journal of Cognitive Neuroscience, 23(7), 1664-1680.

Santi, A., & Grodzinsky, Y. (2012). Broca's area and sentence comprehension: A relationship parasitic on dependency, displacement or predictability?. Neuropsychologia, 50(5), 821-832.

Santi, A., Friederici, A. D., Makuuchi, M., & Grodzinsky, Y. (2015). An fMRI Study Dissociating Distance Measures Computed by Broca’s Area in Movement Processing: Clause boundary vs Identity. Frontiers in psychology, 6, 654.

Sprouse, J. (2015). Three open questions in experimental syntax. Linguistics Vanguard, 1(1), 89-100.

Sprouse, J., Wagers, M., & Phillips, C. (2012). A test of the relation between working-memory capacity and syntactic island effects. Language, 88(1), 82-123.

Stowe, L. A., Haverkort, M., & Zwarts, F. (2005). Rethinking the neurological basis of language. Lingua, 115(7), 997-1042.

Stowe, L. A. (1986). Parsing WH-constructions: Evidence for on-line gap location. Language and cognitive processes, 1(3), 227-245.

Vosse, T., & Kempen, G. (2000). Syntactic structure assembly in human parsing: a computational model based on competitive inhibition and a lexicalist grammar. Cognition, 75(2), 105-143.

Wilson, S. M., & Saygın, A. P. (2004). Grammaticality judgment in aphasia: Deficits are not specific to syntactic structures, aphasic syndromes, or lesion sites. Journal of Cognitive Neuroscience, 16(2), 238-252.

Zaccarella, E., & Friederici, A. D. (2015). Merge in the human brain: A sub-region based functional investigation in the left pars opercularis. Frontiers in psychology, 6.