Comments

Showing posts with label AGREE. Show all posts
Showing posts with label AGREE. Show all posts

Thursday, February 4, 2016

David Adger; Baggett Lecture 3

Well, he did it! Three excellent and provocative talks on syntactic theory. Here is the third set of slides. In this talk, David went after sidewards movement (SWM), a favorite idea of mine and, though you might not believe this, I sat quite demurely through the whole talk and basically agreed with much of the what David had to say. I, not surprisingly, did not buy the conclusion, but I did buy the way that he set up the problem and the way that he approached a solution given his empirical judgments about the empirical viability of SWM. How so?

David makes two important points (i.e. points that I completely agree with).

First, that contrary to what is sometimes said, the current definition(s) of Merge does not by itself rule out SWM as an instance of Merge. It is sometimes claimed that whereas Merge when applied to E and I instances is a binary operation, when extended to SWM becomes a 3-place operation. What is correct is that one can define SWM to be 3-place and thereby invidiously distinguish it from the other applications, BUT, this is not a definition forced by any notions of conceptual simplicity or computational elegance. It is just something one can do if one wants to rule out SWM.

Moreover, as David seemed to concede, there is no really non ad-hoc way of ruling SWM out by simplifying the definition of Merge. All require further (ahem) refinements (what I would dub, extrinsic machinery designed to rule out a perfectly well defined option). As he noted in the talk, and I have been at pains to emphasize in conversation over the years, SWM is what you get when you leave the simple definition alone. To repeat. It is possible to merge two unconnected expressions together (E-merge) and to merge a subpart of one expression to that expression (I-merge). SO why is it not also possible to merge the subpart of one expression to another that it is not a subpart of (SWM). In other words, you can "look inside" a constituent and you can have multiple constituents in a "workspace" and this is all you need to allow SWM unless you make things more complicated. And that's why I have always thought that SWM is a natural consequence of a very simple definition of Merge and that preventing it requires either complicating the definition or arguing that more goes into Merge than the simple operation.

Moreover, how to complicate matters is not a mystery. The idea that I-merge requires AGREE (i.e. move = Agree + EPP) suffices to block SWM as ex-ante the target does not c-command the mover. Needless to say, someones modus ponens can be someone else's modus tolens and one might conclude from this that AGREE is a suspect operation (e.g. someone like moi) that should be forcefully thrown out of our minimalist Eden. But, if SWM proves to be empirically unpalatable, well this is one way to get rid of it.

Side note: of course if E-merge and I-merge are actually the very same operation then why I-merge needs to be licensed by AGREE but E-merge does not (indeed cannot) have to be becomes a bit of a conceptual mystery. And please don't tell me about having to identify the inputs to Merge as if finding an expression inside a constituent is particularly computationally demanding (indeed, more demanding than finding an element in the lexicon or the numeration).

Second, David thinks that it is important to start thinking about how operations like Merge are computationally realized algorithmically. In fact, his talk presents an algorithm that makes SWM unavailable. In other words, Merge the rule allows it but the computational implementation of Merge inside an architecture with certain kinds of memory restrictions prevents it. So why not SWM on this view? It's the structure of linguistic memory stupid!

I liked the ambitions behind this a lot. I've argued before that this line of thinking is what MP should be endorsing (e.g. see here and here and search from SMT on the site for others). It seems to me that David is on board with this now. In particular, that the SMT should concern itself not merely with restrictions imposed on FL by the interfaces AP and CI but also by the kinds of memory structures that we think are necessary to use Gs generated by FL. We know a little about these things now and it is worth speculating as to what this would mean for FL. David's talk can be viewed as an exercise in this kind of thinking. Great.

So, I loved the lecture, but did not buy the conclusion. Let me note why very briefly.

One thing to look for in a new research program are things that are different or novel from the perspective of older research programs. SWM should it exist is a novel kind of operation that we would not have though reasonable within GB, say. However, if the above is right, then in an MP Merge-centric context this is a kind of operation we might expect to find. So, if we do find it, it constitutes an empirical argument in favor of the new way of looking at things. Thus, if SWM, then it is very interesting.  Of course, this does not mean that it is right. It probably isn't. But it is very interesting and we should not try to get rid of it because it looks novel. Just the opposite. We should look for cases and see how they fare empirically. IMO, SWM analyses have been pretty insightful (e.g. Nunes on parasitic gaps, Uriagereka and Bobaljik and Brown on head movement, moi on adjunct control, and some new stuff on double object constructions whose authors must remain nameless for now). This stuff might all be wrong, but I find the analyses very interesting.

So, thx to David for 3 great lectures. High theory indeed! And also lots of fun. As I noted before, when the lectures become available on video I will link to them.

Tuesday, April 28, 2015

A minimalist evolang story

MIT News (here) mentions a paper that recently appeared in Frontiers in Psychology (here) by Vitor Norbrega and Shigeru Miyagawa (N&M). The paper is an Evolang effort that argues for a rapid (rather than a gradual) emergence of FL. The blessed event was “triggered” by the emergence of Merge which allowed for the “integration” of two “pre-adapted systems,” one relating to outward expression (think AP) and one related to referential meaning (think CI). N&M calls the first the E-system and the second the L-system. The main point of the paper is that the L-system does not correspond to anything like a word. Why? Because words found in Gs are themselves hierarchically structured objects, with structures very like the kind we find in phrases (a DMish perspective). The paper is interesting and worth looking at, though I have more than a few quibbles with some of the central claims. Here are some comments.

N&M has two aims: first to rebut gradualist claims concerning the evolution of FL. The second is to provide a story for the rapid emergence of the faculty. I personally found the criticisms more compelling than the positive proposal. Here’s why.

The idea that FL emerged gradually generally rests on the idea that FL builds on more primitive systems that went from 1-word to 2-word to arbitrarily large n-word sequences.  My problem with these kinds of stories has always been how we get from 2 to arbitrarily large n. As Chomsky has noted, “go on indefinitely” does not obviously arise from “go to some fixed n.” The recursive trick that Merge embodies does not conceptually require priming by finite instances to get it going. Why? Because there is no valid inference from “I can do X once, twice” to “I can do X indefinitely many times.” True, to get to ‘indefinitely many X’ might casually (if not conceptually) require seguing via finite instances of X, but if it does, nobody has explained how it does.[1] Brute facts causing other brute facts does not an explanation make.

Let me put this another way: Perhaps as a matter of historical fact our ancestors did go through a protolanguage to get to FL. However, it has never been explained how going through such a stage was/is required to get to the recursive FL of the kind we have. The gradualist idea seems to be that first we tried 1-word sequences then 2-word and that this prompted the idea to go to 3, 4, n-word sequences for arbitrary n. How exactly this is supposed to have happened absent already having the idea that “going on indefinitely” was ok has never been explained (at least to me). As this is taken to be a defining characteristic of FL, failing to show the link between the finite stages and the unbounded one (a link that I believe is conceptually impossible to show, btw) leaves the causal relevance of the earlier finite stages (should they even exist) entirely opaque (if not worse).  So, the argument that recursion “gradually” emerged is not merely wrong, IMO, it is barely coherent, at least if one’s interest is in explaining how unbounded hierarchical recursion arose in the species.[2]

N&M hints at a second account that, IMO, is not as conceptually handicapped as the one above. Here it is: One might imagine a system in place in our ancestors capable of generating arbitrarily big “flat” structures. Such structures would be different from our FL in not being hierarchical, and the same in being unbounded. These procedures, then, could generate arbitrarily “long” structures (i.e. the flat structures could be indefinitely long (think beads on a string) but have 0-depth).  Now we can ask a question: how can one get from the generative procedures that deliver arbitrarily long strings to our generative procedures which deliver structures that are both long and deep? I confess to having been very attracted to this conception of Darwin’s Problem (DP). DP so understood asks for the secret sauce required to go from “flat” n-membered sets (or sequences for arbitrary n) to the kind of arbitrarily deeply hierarchically structured sets (or graphs or whatever) we find in Gs produced by FL. I have a dog in this fight (see here), though I am not that wedded to the answer I gave (in terms of labeling being the novelty that precipitated change). This version of the problem finesses the question of where recursion came from (after all, it assumes that we have a procedure to generate arbitrarily long flat structures) and substitutes the question where did hierarchical recursion come from. At any rate, the two strike me as different, the second not suffering from the conceptual hurdle besetting the first.

N&M provides more detailed arguments against several current proposals for a gradualist conception for the evolution of FL. Many of these seem to take words as fossils of the earlier evolutionary stages. N&M argues that words cannot be the missing link that gradualists have hoped for. The discussion is squarely based on Distributed Morphology reasoning and observations. I found the points N&M makes very much to the point. However, given the technical requirements needed to follow the details, I fear that tyros (i.e. the natural readership of Frontiers) will remain unconvinced. This said, the points seem dead on target.

This brings us to the second aim of the paper, and here I confess to having a hard time following the logic. The idea seems to be that Merge when added to the E systems we find in bird song and the L system we find in vervets gets us the kinds of generative systems we find in G products of FL  This is a version of the classical Minimalist answer to DP favored by Chomsky. I say “sort of” as Chomsky, at least lately, has been making a big deal of the claim that the mapping to E systems is a late accretion and the real action is in the mapping to thought. I am not sure that N&M disagrees with this (the paper doesn’t really discuss this point) as I am not sure how the L-system and Chomsky’s CI interface relate to one another. The L-system seems closer to concepts than full-blown propositional representations, but I could be wrong here.  At any rate, this seems to be the N&M view.

Here’s my problem; in fact a few. First, this seems to ignore the various observations that whatever our L-atoms are they seem different in kind from what we find in animal communication systems. The fact seems to be that vervet calls are far more “referential” than human “words” are. Ours are pretty loosely tied to whatever humans may use words to refer to. Chomsky has discussed these differences at length (see here for a recent critique of “referentialism”) and if he is in any way correct it suggests that vervet calls are not a very good proxy for what our linguistic atoms do as the two have very different properties. N&M might agree with this, distinguishing roots from words and saying that our words have the Chomsky properties but our concepts are vervetish. But how turning roots into words manages this remains, so far as I can see, a mystery. Chomsky notes that the question of where the properties of our lexical items comes from is at present completely mysterious. But the bottom line, as Chomsky sees it (and I agree with him here), is that “[t]he minimal meaning-bearing elements of human languages – word-like, but not words -- are radically different from anything known in animal communication systems.” And if this is right, then it is not clear to me that Merge alone is sufficient to explain what our language manages to do, at least on the lexical side. There is something “special” about lexicalization that we really don’t yet understand and it does not seem to be reducible to Merge and it does not seem to really resemble the kinds of animal calls that N&M invokes. In sum, if Merge is the secret sauce, then it did more than link to a pre-existing L-system of the kind we find in vervet calls. It radically changed their basic character. How Merge might have done this is a mystery (at least to me (and, I believe, Chomsky)).

Again, N&M might agree, for the story it tells does not rely exclusively on Merge to bridge the gap. The other ingredient involves checking grammatical features. By “grammatical” I mean that these features are not reducible to the features of the E or L systems. Merge’s main grammatical contribution is to allow these grammatical features to talk to one another (to allow valuation to apply). As roots don’t have such features, merging roots would not deliver the kinds of structures that our Gs do as roots do not have the wherewithal to deliver “combinatorial systems.” So it seems that in addition to Merge, we need grammatical features to deliver what we have.

The obvious question is where these syntactic features come from?  More pointedly, Merge for N&M seems to be combinatorically idle absent these features. So Merge as such is not sufficient to explain Gish generative procedures. Thus, the real secret sauce is not Merge but these features and the valuation procedures that they underwrite. If this is correct, the deep Evolang question concerns the genesis of these features, not the operation instructing how to put grammatical objects together given their feature structures. Or, put another way: once you have the features how to put them together seems pretty straightforward: put them together as the features instruct (think combinatorial grammar here or type theory). Darwin’s Problem on this conception reduces to explaining how these syntactic features got a mental toehold. Merge plays a secondary role, or so it seems to me.

To be honest, the above problem is a problem for every Minimalist story addressing DP. The Gs we are playing with in most contemporary work have two separate interacting components: (i) Merge serves to build hierarchy, (ii) AGREE in Probe-Goal configurations check/value features. AGREE operations, to my knowledge, are not generally reducible to Merge (in particular I-merge). Indeed trying to unify them, as in Chomsky’s early minimalist musings, has (IMO, sadly) fallen out of fashion.[3] But if they are not unified and most/many non-local dependencies are the province of AGREE rather than I-merge, then Merge alone is not sufficient to explain the emergence of Gs with the characteristic dependencies ours embody. We also need a story about the etiology of the long distance AGREE operation and a story about the genesis of the syntactic features they truck in.[4] To date, I know of no story addressing this, not even very speculative ones. We could really use some good ideas here (or, as in note 3, begin to rethink the centrality of Probe/Goal Agree).

I don’t want to come off sounding overly negative. N&M, unlike many evolangers know a lot about FL. Their critique of gradualist stories seems to be very well aimed. However, precisely because the authors know so much about FL while trying to give a responsible positive outline of an answer to DP the problem, the paper makes clear the outstanding problems that providing an adequate explanation sketch faces. For this alone, N&M is worth reading.

So what’s the takeaway message here? I think we know what a solution to DP in the domain of language should involve. It should provide an account of how the generative procedures responsible for the G properties we have discovered over the last 60 years arose in the species. The standard Minimalist answer has been to focus on Merge and argue that adding it the capacities of our non-linguistic ancestors suffices to give them our kinds of grammatical powers. Now, there is no doubting that Merge does work wonders. However, if current theoretical thinking is on the right track, then Merge alone is insufficient to account for the various non-local dependencies that we find in Gs. Thus, Merge alone does not deliver what we need to fully explain the origins of our FL (i.e. it leaves out a large variety of agreement phenomena).[5] In this sense, either we need some  ideas about where AGREE comes from, or we need some work showing how to accomodate the phenomena that AGREE does via I-merge. Either way, the story that ties the evolutionary origins of our FL to the emergence of a single novel Merge operation is, at best, incomplete.



[1] Here from Edward St Aubyn in At Last: the final Patrick Melrose Novel:
 “Ok, so who created infinite regress.” That’s the right question.
[2] No less a figure than Wittgenstein had a field day with this observation.  “And so on” is not a concept that finite sequences of anything embody.
[3] I may be one of the last thinking that moving to AGREE systems was a bad idea if one’s interest is in DP. I argue this here. I don't think I’ve convinced many of the virtues either of disagreement in general or dis-AGREE-ment in particular. So it goes.
[4] It is tempting to see Chomsky’s latest discussions of labeling as an attempt to resolve this problem. Agreement on this view is what is required to get interpretable objects at the CI interface. It is not the product of AGREE but of the labeling algorithm. Chomsky does not say this. But this is where he might be heading. It is an attempt to reduce “morphology” to Bare Output Conditions. I personally am not convinced by his detailed arguments, but if this is the intent, I am very sympathetic to the project.
[5] I am currently co-teaching intro to contemporary minimalism with Omer Preminger. He has inundated me with arguments (good ones) that something like AGREE does excellent work in accounting for huge swaths of intricate data. Thus, at the very least, it seems that the current consensus among minimalist syntacticians is that Merge is not the only basic syntactic operation and so an account that ties all of our grammatical prowess to Merge is either insufficient or the current consensus is wrong. If I were a betting person, I would put my money on the first disjunct. But…

Monday, July 14, 2014

A question about feature valuation

I've been working in a "whig history" (WH) of generative grammar. A WH is a kind of rational reconstruction which, if doable, serves to reconstruct the logical development of a field of inquiry. WHs, then, are not “real” histories. Rather, they present the past as “an inevitable progression towards ever greater…enlightenment.” Real history is filled with dead ends, lucky breaks, misunderstandings, confusions, petty rivalries, and more.  WHs are not. They focus on the “successful chain of theories and experiments that led to the present-day science, while ignoring failed theories and dead ends” (see here). The value of WHs is that they expose the cumulative nature of a given trajectory of inquiry. As one sign of a real science is that it has a cumulative structure and given that many think that the history of Generative Grammar (GG) fails to have a cumulative structure, many think that this tells against the GG enterprise. However, the "many" are wrong: GG has a perfectly respectable WH and both empirically and theoretically the development has been cumulative. In a word, we've made loads of progress.  But this is not the topic for this post. What is?

As I went about reconstructing the relation between current minimalist theory and earlier GB theory, I came to appreciate just how powerful the No Tampering Condition (NTC) really is (I know I know, I should have understood this before, but dim bulb that I am, I didn't). I understand the NTC as follows: the inputs to a given grammatical operation must be preserved in the outputs of that operation. In effect, the NTC is a conservation principle that says that structure can be created but not destroyed. Replacing an expression with a trace of that expression destroys (i.e. fails to preserve) the input structure in the output and so the GB conception of traces is theoretically inadmissible in a minimalist theory that assumes the NTC (which, let me remind you is a very nice computational principle and part of most (all?) current minimalist proposals).

The NTC has many other virtues as well. For example, it derives the fact that movement rules cannot "lower" and that movement (at least within a single rooted sub-"tree") is always to a "c-commanding" position. Those of you who have listened to any of the Chomsky lectures I posted earlier will understand why I have used scare quotes above. If you don't know why and don't want to listen to the lectures, as David Pesetsky. He can tell you.  

At any rate, the NTC also suffices to derive the GB Projection Principle and the MP Extension Condition. In addition, it suffices to eliminate trace theory as a theoretical option (viz. co-indexed empty categories that are residues of movement: [e]1). Why? because traces cannot exist in the input to the derivation and so they cannot exist in the output given the NTC. Thus, given the NTC, the only way to implement the Projection Principle is via the Copy Theory. This is all very satisfying theoretically for the usual minimalist reasons. However, it also raises a question in my mind, which I would like to ask here.

Why doesn't the NTC rule out feature valuation?  One of the current grammatical operations within MP grammars is AGREE. What it does is relate two expressions (heads actually) in a Probe/Goal configuration and the goal "values" the features of the probe.  Now, the way I've understood this is that the Probe is akin to a property, something like P(x) (maybe with a lambdaish binder, but who really cares) and the goal serves to turn that 'x' into some value, so turns P(x) into P(phi) for example (if you want, via something like lambda conversion, but again who really cares). At any rate, and here's my question: doesn't this violate the NTC? After all, the input to AGREE is P(x) and the output is, e.g. P(phi). Doesn't this violate a strict version of the NTC?

Note, interestingly, feature checking per se is consistent with the NTC, as no feature changing/valuing need go on to "check" if sets of features are "compatible."  However, if I understand the valuation idea, then it is thought to go beyond mere bookkeeping. It is intended to change the feature composition of a probe based on the feature composition of the goal.  Indeed, it is precisely for this reason that phases are required to strip off the valued yet uninterpretable features before Transfer. But if AGREE changes feature matrices then it seems incompatible with the NTC.

The same line of reasoning suggests that feature lowering is also incompatible with the NTC. To wit: if features really transfer from C to T or from v to V (either by being copied from the former to the latter or actually copied from the higher to the lower and deleted from the higher) then again the NTC in its strongest form seems to be violated. 

So, my question: are theories that adopt feature valuation and feature lowering inconsistent with the NTC or not? Note, we can massage the NTC so that it does not apply to such feature "checking" operations. But then we could massage the NTC so that it does not prohibit traces. We can, after all, do anything we wish. For example, current theory stipulates that pair merge, unlike set merge, is not subject to Extension, viz. the NTC (though I think that Chomsky is not happy with this given some oblique remarks he made in lecture 3). However, if  the NTC is strictly speaking incompatible with these two operations, then it is worth knowing, as it would seem to be theoretically very consequential. For example, a good chunk of phase theory, as currently understood, depends on these operations and would we discover that they are incompatible with the NTC then this might (IMO, likely does) have consequences for Darwin's Problem.

So, all you thoroughly modern minimalists out there: what say you?