Comments

Showing posts with label recursion. Show all posts
Showing posts with label recursion. Show all posts

Tuesday, January 31, 2017

A small addendum to the previous post on Syntactic Structures

Here’s a small addition to the previous post prompted by a discussion with Paul Pietroski. I am always going on about how focusing on recursion in general as the defining property of FL is misleading. The interesting feature of FL is not that it produces (or can produce) recursive Gs but that it produces the kinds of recursive Gs that it does. So the minimalist project is not to explain how recursion arises in humans but how a specific kind of recursion arises in the species. What kind? Well the kind we find in the Gs we find. What kind are these? Well not FSGs nor CFGs, or at least this is what Syntactic Structures (SS) argues.

Let me put this another way: GG has spent the last 60 years establishing that human Gs have a certain kind of recursive structure. In SS, it argued for a transformational grammar arguing that FSGs (which were recursive) were inherently too weak and that PSGs (also recursive) were inadequate empirically. Transformational Gs, SS argued, are the right fit.


So, when people claim that the minimalist problem is to explain the sources of recursion or observe that there may be/is recursion in other parts of cognition thereby claiming to “falsify” the project, it seems to me that they are barking up a red herring (I love the smell of mixed metaphors in the morning!). From where I sit, the problem is explaining how an FL that delivers TGish recursive G arose as this is the kind of FL that we have and the kinds of Gs that it delivers. SS makes clear that “in the earliest days of GG,” not all recursive Gs are created equal and that the FL and Gs of interest have specific properties. It’s the sources for this kind of recursion we want to explain. This is worth bearing in mind when issues of recursion (and its place in minimalist theory) make it to the spotlight.

Friday, January 20, 2017

Tragedy, farce, pathos

Dan Everett (DE) has written once again on his views about Piraha, recursion, and the implications for Universal Grammar (here). I was strongly tempted to avoid posting on it for it adds nothing new of substance to the discussion (and will almost certainly serve to keep the silliness alive), beyond a healthy dose of self-pity and self-aggrandizement. It makes the same mistakes, in almost the same way, and adds a few more irrelevancies to the mix. If history surfaces first as tragedy and the second time as farce (see here) then pseudo debates in their moth eaten n-th iteration are just pathetic. The Piraha “debate” has long since passed its sell-by date. As I’ve said all that I am about to say before, I would urge you not to expend time or energy reading this. But if you are the kind of person who slows down to rubberneck a wreck on the road and can’t help but find the ghoulish fascinating, this post is for you.

The DE piece makes several points.

First, that there is a debate. As you all know this is wrong. There can be no debate if the controversy hinges on an equivocation. And it does, for what the DE piece claims about the G of Piraha, even if completely accurate (which I doubt, but the facts are beyond my expertise) has no bearing on Chomsky’s proposal, viz. that recursion is the only distinctively linguistic feature of FL. This is a logical point, not an empirical one. More exactly, the controversy rests on an equivocation concerning the notion “universal.” The equivocation has been a consistent feature of DE’s discussions and this piece is no different. Let me once again explain the logic.

Chomsky’s proposal rests on a few observations. First, that humans display linguistic creativity. Second, that humans are only accidentally native speakers of their native languages.

The first observation is manifest in the fact that, for example, a native speaker of English, can effortlessly use and understand an unbounded number of linguistic expressions never before encountered. The second is manifest in the observation that a child deposited in any linguistic community will grow up to be a linguistically competent native speaker of that language with linguistic capacities indistinguishable from any of the other native speakers (e.g. wrt his/her linguistic creativity).

These two observations prompt some questions.

First, what underlying mental architecture is required to allow for the linguistic creativity we find in humans?  Answer 1 a mind that has recursive rules able to generate ever more sophisticated expressions from simple building blocks (aka, a G). Question 2: what kind of mental architecture must a such a G competent being have? Answer 2: a mind that can acquire recursive rules (i.e a G) from products of those rules (i.e. generated examples of the G). Why recursive rules? Because linguistic productivity just names the fact that human speakers are competent with respect to an unbounded number of different linguistic expressions.

Second, why assume that the capacity to acquire recursive Gs is a feature of human minds in general rather than simply a feature of those human minds that have actually acquired recursive Gs? Answer: Because any human can acquire any G that generates any language. So the capacity to acquire language in general requires the meta-capacity to acquire recursive rule systems (aka, Gs).  As this meta-capacity seems to be restricted to humans (i.e. so far as we know only humans display the kind of recursive capacity manifested in linguistic creativity) and as this capacity is most clearly manifest in language then Chomsky’s conjecture is that if there is anything linguistically specific about the human capacity to acquire language the linguistic specificity resides in this recursive meta-capacity.[1] Or to put this another way: there may be more to the human capacity to acquire language than the recursive meta-capacity but at least this meta capacity is part of the story.[2] Or, to put this another way, absent the human given (i.e. innate) meta-capacity to acquire (certain specifiable kinds of) recursive Gs, humans would not be able to acquire the kinds of Gs that we know that they in fact do acquire (e.g. Gs like those English, French, Spanish, Tagalog, Arabic, Inuit, Chinese … speakers have in fact acquired). Hence, humans must come equipped with this recursive meta-capacity as part of FL.

Ok, some observations: recursion in this story is principally a predicate of FL, the meta-capacity. The meta-capacity is to acquire recursive Gs (with specific properties that GG has been in the business of identifying for the last 50 years or so). The conjecture is that humans have this meta-capacity (aka FL) because they do in fact display linguistic creativity (and, as the DE paper concedes, native speakers of non-Piraha do regularly display linguistic creativity implicating the internalization of recursive language specific Gs) and because the linguistic creativity a native speaker of (e.g.) English displays could have been displayed by any person raised in an English linguistic milieu. In sum, FL is recursive in the sense that it has the capacity to acquire recursive Gs and speakers of any language have such FLs.

Observe that FL must have the capacity to acquire recursive Gs even if not all human Gs are recursive. FL must have this capacity because all agree that many/most (e.g.) non-Piraha Gs are recursive in the sense that Piraha is claimed not to be. So, the following two claims are consistent: (1) some languages have non-recursive Gs but (2) native speakers of those languages have recursive FLs. This DE piece (like all the other DE papers on this topic) fails, once again, to recognize this. A discontinuous quote (4):

 If there were a language that chose not to use recursion, it would at the very least be curious and at most would mean that Chomsky’s entire conception of language/grammar is wrong….

Chomsky made a clear claim –recursion is fundamental to having a language. And my paper did in fact present a counterexample. Recursion cannot be fundamental to language if there are languages without it, even just one language.

First an aside: I tend to agree that it would indeed be curious if we found a language with a non-recursive G given that virtually all of the Gs that have been studied are recursive. Thus finding one that is not would be odd for the same reason that finding a single counter example to any generalization is always curious (and which is why I tend not to believe DE’s claims and tend to find the critique by Nevins, Pesetsky and Rodrigues compelling).[3] But, and this is the main take home message, whether curious or not, it is at right angles to Chomsky’s claim concerning FL for the reasons outlined above. The capacity to acquire recursive Gs is not falsified by the acquisition of a non-recursive one. Thus, logically speaking, the observation that Piraha does not have embedded clauses (i.e. does not the display one of the standard diagnostics of a recursive G) does not imply that Piraha speakers do not have recursive FLs. Thus, DE’s claims are completely irrelevant to Chomsky’s even if correct. That point has been made repeatedly and, sadly, it has still not sunk in. I doubt that for some it ever will.

Ok, let’s now consider some other questions. Here’s one: is this linguistic meta-capacity permanent or evanescent? In other words, one can imagine that FL has the capacity to acquire recursive Gs but that once it has acquired a non-recursive G it can no longer acquire a recursive one. DE’s article suggests that this is so for Piraha speakers (p. 7). Again, I have no idea if this is indeed the case (if true it constitute evidence for a strong version of the Sapir-Whorf hypothesis) but this claim even if correct is at right angles to Chomsky’s claim about FL. Species specific dedicated capacities need not remain intact after use. It could be true that FL is only available for first language acquisition and this would mean that second languages are acquired in different ways (maybe by piggy backing on the first G acquired).[4] However so far as I know, neither Chomsky nor GG has ever committed hostages to this issue. Again, I am personally skeptical that having a Piraha G precludes you from the recursive parts of a Portuguese G, but I have nothing but prejudicial hunches to sustain the skepticism. At any rate, it doesn’t bear on Chomsky’s thesis concerning FL. The upshot: DE’s remarks once again are at right angles to Chomsky’s claims so interesting as the possibility it raises might be for interesting issues relating to second language acquisition, it is not relevant to Chomsky’s claims about the recursive nature of FL.

A third question: is the meta-capacity culturally relative? DE’s piece suggests that it is because the actual acquisition of recursive Gs might be subject to cultural influences. The point seems to be that if culture influences whether an acquired G is recursive or not implies that the meta-capacity is recursive or not as well. But this does not follow.  Let me explain.

All agree that the details of an actual G are influenced by all sorts of factors, including culture.[5] This must be so and has been insisted upon since the earliest days of GG. After all, the G one acquires is a function of FL and the PLD used to construct that G. But the PLD is itself a function of what is actually gets and there is no doubt that what utterances are performed is influenced by the culture of the utterers.[6] So, that culture has an effect on the shape of specific Gs is (or should be) uncontroversial. However, none of this implies that the meta-capacity to build recursive Gs is itself culturally dependent, nor does DE’s piece explain how it could be. In fact, it has always been unclear how external factors could affect this meta-capacity. You either have a recursive meta-capacity or you don’t. As Dawkins put it (see here for discussion and references):

… Just as you can’t have half a segment, there are no intermediates between a recursive and a non-recursive subroutine. Computer languages either allow recursion or they don’t. There’s no such thing as half-recursion. It’s an all or nothing software trick… (383)

Given this “all or nothing” quality, what would it mean to say that the capacity (i.e. the innately provided “computer language” of FL) was dependent on “culture.”? Of course, if what you mean is that the exercise of the capacity is culture dependent and what you mean by this is that it depends on the nature of the PLD (and other factors) that might themselves be influenced by “culture” then duh! But, if this is what DE’s piece intends, then once again it fails to make contact with Chomsky’s claim concerning the recursive nature of FL. The capacity is what it is though of course the exercise of the capacity to produce a G will be influenced by all sorts of factors, some of which we can call “culture.”[7]

Two more points and we are done.

First, there is a source for the confusion in DE’s papers (and it is the same one I have pointed to before). DE’s discussion treats all universals as if Greenbergian. Here’s a quote from the current piece that shows this (I leave it as an exercise to the reader to uncover the Greenbergian premise):

The real lesson is that if recursion is the narrow faculty of language, but doesn’t actually have to be manifested in a given language, then likely more languages than Piraha…could lack recursion. And by this reasoning we derive the astonishing claim that. Although, recursion would be the characteristic that makes human language possible, it need not actually be found in any given language. (8)

Note the premise: unless every G is recursive then recursion cannot be “that which makes human languages possible.” But this only makes sense if you understand things as Greenberg does. If you understand the claim as being about the capacity to acquire recursive Gs then none of this follows.

Nor are we led to absurdity. Let me froth here. Of course, nobody would think that we had a capacity for constructing recursive Gs unless we had reason to think that some Gs were so. But we have endless evidence that this is the case. So, given that there is at least one such G (indeed endlessly many), humans clearly must have the capacity to construct such Gs. So, though we might have had such a capacity and never exercised it (this is logically possible), we are not really in that part of the counterfactual space. All we need to get the argument going for a recursive meta-capacity is mastery of at least one recursive G and there is no dispute that there exists such a G and that humans have acquired it. Given this, the only coherent reason for thinking a counterexample (like Piraha) could be a problem is if one understood the claim to universality as implying that a universal property of FL (i.e. a feature of FL) must manifest itself in every G. And this is to understand ‘universal’ a la Greenberg and and not as Chomsky does. Thus we are back to original sin in DE’s oeuvre; the insistence on a Greenberg conception of universal.

Second, the piece makes another point. It suggests that DE’s dispute with Chomsky is actually over whether recursion is part of FL or part of cognition more generally. Here’s the quote (10):

…the question is not whether humans can think recursively. The question is whether this ability is linked specifically to language or instead to human cognitive accomplishments more generally…

If I understand this correctly, it is agreed that recursion is an innate part of human mental machinery. What’s at issue is whether there is anything linguistically proprietary about it. Thus, Chomsky could be right to think that human linguistic capacity manifests recursion but that this is not a specifically linguistic fact about us as we manifest recursion in our mental life quite generally.[8]

Maybe. But frankly it is hard to see how DE’s writings bear on these very recondite issues. Here’s what I mean: Human Gs are not merely recursive but exhibit a particular kind of recursion. Work in GG over the last 60 years has been in service of trying to specify what kind of recursive Gs humans entertain. Now, the claim here is that we find the kind of structure we find in human Gs in cognition more generally. This is empirically possible. Show me! Show me that other kinds of cognition have the same structures as those GGers have found occur in Gs.  Nothing in DE’s arguments about Piraha have any obvious bearing on this claim for there is no demonstration that other parts of cognition have anything like the recursive structure we find in human Gs.

But let’s say that we establish such a parallelism. There is still more to do. Here is a second question: is FL recursive because our mental life in general is or is our mental life in general recursive because we have FL.[9] This is the old species specificity question all over again. Chomsky’s claim is that if there is anything species special about human linguistic facility it rests in the kind of recursion we find in language. To rebut this species specificity requires showing that this kind of recursion is not the exclusive preserve of linguistically capable beings. But, once again, nothing in DE’s work addresses this question. No evidence is presented trying to establish the parallel between the kind of recursion we find in human Gs and any animal cognitive structures.

Suffice it to say that the kind of recursion we find in language is not cognitively ubiquitous (so far as we can tell) and that if it occurs in other parts of cognition it does not appear to be rampant in non-human animal cognition. And, for me at least, that is linguistically specific enough. Moreover, and this is the important point as regards DE’s claims, it is quite unclear how anything about Piraha will bear on this question. Whether or not Piraha has a recursive G will tell us nothing about whether other animals have recursive minds like ours.

Conclusion? The same as before. There is no there there. We find arguments based on equivocation and assertions without support. The whole discussion is irrelevant to Chomsky’s claims about the recursive structure of FL and whether that is the sole UGish feature of FL.[10]

That’s it. As you can see, I got carried away. I didn’t mean to write so much. Sorry. Last time? Let’s all hope so.


[1] Here you can whistle some appropriate Minimalist tune if you would like. I personally think that there is something linguistically specific about FL given that we are the only animals that appear to manifest anything like the recursive structures we find in language. But, this is an empirical question. See here for discussion.
[2] Chomsky’s minimalist conjecture is that this is the sole linguistically special capacity required.
[3] Indeed such odd counterexamples place a very strong burden of proof on the individual arguing for it. Sometimes this burden of proof can be met. But singular counterexamples that float in a sea of regularity are indeed curious and worthy of considerable skepticism. However, that’s not my point here. It is a different one: the Piraha facts whatever they turn out to be are irrelevant to the claim the FL has the capacity to acquire recursive Gs. As this is what Chomsky has been proposing. Thus, the facts regarding Piraha whatever they turn out to be are logically irrelevant to Chomsky’s proposal.
[4] This seems to be the way that Sakel conceives of the process (see here). Sakel is the person the DE piece cites as rebutting the idea that Piraha speakers with Portuguese as a second language behave. That speakers build their second G on the scaffolding provided by a first G is quite plausible a priori (though whether it is true is another matter entirely). And if this is so, then features of one’s first G should have significant impact on properties of one’s second G. Sakel, btw, is far less categorical in her views than what DE’s piece suggests. Last point: a nice “experiment” if this interests you is to see what happens if a speaker is acquiring Portuguese and Piraha simultaneously; both as first Gs. What should we expect? I dunno, but my hunch is that both would be acquired swimmingly.
[5] So, for example, dialects of English differ wrt the acceptability of Topicalization. My community used it freely and I find them great. My students at UMD were not that comfortable with this kind of displacement. I am willing to bet that Topicalization’s alias (i.e. Yiddish Movement) betrays a certain cultural influence.
[6] Again, see note 4 and Sakel’s useful discussion of the complexity of Portuguese input to the Piraha second language acquirer.
[7] BTW, so far as I can tell, invoking “culture” is nothing but a rhetorical flourish most of the time. It usually means nothing more than “not biology.” However, how culture affects matters and which bits do what is often (always?) left unsettled. It often seems to me that the word is brandished a bit like garlic against vampires, mainly there to ward off evil biological spirits.
[8] On this view, DE agrees that there is FLB but no FLN, i.e. a UGish part of FL.
[9] In Minimalist terms, is recursion a UGish part of FL or is there no UG at all in FL.
[10] There is also some truly silly stuff in which DE speculates as to why the push back against his views has been so vigorous. Curiously, DE does not countenance the possibility that it is because his arguments though severely wanting have been very widely covered. There is some dumb stuff on Chomsky’s politics, Wolfe junk, and general BS about how to do science. This is garbage and not worth your time, except for psycho-sociological speculation.

Wednesday, July 13, 2016

Linguistic creativity 1

Once again, this post got away from me, so I am dividing it into two parts.

As I mentioned in a recent previous post, I have just finished re-reading Language & Mind (L&M) and have been struck, once again, about how relevant much of the discussion is to current concerns. One topic, however, that does not get much play today, but is quite well developed in L&M is it’s discussion of Descartes’ very expansive conceptions of linguistic creativity and how it relates to the development of the generative program. The discussion is surprisingly complex and I would like to review its main themes here. This will reiterate some points made in earlier posts (here, here) but I hope it also deepens the discussion a bit.

Human linguistic creativity is front and center in L&M as it constitutes the central fact animating Chomsky’s proposal for Transformational Generative Grammar (TGG). The argument is that a TGG competence theory is a necessary part of any account of the obvious fact that humans regularly use language in novel ways. Here’s L&M (11-12):

…the normal use of language is innovative, in the sense of much of what we say in the course of normal use is entirely new, not a repetition of anything that we have heard before and not even similar in pattern - in any useful sense of the terms “similar” and “pattern” – to sentences or discourse that we have heard in the past. This is a truism, but an important one, often overlooked and not infrequently denied in the behaviorist period of linguistics…when it was almost universally claimed that a person’s knowledge of language is representable as a stored set of patterns, overlearned through constant repetition and detailed training, with innovation being at most a matter of “analogy.” The fact surely is, however, that the number of sentences in one’s native language that one will immediately understand with no feeling of difficulty or strangeness is astronomical; and that the number of patterns underlying our normal use of language and corresponding to meaningful and easily comprehensible sentences in our language is order of magnitudes greater than the number of seconds in a lifetime. It is in this sense the normal use of language is innovative.

There are several points worth highlighting in the above quote. First, note that normal use is “not even similar in pattern” to what we have heard before.[1] In other words, linguistic competence is not an instance of pattern matching or recognition in any interesting sense of “pattern” or “matching.”  Native speaker use extends both to novel sentences and to novel sentence patterns effortlessly. Why is this important?

IMO, one of the pitfalls of much work critical of GG is the assimilation of linguistic competence to a species of pattern matching.[2] The idea is that a set of templates (i.e. in L&M terms: “a stored set of patterns”) combined with a large vocabulary can easily generate a large set of possible sentences in the sense of templates saturated by lexical items that fit. [3] Note, that such templates can be hierarchically organized and so display one of the properties of natural language Gs (i.e. hierarchical structures).[4] Moreover, if the patterns are extractable from a subset of the relevant data then these patterns/templates can be used to project novel sentences. However, what the pattern matching conception of projection misses is that the patterns we find in Gs are not finite and the reason for this is that we can embed patterns within patterns within patterns within…you get the point. We can call the outputs of recursive rules “patterns” but this is misleading for once one sees that the patterns are endless, then Gs are not well conceived of as collections of patterns but collections of rules that generate patterns. And once one sees this, then the linguistic problem is (i) to describe these rules and their interactions and (ii) to further explain how these rules are acquired (i.e. not how the patterns are acquired).

The shift in perspective from patterns (and patternings in the data (see note 5)) to generative procedures and the (often very abstract) objects that they manipulate changes what the acquisition problem amounts to. One important implication of this shift of perspective is that scouring strings for patterns in the data (as many statistical learning systems like to do) is a waste of time because these systems are looking for the wrong things (at least in syntax).[5] They are looking for patterns whereas they should be looking for rules. As the output of the “learning” has to be systems of rules, not systems of patterns, and as rules are, at best, implicit in patterns, not explicitly manifest by them, theories that don’t focus on rules are going to be of little linguistic interest.[6]

Let me make this point another way: unboundedness implies novelty, but novelty can exist without unboundedness. The creativity issue relates to the accommodation of novel structures. This can occur even in small finite domains (e.g. loan words in phonology might be an example). Creativity implies projection/induction, which must specify a dimension of generalization along which inputs can be generalized so as to apply to instances beyond the input. This, btw, is universally acknowledged by anyone working on learning. Unboundedness makes projection a no-brainer. However, it also has a second important implication. It requires that the generalizations being made involve recursive rules. The unboundedness we find in syntax cannot be satisfied via pattern matching. It requires a specification of rules that can be repeatedly applied to create novel patterns. Thus, it is important to keep the issue of unboundedness separate from that of projection. What makes the unboundedness of syntax so important is that it requires that we move beyond the pattern-template-categorization conception of cognition.

Dare I add (more accurately, can I resist adding) that pattern matching is the flavor of choice for the Empricistically (E) inclined. Why? Well, as noted, everyone agrees that induction must allow generalization beyond the input data. Thus even Es endorse this for Es recognize that cognition involves projection beyond the input (i.e. “learning”). The question is the nature of this induction. Es like to think that learning is a function from input to patterns abstracted from the input, the input patterns being perceptually available in their patternings, albeit sometimes noisily.[7] In other words, learning amounts to abstracting a finite set of patterns from the perceptual input and then creating new instances of those patterns by subbing novel atoms (e.g. lexical items) into the abstracted patterns. E research programs amount to finding ways to induce/abstract patterns/templates from the perceptual patternings in the data. The various statistical techniques Es explore are in service of finding these patterns in the (standardly, very noisy) input. Unboundedness implies that this kind of induction is, at best, incomplete. Or, more accurately, the observation that the number of patterns is unbounded implies that learning must involve more than pattern detection/abstraction. In domains where the number of patterns is effectively infinite, learning[8] is a function from inputs to rules that generate patterns, not to patterns themselves. See link in note 6 for more discussion.

An aside: Most connectionist learners (and deep learners) are pattern matchers and, in light of the above, are simply “learning” the wrong things. No matter how many “patterns” the intermediate layers converge on from the (mega) data they are exposed to they will not settle on enough given that the number of patterns that human native speakers are competent in is effectively unbounded. Unless the intermediate layers acquire rules that can be recursively applied they have not acquired the right kinds of things and thus all of this modeling is irrelevant no matter how much of the data any given model covers.[9]

Another aside: this point was made explicitly in the quote above but to no avail. As L&M notes critically (11): “it was almost universally claimed that a person’s knowledge of language is representable as a stored set of patterns, overlearned through constant repetition and detailed training.” Add some statistical massaging and a few neural nets and things have not changed much. The name of the inductive game in the E world is to look for perceptual available patterns in the signal, abstract them and use them to accommodate novelty. The unboundedness of linguistic patterns that L&M highlights implies that this learning strategy won’t suffice the language case, and this is a very important observation.

Ok, back to L&M

Second, the quote above notes that there is no useful sense of “analogy” that can get one from the specific patterns one might abstract from the perceptual data to the unbounded number of patterns with which native speakers display competence. In other words, “analogy” is not the secret sauce that gets one from input to rules So, when you hear someone talk about analogical processes reach for your favorite anti-BS device. If “analogy” is offered as part of any explanation of an inferential capacity you can be absolutely sure that no account is actually being offered. Simply put, unless the dimensions of analogy are explicitly specified the story being proffered is nothing but wind (in both the Ecclesiastes and the scatological sense of the term).

Third, the kind of infinity human linguistic creativity displays has a special character: it is a discrete infinity. L&M observes that human language (unlike animal communication systems) does not consist of a “fixed, finite number of linguistic dimensions, each of which is associated with a particular nonlinguistic dimension in such a way that selection of a point along the linguistic dimension determines and signals selection of a point along the associated nonlinguistic dimension” (69). So, for example, higher pitch or chirp being associated with greater intention to aggressively defend territory or the way that “readings of a speedometer can be said, with an obvious idealization, to be infinite in variety” (12). 

L&M notes that these sorts of systems can be infinite, in the sense of containing “an indefinitely large range of potential signals.” However, in such cases the variation is “continuous” while human linguistic expression exploits “discrete” structures that can be used to “express indefinitely many new thoughts, intentions, feelings, and so on.”  ‘New thoughts’ in the previous quote clearly meaning new kinds of thoughts (e.g. the signals are not all how fast the car is moving). As L&M makes clear, the difference between these two kinds of systems is “not one of “more” or “less,” but rather of an entirely different principle of organization,” one that does not work by “selecting a point along some linguistic dimension that signals a corresponding point along an associate nonlinguistic dimension.” (69-70).

In sum, human linguistic creativity implicates something like a TGG that pairs discrete hierarchical structures relevant to meanings with discrete hierarchical structures relevant to sounds and does so recursively. Anything that doesn’t do at least this is going to be linguistically irrelevant as it ignores the observable truism that humans are, as matter of course, capable of using an unbounded number of linguistic expressions effortlessly.[10] Theories that fail to address this obvious fact are not wrong. They are irrelevant.

Is hierarchical recursion all that there is to linguistic creativity? No!! Chomsky makes a point of this in the preface to the enlarged edition of L&M. Linguistic creativity is NOT identical to the “recursive property in generative grammars” as interesting as such Gs evidently are (L&M: viii). To repeat, recursion is a necessary feature of any account aiming to account for linguistic creativity, BUT the Cartesian conception of linguistic creativity consists of far more than what even the most explanatorily adequate theory of grammar specifies.  What more?



[1] For an excellent discussion of this see Jackendoff’s very nice (though unfortunately (mis)named) Patterns in the mind (here).  It is a first rate debunking of the idea that linguistic minds are pattern matchers.
[2] This is not unique to the linguistic cognition. Lots of work in cog sci seems to identify higher cognition with categorization and pattern matching. One of the most important contributions of modern linguistics to cog sci has been to demonstrate that there is much more to cognition than this. In fact, the hard problems have less to do with pattern recognition than with pattern generation via rules of various sorts.  See notes 5 and 6 for more off handed remarks of deep interest.
[3] I suspect that some partisans of Construction Grammar fall victim to the same misapprehension.
[4] Many cog-neuro types confuse hierarchy with recursion. A recent prominent example is in Frankland and Greene’s work on theta roles. See here for some discussion. Suffice it to say, that one can have hierarchy without recursion, and recursion without hierarchy in the derived objects that are generated. What makes linguistic objects distinctive is that they are the products of recursive processes that deliver hierarchically structured objects.
[5] Note that unbounded implies novelty, but novelty can exist without unboundedness. The creativity issue relates to easy handling of novel structures. This can occur even in small finite domains. Creativity implies projection, which must specify a dimension of generalization along which inputs can be extended to apply to instances beyond the input. Unboundedness makes projection a no-brainer. It further implies that the generalization involves recursive rules. Unboundedness cannot be pattern matching. It requires a specification of rules that can be repeatedly applied to create novel patterns. Thus, it is important to keep the issue of unboundedness separate from that of projection. What makes the unboundedness of syntax so important is that it requires that we move beyond the pattern-template-categorization conception of cognition.
[6] It is arguable that some rules are more manifest in the data that others are and so are more accessible to inductive procedures. Chomsky makes this distinction in L&M, contrasting surface structures which contains “formal properties that are explicit in the signal” to deep structure and transformations for which there is very little to no such information in the signal (L&M:19). For another discussion of this distinction see (here).
[7] Thus the hope of unearthing phrases via differential intra-phrase versus inter-phrase transition probabilities.
[8] We really should distinguish between ‘learning’ and ‘acquisition.’ We should reserve the first term for the pattern recognition variety and adopt the second for the induction to rules variety. Problems of the second type call for different tools/approaches than those in the first and calling both ‘learning’ merely obscures this fact and confuses matters.
[9] Although this is a sermon for another time, it is important to understand what a good model does: it characterizes the underlying mechanism. Good models model mechanism, not data. Data provides evidence for mechanism, and unless it does so, it is of little scientific interest. Thus, if a model identifies the wrong mechanism not matter how apparently successful in covering data, then it is the wrong model. Period. That’s one of the reasons connectionist models are of little interest, at least when it comes to syntactic matters.
            I should add, that analogous creativity concerns drive Gallistel’s arguments against connectionist brain models. He notes that many animals display an effectively infinite variety of behaviors in specific domains (caching behavior in birds or dead reckoning in ants) and that these cannot be handled by connectionist devices that simply track the patterns attested. If Gallistel is right (and you know that I think he is) then the failure to appreciate the logic of infinity makes many current models of mind and brain beside the point.
[10] Note that unbounded implies novelty, but novelty can exist without unboundedness. The creativity issue relates to easy handling of novel structures. This can occur even in small sets. Creativity implies projection which must specify a dimension of generalization along which inputs can be extended to apply to instances beyond the input. Unboundedness makes projection a no-brainer. It further implies that the generalization is due to recursive rules that require more than establishing a fixed number of patterns that can be repeatedly filled to create novel instances of that pattern.

Wednesday, January 15, 2014

Jeff W comments on comments on recursion

I asked Jeff Watumall to respond to some of the points made concerning our previously flagged paper. He was the real driving force behind our joint effort. Thx Jeff. Here's what Jeff has to say.

******


On “On Recursion”

Our paper (http://www.frontiersin.org/Journal/10.3389/fpsyg.2013.01017/abstract) has generated interesting discussion in a previous post (http://facultyoflanguage.blogspot.com/2014/01/more-on-recursion.html).  Here I comment on those comments.

Turing and Gödel:

It is no error to equate Turing computability with Gödel recursiveness.  Gödel was explicit on this point (I am quoting from numerous Gödel papers in his Collective Works; I can furnish references if requested): “A formal system can simply be defined to be any mechanical procedure for producing formulas, called provable formulas[...].  Turing’s work gives an analysis of the concept of ‘mechanical procedure’ (alias ‘algorithm’ or ‘computation procedure’ or ‘finite combinatorial procedure’).  This concept is shown to be equivalent with that of a ‘Turing machine.’”  It was important to Gödel that the notion of formal system be defined so that his incompleteness results could be generalized: “That my [incompleteness] results were valid for all possible formal systems began to be plausible for me[.]  But I was completely convinced only by Turing’s paper.”  This clearly holds for the primitive recursive functions: “[primitive] recursive functions have the important property that, for each given set of values of the arguments, the value of the function can be computed by a finite procedure.”  And even prior to Turing, Gödel saw that “the converse seems to be true if, besides [primitive] recursions [...] recursions of other forms (e.g., with respect to two variables simultaneously) are admitted [i.e., general recursions].”  However, pre-Turing, Gödel thought that “[t]his cannot be proved, since the notion of finite computation is not defined, but it serves as a heuristic principle.”  But Turing proved the true generality of Gödel recursiveness.  As Gödel observed: “The greatest improvement was made possible through the precise definition of the concept of finite procedure, which plays a decisive role in these results [on the nature of formal systems].  There are several different ways of arriving at such a definition, which, however, all lead to exactly the same concept.  The most satisfactory way, in my opinion, is that of reducing the concept of finite procedure to that of a machine with a finite number of parts, as has been done by the British mathematician Turing.”  Elsewhere Gödel wrote: “In consequence of [...] the fact that due to A.M. Turing’s work a precise and unquestionably adequate definition of the general notion of formal system can now be given, a completely general version of Theorems VI and XI [of the incompleteness proofs] is now possible.” 

Intension and Extension:

Properly formulated formal systems can be understood as intensionally and extensionally equivalent to Turing machines.  In such systems the axiomatic derivations correspond to the elementary computation steps (e.g., reading/writing); this is as constructive as a Turing machine.  (There exists a machine that directly performs derivations in the formal system rather than encoding the information in binary strings to be manipulated by the machine.)  Accordingly, Gödel did not see formal systems and Turing machines as simply extensionally equivalent: a formal system is as constructive as a proof: “We require that the rules of inference, and the definitions of meaningful formulas and axioms, be constructive; that is, for each rule of inference there shall be a finite procedure for determining whether a given formula B is an immediate consequence (by that rule) of given formulas A1, ..., An[.]  This requirement for the rules and axioms is equivalent to the requirement that it should be possible to build a finite machine, in the precise sense of a ‘Turing machine,’ which will write down all the consequences of the axioms one after the other.”  This equivalence of formal systems with Turing machines established an absoluteness: “It may be shown that a function which is computable in one of the systems Si or even in a system of transfinite type, is already computable in S1.  Thus, the concept ‘computable’ is in a certain definite sense ‘absolute,’ while practically all other familiar metamathematical concepts depend quite essentially on the system with respect to which they are defined.”  Gödel saw it as “a kind of miracle that”, in this equivalence of computability and recursiveness, “one has for the first time succeeded in giving an absolute definition of an interesting epistemological notion, i.e., one not depending on the formalism chosen.”  Emil Post went further into ontology: The success of proving these equivalences raises Turing-computability/Gödel-recursiveness “not so much to a definition or to an axiom but to a natural law” (Post 1936: 105).  As a natural law, computability/recursiveness applies to any computational system, including a generative grammar.

Rules and Lists:

The important aspect of the recursive-function/lookup-table distinction is not computability per se (table look-up is trivially computable) but explanation.  A recursive function derives--and thus explains--a value.  A look-up table stipulates--and thus does not explain--a value.  (The recursive function establishes epistemological and ontological foundations.)  Turing emphasized this distinction, with characteristic wit, in discussing “Solvable and Unsolvable Problems” (1954).  Imagine a puzzle-game with a finite number of movable squares. “Is there a systematic way of [solving the puzzle?]  It would be quite enough to say: ‘Certainly [b]y making a list of all the positions and working through all the moves, one can divide the positions into classes, such that sliding the squares allows one to get to any position which is in the same class as the one started from.  By looking up which classes the two positions belong to one can tell whether one can get from one to the other or not.’  This is all, of course, perfectly true, but one would hardly find such remarks helpful if they were made in reply to a request for an explanation of how the puzzle should be done.  In fact they are so obvious that under the circumstances one might find them somehow rather insulting.”  Indeed.  A look-up table is arbitrary; it is equivalent to a memorized or genetically preprogrammed list.  This may suffice for, say, nonhuman animal communication, but not natural language.  This is particularly important for an infinite system (such as language), for as Turing explains: “A finite number of answers will deal with a question about a finite number of objects,” such as a finite repertoire of memorized/preprogrammed calls.  But “[w]hen the number is infinite, or in some way not yet completed[...], a list of answers will not suffice.  Some kind of rule or systematic procedure must be given.”  Gallistel and King (2009: xi) follow Turing’s logic: “a compact procedure is a composition of functions that is guaranteed to generate (rather than retrieve, as in table look-up) the symbol for the value of an n-argument function, for any arguments in the domain of the function.  The distinction between a look-up table and a compact generative procedure is critical for students of the functional architecture of the brain.  One widely entertained functional architecture, the neural network architecture, implements arithmetic and other basic functions by table look-up of nominal symbols rather than by mechanisms that implement compact procedures on compactly encoded symbols.”

Iteration and Tail Recursion:


This is mathematics, not computer science.  (Or, rather, I am a mathematician, now interloping in linguistics.  In mathematics, iteration--a general notion applicable to a pattern of succession--is seen as a form of recursion: the function f is defined for an argument x by a previously defined value (e.g., f(y), y < x); but iteration is “tail” recursion given that the previously defined value y is the immediately previously define value.)  We are on the computational level, not the level of mechanisms.  It is important to recall that Marr and Nishihara (1978) distinguished four--not three--levels: “At the lowest, there is the basic component and circuit analysis--how do transistors (or neurons), diodes (or synapses) work?  The second level is the study of particular mechanisms: adders, multipliers, and memories, these being assemblies made from basic components.  The third level is that of the algorithm, the scheme for a computation; and the top level contains the theory of computation.”  (The theory of computation is mathematical.)  Much of the muddling of iteration and tail recursion in the comments on the previous post is the result of misclassifying the level of analysis.  “[W]e may consider the study of grammar and UG to be at the level of the theory of computation” (Chomsky 1980: 48).  Thus discussion of loops, arrays, etc. is irrelevant.  In fact, algorithms and mechanisms are arguable irrelevant in principle.  We concur with Chomsky that, for the computational system of language “there’s no algorithm for the system itself; it’s kind of a category mistake.  [T]here’s no calculation of knowledge; it’s just a system of knowledge[...].  You don’t ask the question what’s the process defined by Peano’s axioms and the rules of inference, there’s no process” (Chomsky 2013a).  Analogously, a Turing machine is not a description of a process or algorithm or mechanism but “a mathematical characterization of a class of numerical functions” (in the words of Martin David (1958: 3), one of the founders of computability theory).  Thus to define the faculty of language as a type of Turing machine as we did in our paper, “On Recursion,” is to give a function: “a finite characterization of an infinite set” (Chomsky 2013b).  A Turing machine--and thus the language faculty--is defined by a tuple containing a finite set of symbols (axioms), a set of states (with “states” defined as “structures” in the sense of mathematical logic), and a transition function (rule of inference) mapping from state/symbol to state/symbol.  “A derivation is thus roughly analogous to a proof with Σ,” a finite set of initial symbols, “taken as the axiom system and F,” the finite set of rewrite rules (or Merge), “[taken] as the rules of inference” (Chomsky 1956: 117), consistent with Gödel’s characterization: “We require that the rules of inference, and the definitions of meaningful formulas and axioms, be constructive; that is, for each rule of inference there shall be a finite procedure for determining whether a given formula B is an immediate consequence (by that rule) of given formulas A1, ..., An[.]  This requirement for the rules and axioms is equivalent to the requirement that it should be possible to build a finite machine, in the precise sense of a ‘Turing machine,’ which will write down all the consequences of the axioms one after the other.