Comments

Showing posts sorted by relevance for query domain specificity. Sort by date Show all posts
Showing posts sorted by relevance for query domain specificity. Sort by date Show all posts

Monday, October 7, 2013

Domain Specificity


One of the things that seems to bug many about FL/UG is the supposition that it is a domain specific module dedicated to ferreting out specifically linguistic information. Even those that have been reconciled to the possibility that minds/brains come chock full of pre-packaging, seem loath to assume that these natively provided innards are linguistically dedicated.  This antipathy has even come to afflict generativists of the MP stripe for there is believed to be an inconsistency between domain specificity and the minimalist ambition to simplify UG by removing large parts of its linguistically idiosyncratic structure. I have suggested elsewhere (here) that this tension is only apparent, and that it is entirely possible to pursue the MP cognitive leveling strategy without abandoning the idea that there is a dedicated FL/UG as part of human biological endowment. In a new paper, Gallistel and Matzel (G&M) (here) argue that domain specificity is the biological default once one rejects associationism and adopts an information processing model of cognition. Put more crudely, allergy to domain specificity is just another symptom of latent empiricism (i.e. a sad legacy of associationism).

I, of course, endorse G&M’s position that there is nothing inconsistent between accepting functionally differentiated modules and the assumption that these are largely constructed using common basic operations. And, I of course love G&M’s position that once one drops any associationist sympathies (and I urge you all to immediately do this for your own intellectual well being!), then the hunt for general learning mechanisms looks, at least in biological domains, ill-advised.  Or put more positively: once one adopts an information processing perspective then domain specificity seems obviously correct. Let’s consider G&M’s points in a little detail.

G&M contrasts associationist (A) and information processing (IP) models of learning and memory. The G&M paper is divided into two parts, more or less. The first several pages comprise a concise critique of associationist/neural net models in which learning is “the rewiring of a plastic nervous system by experience, and memory resides in the changed wiring (170).” The second part develops the evidence for an IP perspective on neural computations. The IP models contrast with A-models in distinguishing the mechanisms for learning (whose function is to “extract potentially useful information from experience”) and those for memory (whose function is to “carr[y] the acquired information forward in time in a computationally accessible form that is acted upon by the animal at the time of retrieval”) (170).  Here are some of their central points.

A-models are “recapitulative.” What G&M (170) intend here is that learning consists in finding the pattern in the data (see here): “An input that is part of the training input, or similar to it, evokes the trained output, or an output similar to it.”  IP models are “in no way recapitulations of the mappings (if any) that occurred during the learning.” This is the classical difference between rationalist vs empiricist conceptions of learning. A- models conceive of environmental input as adapting “behavior” to environmental circumstances. IP-models conceive of learning as building a “representation of important aspects of the experienced world.”

A-models gain a lot of their purchase within neuro-science (and psychology) by appearing to link so directly to a possible neural mechanism; long-term potentiation (LTP). However, G&M argue vigorously that LTP support for A-models entirely evaporates when the evidence linking LTP to A-models is carefully evaluated.  G&M walk us slowly through the various disconnects between standard A-processes of learning and LTP function; their time scales are completely different (“…temporal properties of LTP do not explain the temporal properties of behaviorally measured association formation (172).”), their persistence (i.e. how long the changes in LTP vs associations last) is entirely different and so “LTP does not explain the persistence of associative learning (172),” their reactivation schedules are entirely different (i.e. If L is learned and then extinguished, L is reacquired more quickly, but “LTP is neither more easily nor more persistent than it was after previous inductions.”), nor do LTP models provide any mechanism for solving the encoding problem (viz. A-learning is mediated by the comparison of different kinds of temporal intervals and there is no obvious way for LTP nets to do this) except by noting that what gets encoded is emergent (and this amounts to punting on the encoding problem, rather than addressing it).

In short, there is no support for A-models from neural LTP models. Indeed, the latter seem entirely out of synch with what’s needed to explain memory and learning. As G&M put it: “…if synaptic LTP is the mechanism of associative learning- and more generally, of memory- then it is disappointing that its properties explain neither the basic properties of associative learning nor the essential properties of a memory mechanism (173).” So much for the oft insinuated claim that connectionist models are preferable because they are neurally plausible (indeed, obvious!).

General conclusion: A-models have no obvious support from standard LTP models and these standard LTP models are inadequate for handling the simplest behavioral data. In effect, A- and LTP- accounts are the wrong kinds of theories (not wrong in detail, but in conception and hence without much (if any) redeeming scientific value) if one is interested in understanding the neural bases of cognition.

So what’s the right approach? IP-models. In the last parts of the paper G&M go over some examples.  They note that biologically plausible IP models will all share some important features:

1.     They will involve domain specific computations. Why? “Because no general purpose computation could serve the demands of all types of learning (175),” i.e. domain specificity is the natural expectation for IP models of neuro-cognition.
2.     The different computations will apply the same “primitive operations” in achieving functionally different results (175).[1]
3.     The IP approach to learning mechanisms “requires an understanding of the rudiments of the different domains in which the different learning mechanisms operate” (175). So, for example, figuring out if A is cause of B, or A is the edge of B will involve different computations from each other and from those that mediate the pairing of meanings and/with sounds.
4.     Though the neuro-science is at a primitive stage right now, “…if learning is the result of domain-specific computations, then studying the mechanism of learning is indistinguishable from studying the neural mechanisms that implement computations (175).”

Note that this will hold as much in the domain of language as in navigation and spatial representation.  In other words, once one dumps Associationism (as one must as it is empirically completely inadequate and intellectually toxic) then domain specificity is virtually ineluctable. There exist no interesting general purpose learning systems (just as there it no general sensing mechanism, as Gallistel has been wont to observe). That’s the G&M message. Cognitive computation, if it’s to be neurally based, will be quite specifically tailored to the cognitive tasks at hand, even if built from common primitive circuits.

The most interesting part of G&M, at least to me, was the review of the specific neural cells implicated in animal capacities for locating oneself in space and moving around within it. It seems that neuro-scientists are finding “functionally specialized neurons [that] signal abstract properties of the animal’s relation to its spatial environment (185).” These are genetically controlled and, as G&M note, their functional specialization provide “compelling evidence for problem-specific mechanisms.”

Note that the points G&M make above fit very snugly with standard assumptions within the Chomsky version of the generative tradition. In other words, the assumptions that generative linguists make concerning domain specific computations and mechanisms (though not necessarily primitive operations) simply reflects what is, or at least should be, the standard assumption in the study of cognition once Associationism is dumped (may it’s baneful influence soon disappear). If G&M are right, then there are no good reasons from neuro-biology for thinking that the standard assumptions concerning native domain specific structures for language are exotic or untoward.  They are neither. The problem is not with these assumptions, but with the unholy alliance between some parts of contemporary neuroscience and the A-models of learning and cognition that neuro types have uncritically accepted.

If you read the whole G&M paper (some parts involve some heavy lifting) and translate it into a linguistics framework, it is very hard to avoid the conclusion that if G&M are correct, (and, in case you’ve missed it, IMO they are) then the Chomskyan conception of language, mind, and brain is both anodyne and the only plausible game in cog-neuro town.


[1] A similar point wrt MP and UG is made here.

Wednesday, September 12, 2018

The neural autonomy of syntax

Nothing does language like humans do language. This is not a hypothesis. It is a simple fact. Nonetheless, it is often either questioned or only reluctantly conceded. Therefore, I urge you to repeat the first sentence of this post three times before moving forward. It is both true and a truism. 

Let’s go further. The truth of this observation suggests the following non-trivial inference: there is something biologically special about humans that enables them (us) to be linguistically proficient andthis special mental power is linguistically specific. In other words, humans are uniquely cognitively endowed as a matter of biology when it comes to language and this biological gift is tailored to track some specific cognitive feature of language rather than (for example) being (just!) a general increase in (say)generalbrain power. On this view, the traditional GG conception stemming from Chomsky takes FL to be both species specific and domain specific. 

Before proceeding, let me at once note that these are independent specificity theses. I do this because every time I make this point, others insist in warning me that the fact mentioned in the first sentence does not imply the inference I just drew in the second paragraph. Quite right. In fact: 

It is logically possible that linguistic competence supervenes on no domain specific capacities but is still species specific in that only humans have (for example) sufficiently powerful general brains to be linguistically proficient. Say, for example, linguistic competence requires at least 500 units of cognitive power (CP) and only human brains can generate this much CP. However, modulo the extra CPs, the mental “programs” the CPs drive are the same as those that (at least some) other cognitive creatures enjoy, they just cannot drive them as fast or as far because of mileage restrictions imposed by low CP brains.

Similarly, it is logically possible that animals other than humans have domain specific linguistic powers. It is conceivable that apes, corvids, platypuses, manatees, and Portuguese water dogs all have brains that include FLs just like ours that are linguistically specific (e.g. syntax focused and not exercised in other cognitive endeavors). Were this so, then both they and we would have brains with specific linguistic sensitivities in virtue of having brains with linguistically bespoke wiring/circuitry or whatever specially tailored brain ware makes FL brains special. Of course, were I one of them I would keep this to myself as humans have the unfortunate tendency of dismembering anything that might yield scientific insight (or just might be tasty). If these other animals actually had an FL I am pretty sure some NIH scientist would be trying to figure out how to slice and dice their brains in order to figure out how its FL ticks.

So, both options are logically possible, but, the GG tradition stemming from Chomsky (and this includes yours truly, a fully paid up member of this tribe) has doubted that these logical options are live and that when it comes to language onlywe humans are built for it and what makes our cognitive profile special is a set of linguistically specific cognitive functions built into FL and dedicated to linguistic cognition. Or, to put this another way, FL has some special cognitive sauce that allows us to be as linguistically adept as we evidently are and we alone have minds/brains with this FL.

Nor do the exciting leaps of inference stop here. GG has gone even further out on the empirical limb and suggested that the bespoke property of FL that makes us linguistically special involves an autonomous SYNTAX (i.e. a syntax irreducible to either semantics or phonology and with its own special combinatoric properties). That’s right readers, syntax makes the linguistic world go round and only we got it and that’s why we are so linguistically special![1]Indeed, if a modern linguistic Ms or Mr Hillel were asked to sum up GG while standing on one foot s/he could do worse than say, only humans have syntax, all the rest is commentary.

This line of reasoning has been (and still is) considered very contentious. However, I recently ran across a paper by Campbell and Tyler (here, henceforth C&T) that argues for roughly this point (thx to Johan Bolhuis and William Matchin for sending it along). The paper has several interesting features, but perhaps the most intriguing (to me) is that Tyler is one of the authors. If memory serves, when I was growing up, Tyler was one of those who were very skeptical that there was anything cognitively special about language. Happily, it seems that times have changed.

C&T argues that brain localizes syntactic processing in the left frontotemporal lobe and “makes a strong case for the domain specificity of the frontotemporal syntax system and its autonomy from domain-general networks” (132). So, the paper argues for a neural version of the autonomy of syntax thesis. Let me say a few more words about it.

First, C&T notes that (of course) the syntax dedicated part of the brain regularly interacts with the non-syntactic domain general parts of the brain. However, the paper rightly notes that this does not argue against the claim that there is an autonomous syntactic system encoded in the brain. It merely means that finding it will be hard as this independence will often be obscured.  More particularly C&T says the activation of the domain general systems only arise “during task based language comprehension” (133). Tasks include having to make an acceptability judgment. When we focus on pure comprehension, however, without requiring any further “task” we find that “only the left-laterilized frontotemporal syntax system and auditory networks are activated” (133). Thus, the syntax system only links to the domain general ones during “overt task performance” and otherwise activates alone. C&T note that this implies that the syntactic system alone is sufficient for syntactic analysis during language comprehension.

Second, C&T argue that arguments against the neural autonomy of syntax rest on bad definitions of domain specificity. More particularly, according to C&T the benchmarks for autonomy in other studies beg the autonomy question by embedding a “task” in the measure and so “lead to the activation of additional domain-general regions” (133). As C&T notes, when such “tasks” are controlled for, we only find activation in the syntax region.

Third, the relevant notion of syntax is the one GGers know and love. For C&T takes syntax to be prime species specific feature of the brain and understands syntax in GGish terms to be implicated in “the construction of hierarchical syntactic structures.” C&T contrasts hierarchical relations with “adjacency relationships” which it claims “both human and non-human primates are sensitive to” (134). This is pretty much the conventional GG view and C&T endorses it.

And there is more. C&T endorses the Hauser, Chomsky, Fitch distinction between FLN and FLB. This is not surprising for once one adopts an autonomy of syntax thesis and appreciates the uniqueness of syntax in human minds/brains the distinction follows pretty quickly. Let me quote C&T (135):

In this brief overview, we have suggested that it is necessary to take a more nuanced view to differentiating domain-general and domain-specific components involved in language. While syntax seems to meet the criteria for domain-specificity….there are other key components in the wider language system which are domain-general in that they are also involved in a number of cognitive functions which do not involve language.

C&T has one last intriguing feature, at least for a GGer like me. The name ‘Chomsky’ or the terms ‘generative grammar’ are never mentioned, not even once (shades of Voldemort!). Quite clearly, the set of ideas that the paper explores presupposes the basic correctness of the Chomskyan generative enterprise. C&T arugues for a neural autonomy of syntax thesis and, in doing so, it relies on the main contours of the Chomsky/GG conception of FL. Yes, if C&T is correct it adds to this body of thought. But it clearly relies on it’s main claims and presupposes their essential correctness. A word to this effect would have been nice to see. That said, read the paper. Contrary to the assumptions of many, it argues that for a cog-neuro conception of the Chomsky conception of language. Even if it dares not speak his name.


[1]I suspect that waggle dancing bees and dead reckoning insects also non verbally advance a cognitive exceptionalism thesis and preen accordingly.

Sunday, August 7, 2016

Pullum on Putnam

Geoff Pullum has a recent piece in The Chronicle (here) in which he praises a deservedly famous man, Hilary Putnam. Putnam was an important 20-21st century philosopher who compiled what is arguably the best collection of essays ever in analytic philosophy. Pullum notes all of this and I cannot fault him for his judgment. However, he then takes one more step, he lauds Putnam as the “world’s most brilliant, insightful, and prescient philosopher of linguistics.” That Putnam was brilliant and insightful and (maybe) prescient is not something that I would (and did not (here)) contest. That any of this extended to his discussions of linguistics topics strikes me as either a sad commentary on the state of the philosophy of linguistics (this gets my vote) or hyperbole (it was a belated obit after all). At any rate, I want to make clear why I think that Putnam’s writings on these matters are best treated as object lessons rather than insights.  Happily, this coincides with my re-reading of Language and Mind.  Chomsky takes on some of Putnam’s more (sadly) influential criticisms of GG and (I am sure you will not be surprised to hear from me) demolishes them. The gist of Chomsky’s reply is that there is very little there there. He is right. This has not stopped analogous criticisms from repeatedly being advanced, but they have not gotten more convincing by repetition. Let me elaborate.

Putnam’s most directed critique of the Chomsky program in GG were his 1967 Synthese paper (“The ‘Innateness Hypothesis’ and Explanatory Models in Linguistics”) and a later companion piece “What is innate and why.” Chomsky considers Putnam’s arguments in detail in chapter 6 of the expanded edition of Language and Mind entitled “Linguistics and Philosophy.” Here is the play by play.

Chomsky’s critique has three parts:

1. Putnam’s specific critiques “enormously underestimate and misdescribe, the richness of structure, the particular and detailed properties of grammatical form and organization that must be accounted for by a “language acquisition model,” that are acquired by the normal speaker-hearer and that appear to be uniform among speakers and across languages” (179-180).
2. Putnam’s computational claims concerning grammatical simplicity are unfounded (181-2).
3. There is no argument for Putnam’s claim that “general multipurpose learning strategies” are sufficient to account for G acquisition and there is no current reason to think that any such exist when one looks at the grammatical details (184-5).

These are all closely related points, and Pullum is correct in suggesting that these points have repeatedly reappeared in critiques of GG. Thus, it is still true that simplistic views of what is required for G acquisition rely on underestimating and misdescribing what must be explained. It is still true that claims made on behalf of general learning strategies eschew the hard work of showing how the many “laws” linguists have discovered over the last 60 years are to be acquired without quite a bit of what looks like language specific software. Pullum is right: the critics have repeatedly picked up Putnam’s objections even after these have been shown to be inadequate and/or beside the point. Putnam has indeed been influential, and we are the worse for it.

Let me lightly elaborate on these three points.

Critics regularly avoid the hard problems. For example, look at virtually any takedown in the computer science literature of, for example, structure dependence, and you will observe this (see here and here for two recent reiterations of this complaint). All of these miss the point of the argument for structure dependence by concentrating on easily understandable illustrative toy examples intended for the general public (Reali and Christiensan) or misconstruing what the term actually denotes (Perfors et. al.).

I have said this before and I will do so again: GGers have discovered many non-trivial mid level generalizations that are the detailed fodder that fuel Poverty of Stimulus (PoS) arguments that implicate linguistically specific structure for FL. There can be no refutation of these arguments if the generalizaions are not addressed. So, Island Effects, ECP effects, Binding effects, Cross Over effects etc. constitute the “hard problems” for non-domain specific learning architectures. If you think that a general learner is the right way to go you need to account for these sorts of data. And there is, by now, a lot of this (see here for a partial list). However, advocates of “simpler” less domain specific accounts (almost) never broach these details, though absent this the counter proposals are at best insufficient and at worst idle.

It seems that Putnam is the first in a continuing line of critics that have decided that one can ignore the linguistic details when arguing against undesired cognitive conclusions. As Chomsky notes contra Putnam, there is more to phonology than a “short list of phonemes” from which languages can choose (e.g. there is also cyclic rule application) and there is more to syntax than proper names (e.g. there are also Island effects). Putnam failed to engage with the details (as discussed in Chomsky’s work at the time) and in doing so established a tradition that many have followed. It is not, however, a tradition that anyone should be proud to be part of, whatever its pedigree.

Putnam advanced another argument that is sadly still alive today. He argued that invoking innateness doesn’t solve the acquisition problem, but only “postpones” it. What’s this mean? The argument seems to be that stuffing FL with innate structure is explanatorily sterile as it simply pushes the problem back one step: how did the innate structure get there?[1] Frankly, I find this claim philosophically embarrassing. Why?

One of the main professional requirements of a card-carrying philosopher is that her/his work clarify what point an argument is aiming to make; what question is it trying to answer? Assuming an FL that is structured with domain specific linguistic structure addresses the question of how an LAD can acquire its language specific G despite the poverty of the relevant PLD (here’s Chomsky 184-5: “Invoking an innate representation of universal grammar does solve the problem of learning (at least partially), in this case.”) If such a UG structured FL suffices to solve the PoS problem it raises a second question: how did the relevant (domain specific) mental structure get there (i.e. why is FL structured with language proprietary UGish principles). Note, these are two different questions (viz. “what FL is required to project a GL from PLDL?” is different from “how did the FL we in fact have get embedded in our mental architecture in the first place?”). Consequently failing to answer the evolutionary questions concerning the etiology of rich innate mental structure does not imply a failure to answer/address the question of how an individual LAD acquires its GL.

Of course, it is not an irrelevant either, or might not be. If we could show that a given domain specific FL could not possibly have evolved then we are pretty sure that the postulated innate mental mechanism in the individual cannot be a causal factor in G acquisition. After all, if such an FL cannot be there then it isn’t there and if it isn’t there then it cannot help with the acquisition problem. But, and this is very important, nobody has even the inklings of an argument against the assumption that even a very rich domain specific FL could not have arisen in humans. Right now, this impossibility claim is at best a hunch (viz. an ungrounded prejudice). Why? Because we currently have very few ideas about how any cognitive structures evolve (as Lewontin has famously noted). Indeed, even the evolution of seemingly simple non-cognitive structures remains mysterious (see here for a recent example). So, any confident claims that even a richly domain specific FL is evolutionarily impossible is not on the cards right now and is thus a weak counter-argument against an FL that can solve the acquisition problem.[2]

A sidebar: now of this is meant to imply that this evolutionary question is uninteresting. I am an unrepentant Minimalist and take seriously the minimalist problematic: how could an FL such as ours arisen in the species. As such I am all in favor of purging FL of as much UG as possible and trading this for general cognitive mechanisms. However, because I consider this an interesting problem I resist fiat solutions; you know, bold yet vacuous declarations that a general learner can do it all without any detailed indications dealing with specific claims resting on bland assurances that it is in principle possible. I like the question so much that I want to see details; actual explanations engaging with specific proposed UG structures. I love reduction, I just don’t like the cheap variety. So derive your favorite UG based accounts from more general principles and watch me snap to attention.

BTW, Chomsky makes just this point as early as 1972. Here is a quote from his discussion of Putnam (182):

I would, naturally, assume that there is some more general basis in human mental structure for the fact (if it is a fact) that languages have transformational grammars; one of the primary scientific reasons for studying language is that this study may provide some insight into general properties of mind. Given those specific properties, we may then be able to show that transformational grammars are “natural.” This would constitute real progress, since it would enable us to raise the problem of innate conditions on acquisition of knowledge and belief in a more general framework. But it must be emphasized that, contrary to what Putnam asserts, there is no basis for assuming that “reasonable computing systems” will naturally be organized in the specific manner suggested by transformational grammar.

One might argue that Chomsky’s version of minimalism is his way of making good on Putnam’s computational conjecture, though I doubt that Putnam would see it that way. At any rate, Minimalism starts from the recognition that domain specific FLs can solve standard linguistic acquisition problems (i.e. PoS problems) and then tries to reduce the linguistic specificity of the various principles. It does not solve the domain specificity problem by ignoring the relevant domain specific principles.

One more point and I end. In his reply to Putnam Chomsky outlines a very reasonable strategy for eliminating domain specificity in favor of something like general learning.[3] In his words 184):

A non dogmatic approach to this problem [i.e. the acquisition of language NH] can be pursued, though the investigation of specific areas of human competence, such as language, followed by the attempt to devise a hypothesis that will account for the development of such competence. If we discover that the same “learning strategies” are involved in a variety of cases, and that these suffice to account for the acquired competence, then we will have good reason to believe Putnam’s empirical hypothesis is correct. If, on the other hand, we discover that different innate systems…have to be postulated, then we will have good reason to believe that an adequate theory of mind will incorporate separate “faculties,’ each with unique or partially unique properties.

See here for another discussion elaborating these themes.

To sum up: The problem with Putnam’s philosophical discussions of linguistics is that they entirely missed the mark. They were based on very little detailed knowledge of the GG of the time. They confused several questions that needed to be kept separate and they philosophically begged questions that were (and still are) effectively empirical. The legacy has been a trail of really bad arguments that seem to arise zombie like despite their inadequacy. Putnam wrote many interesting papers. Unfortunately his papers on linguistics are not among these. Let these rest in peace.[4]



[1] There are actually two points being run together here. The first is that any innate structure whether it is domain specific or not begs the explanatory question. The second is that only a domain specific “rich” FL does so. The form of the argument Putnam presents applies to either for both call for an evolutionary account of how the mental capacities arose. Humans might after all have a richer general cognitive apparatus than our ape cousins and how it arose would demand explanation ever if it were not domain specific. However, the thinking usually is that only domain specific richness is problematic. In what follows I abstract from this ambiguity.
[2] Gallsitel has noted that cognitive domain specificity is biologically quite reasonable (see here for discussion and links).
[3] See here for another discussion along the same lines channeling Reflections on Language
[4] Perhaps it is not surprise that Dan Everett loved this Pullum post. In his words: “Glad you noticed this! He was indeed one of the best of the last 100 years.” This comment does not indicate what Everett found so wonderful, but given the topic of the Pullum’s post and Everett’s own added confusions to the philosophical issues, it is reasonable to assume that he found the Putnam critiques against domain specific nativism compelling. But you knew he would, right?

Tuesday, November 17, 2015

Never thought I would say this

Never thought I would say this, but I found that I resonated positively to a recent small comment by Chris Manning on Deep Learning (DL) that Aaron White sent my way (here). It seems that the DL has computational linguistics (CL) of the Manning variety in its sights. Some DLers apparently believe that CL is just is nano-moments away from extinction. Here’s a great quote from one of the DL doyens:

NLP is kind of like a rabbit in the headlights of the Deep Learning machine, waiting to be flattened.

DL wise men like Geoff Hinton have already announced that they expect that machines will soon be able to watch videos and “tell a story about what happened” and be downsized onto an in-your-ear chip that can translate into English on the fly. Great things are clearly expected. Personally, I am skeptical as I’ve heard such hyperbole before. We have been five years away from this sort of stuff for a very long time.

Moreover, I am not alone. If I read Manning correctly, he is skeptical (though very politely so) as well.[1] But, like me, he sees an opportunity here, one I noted before (here and here). Of course we likely disagree about what kind of linguistics will be most useful for advancing these technological ends,[2] but when it comes to engineering projects I am very catholic in my tastes.

What’s the opportunity consist in? It relies on a bet: that generic machine learning (even of the DL variety) will not be able to solve the “domain problem.” The latter is the belief that how a domain of knowledge is structured matters a lot even if one’s aim is to solve an engineering problem.

An aside: shouldn’t those that think that the domain problem is a serious engineering hurdle also think that modularity is a good biological design feature? And shouldn’t these people therefore think that the domain specificity of FoL is a no-brainer? In other words, shouldn’t the idea that humans have domain specific knowledge that allows them to “solve” language problems (and support human facile acquisition and use) be the default position? Chris?  What think you? Dump general learning approaches and embrace domain specificity?

Back to the main point: The bet. So, if you think that using word contexts can only get you so far (and not interestingly far either), then you are ready to bet that knowing something about language will be useful in solving these engineering problems. And that provides linguists with an opportunity to ply their trade. In fact, Manning points to a couple of projects aimed at developing “a common syntactic dependency representation and POS (‘part of speech,’ NH) and feature label sets which can be used with reasonable linguistic fidelity and human usability across all human languages” (3).[3] He also advocates developing analogous representations for “Abstract Meaning.” This looks like the kind of thing that GGers could usefully contribute to. In other words, what we do directly fits into the Manning project.

Another aside: do not confuse this with investigating the structure of FL.  What matters for this project is a reasonable set of Greenberg “Universals.” Indeed, being too abstract might not be that useful practically, and being truly universal is not that important (what is important is finding those categories that best fit the particular languages of interest). This is not a bad thing. Engineering is not to be disparaged. It’s just not the same project as the one that GG has scientifically set for itself. Of course, should the Chomsky version of GG succeed, it is possible that it will contribute to the engineering problem. But then again, it might not. As I understand it, General Relativity has yet to make a big impact on land surveying. It really all depends (to fix ideas think birds and planes or fish and submarines. Last time I looked plane wings don’t flap and sub bodies don’t undulate).

Manning makes lots of useful comments about DL, many of which I didn’t understand. He makes some, however, that I did. For example, his the observation that DL has mainly proved useful in signal processing contexts (2) (i.e. where the problem is to get the generalization that is in the data, the pattern from (noisy) patternings). The language problem, as I’ve argued, is different from this (see here) so the limits of brute force DL will, I predict, become evident when the new wise men turn their attention to these. In fact, I make a more refined prediction: to “solve” this problem DLers will either (i) ignore it, (ii) restrict the domain of interest to finesse it or (iii) promise repeatedly that the solution is but 5 years away. This has happened before and will happen again unless the intricate structural constraints that characterize language are recognized and incorporated.

Manning also makes several points that I would take issue with. For example, IMO he (like many others) confuses squishy data for squishy underlying categories. See, in particular, Manning’s discussion of gerunds on p. 4. That the data does not exhibit sharp boundaries does not imply that the underlying structures are not sharp. In fact, at some level they must be for under every probabilistic theory there is a categorical algebra.  I leave it to you out there to come up with an alternative analysis of Manning’s observed data set. I give you a 30 second time limit to make it challenging.

At any rate, you will not be surprised to find out that I disagree with many of Manning’s comments. What might surprise you is that I think he is right in his reaction to DL hubris and he is right that there is an opportunity for what GGers know to be of practical value. There is no reason for DL (or Bayes or stats) to be inimical to GG. It’s just technology. What makes its practice often anathema is the hard-core empiricism gratuitously adopted by its practitioners. But this is not inherent to the technology. It is only a bias of the technologists. And there are some like Jordan and Manning and Reisinger who seem to get this. It looks like an opportunity for GGers to make a contribution? One, incidentally, that can have positive repercussions for the standing of GG. Scientific success does not require technological application. But having technological relevance does not hurt either.



[1] I confess to a touch of schadenfreude given that this is the kind of thing that Manning and Co like to say about my kind of linguistics wrt to their CL approaches.
[2] Though I am not confident about this. I am pretty confident about what kind of linguistics one needs to advance the cognitive project. I am far less sure about what one needs to advance the engineering one. In fact, I suspect that a more “surfacy” syntax will fit the latter’s design requirements better than a more abstract one given its NLPish practical aims. See below for a little more discussion.
[3] I have it from a reliable source that this project is being funded by Google to the tune of millions. I have no idea how many millions, but given that billions are rounding errors to these guys, I suspect that there is real gold in them thar hills.

Thursday, December 6, 2012

Reply To Alex

This is a reply to Alex's reply here. It did not fit into the limited space the comment section makes available. Sorry.

Sigh. Why the bait and switch there Alex?  Where't this talk about categories coming from? But let me get there.  It seems to me that you really don't get the argument, so let me illustrate it with an example you will be familiar with as I gave it to you before.

The question on the table is whether I am entitled to a domain specific UG built with largely domain general "circuits."  Now, a priori this seems reasonable. I can build a "will read windows only" machine using the same chips that will build a "will read OS only" system. The very same chips can be used to exclusively read/use two different programming formats. It's not only disable, it has been done.  So the conceptual possibility exists.

In the grammar domain: Say I can show how using Merge and other simple non linguistically proprietary operations I can "derive" the binding theory (I try to show this in 'A Theory of Syntax' but differently (and more idiosyncratically) than I sketch here). Here's the proposal:
(i) If A is antecedent to B then A and B form a constituent.
(ii) Merge in both E and I forms is a basic operation
(iii) Full interpretation holds: A DP must be interpretable at both interfaces, this means bears both a theta role and a case value.
(iv) There is no DS and so movement into theta positions is ok.
(v) Minimality holds of movement
(vi) Extension regulates Merge
The net effect of (i)-(v) is to have reflexivization "live on" A-chains. A-chain properties follow from minimality, Extension, and full interpretation.  I take these latter two properties to reflect domain general/computationally general features of FL and so NOT special to FL (this may be wrong, but I argue for it, so let me get away with that here).

The effect of having reflexivization live on A-chains derives binding principle A (this is easy to see given the LGB relation between NP-trace and movement via binding theory. I reverse the relation relating them via movement theory).  The locality follows from (iii) and (v). The C-command condition holds from (vi).

Say for purposes of discussion this indeed derives Principle A of the binding theory as I said. Now what does the kid have to learn to master principle A? Well all but the fact that 'himself' is spelled out as the tail of the chain is "given." So that's what the kid has to "learn," i.e. that reflexives are spell outs of A-chain tails (roughly the old Lees-Klima account in gussied up form).  Note, as I indicated in a reply to an earlier question, this is all the kid has to learn on the GB theory as well (i.e. that reflexives fall under A). The same thing. This is not surprising as if successful we have derived principle A as the product of Merge plus these other principles. In other words, if I reduce Principle A to movement theory, then if FL is structured as the reducing picture envisages I am in the same position I was in wrt Plato's problem and Binding Theory as I was in the GB era. The answer to Plato's problem has not changed. The information is domain specific though the computational circuits used to build the FL circuit board that embodies the competence are largely domain general (i.e. circuits and properties available domain generally) in their properties and modes of operation.

Now, I am not saying that this is correct (though I do like it). I am asking ASSUMING IT OR SOMETHING LIKE IT CAN BE DONE  whether the fact of a Minimalist reduction means that all learning is domain general and the answer I give is no if you see that the Minimalist proposal is not competitor to the GB one but an attempt to place it on more solid foundations. So there's my eaten cake and I plan another big helping. 

Now your categorization question:MP (and GB for that matter) had very little to say about words and their categories (i.e. the generalizations adduced were not nearly as impressive as what we had to say about syntax IMHO).  Thus, what I said did not address these questions.  Truth be told, IMHO we know very little about the intricacies of word learning and the innate knowledge required to get it off the ground. Chomsky's discussion of these matters (riffing on Austin and the later Wittgenstein) is fascinating but so far theoretically inconclusive.  So, the short answer is that NOTHING I KNOW ABOUT MP HAS ANYTHING ENLIGHTENING TO SAY ABOUT THIS.  I also know that Chomsky believes the same thing.  So, as far as I can tell, we have no answer to this question from an MP point of view.  However, most arguments for rich UG were made using syntactic facts like those GB and MP do deal with so the fact that we have no MP story here strikes me as of little relevance.

In sum, what you are pointing out is that there are other important poorly understood questions. Yup, many.  Do these require domain specific innate knowledge? Who knows?  I am not being entirely flippant (though I am being a teensy bit). Here's why. MP makes sense because we have theories like GB.  Till GB came up with its laws of grammar the question of how to reduce them to simple principles was way premature. Ok, what do we know about word learning and categorization that comes even close to being interesting. Not much. So the Minimalist question is entirely out of place.  The thing about research questions is that they make sense in some areas and not in others. They make sense for syntax and so we are making some interesting progress in answering them there. I have no reason to think that they make sense for the problems you mention and so am not surprised that there is not much to say. Of course, should categorization and word acquisition be subject to domain general procedures, I would be delighted. If not, I would start to ask what makes it possible and how much domain specificity we need. But, till we have interesting "laws" here I will refrain from indulging minimalist confabulations.