Comments

Showing posts with label Jerry Fodor. Show all posts
Showing posts with label Jerry Fodor. Show all posts

Monday, August 27, 2018

Revolutions in science; a comment on Gelman

In what follows I am going to wander way beyond my level of expertise (perhaps even rudimentary competence). I am going to discuss statistics and its place in the contemporary “replication crisis” debates. So, reader be warned that you should take what I write with a very large grain of salt. 

Andrew Gelman has a long post (here, AG) where he ruminates about a comparatively small revolution in statistics that he has been a central part of (I know, it is a bit unseemly to toot your own horn, but heh, false modesty is nothing to be proud of either). It is small (or “far more trivial”) when compared to more substantial revolutions in Biology (Darwin) or Physics (Relativity and Quantum mechanics), but AG argues that the “Replication revolution” is an important step in enhancing our “understanding of how we learn about the world.” It may be right. But…

But, I am not sure that it has the narrative quite right. As AG portrays matters, the revolution need not have happened. The same ground could have been covered with “incremental corrections and adjustments.” Why weren’t they? The reactionaries forced a revolutionary change because of their reactions to reasonable criticisms by the likes of Meehl, Mayo, Ioannidis, Gelman, Simonsohn, Dreber, and “various other well-known skeptics.” Their reaction to these reasonable critiques was to charge the critics with bullying or insist that the indicated problems are all part of normal science and will eventually be removed by better training, higher standards etc. This, AG argues, was the wrong reaction and required a revolution, albeit a minor one relatively speaking, to overturn.

Now, I am very sympathetic to a large part of this position. I have long appreciated the work of the critics and have covered their work in FoL. I think that the critics have done a public service in pointing out that stats has served to confuse as often (maybe more often) than it has served to illuminate. And some have made the more important point (AG prominently among them) that this is not some mistake, but serves a need in the disciplines where it is most prominent (see here). What’s the need? Here is AG:[1]

Not understanding statistics is part of it, but another part is that people—applied researchers and also many professional statisticians—want statistics to do things it just can’t do. “Statistical significance” satisfies a real demand for certainty in the face of noise. It’s hard to teach people to accept uncertainty. I agree that we should try, but it’s tough, as so many of the incentives of publication and publicity go in the other direction.

And observe that the need is Janus faced. It faces inwards to relieve the anxiety of uncertainty and it faces outwards in relieving professional publish-or-perish anxiety. Much to AG’s credit he notices that these are different things, though they are mutually supporting. I suspect that the incentive structure is important, but secondary to the desire to “get results” and “find the truth” that animates most academics. Yes, lucre, fame, fortune, status are nice (well, very nice) but I agree that the main motivation for academics is the less tangible one, wanting to get results just for the sake of getting them. Being productive is a huge goal for any academic, and a big part of the lure of stats, IMO, is that it promises to get one there if one just works hard and keeps plugging away. 

So, what AG says about the curative nature of the mini-revolution rings true, but only in part. I think that the post fails to identify the three main causal spurs to stats overreach when combined with the desire to be a good productive scientist.

The first it mentions, but makes less off than perhaps others have. It is that stats are hard and interpreting them and applying them correctly takes a lot of subtlety. So much indeed that even experts often fail (see here). There is clearly something wrong with a tool that seems to insure large scale misuse. AG in fact notes this (here), but it does not play much of a role in the post cited above, though IMO it should have. What is it about stats techniques that make them so hard to get right? That I think is the real question. After all, as AG notes, it is not as if all domain find it hard to get things right. As he notes, psychometricians seem to get their stats right most of the time (as do those looking for the Higgs boson). So what is it about those domains where stats regularly fails to get things right that makes it the case that they so generally fail to get things right?  And this leads me to my second point.

Stats techniques play an outsized role in just those domains where theory is weakest. This is an old hobby horse of mine (see here for one example). Stats, especially fancy stats, induces the illusion that deep significant scientific insights are for the having if one just gets enough data points and learns to massage them correctly (and responsibly, no forking paths for me thank you very much). This conception is uncomfortable with the idea that there is no quick fix for ignorance. No amount of hard work, good ethics, or careful application suffices when we really have no idea what is going on. Why do I mention this? Because, in many of the domains where the replication crisis has been ripest are domains that are very very hard and where we really don’t have much of an understanding of what is happening. Or maybe to put this more gracefully, either the hypotheses of interest are too shallow and vague to be taken seriously (lots of social psych) or the effects of interest are the results of myriad interactions that are too hard to disentangle. In either case, stats will often provide an illusion of rigor while leading one down a forking garden path. Note, if this is right, then we have no problem seeing why psychometricians were in no need of the replication revolution. We really do have some good theory in the domains like sensory perception, and here stats have proven to be reliable and effective tools. The problem is not with stats, but with stats applied where they cannot be guided (and misapplications tamed) by significant theory.

Let me add two more codicils to this point.

First, here I part ways with AG. The post suggests that one source of the replication problem is with people having too great “an attachment to particular scientific theories or hypotheses.” But if I am right this is not the problem, at least not the problem behind the replication crisis. Being theoretically stubborn may make you wrong, but it is not clear why it makes your work shoddy. You get results you do not like and ignore them. That may or may not be bad. But with a modicum of honesty, the most stiff necked theoretician can appreciate that her/his favorite account, the one true theory, appears inconsistent with some data. I know whereof I speak, btw. The problem here, if there is one, is not generating misleading tests and non-replicable results, but if ignoring the (apparent) counter data. And this, though possibly a problem for an individual, may not be a problem for a field of inquiry as a whole. 

Second, there is a second temptation that today needs to be seriously resisted but that severely leads to replication problems: because of the ubiquity and availability of cheap “data” nowadays, the temptation to think that this time it’s different is very alluring. Big Data types often seem to think that get a large enough set of numbers, apply the right stats techniques (rinse and repeat) and out will plop The Truth. But this is wrong. Lars Syll puts it well here in a post entitled correctly “Why data is NOT enough to answer scientific questions”:

The central problem with present ‘machine learning’ and ‘big data’ hype is that so many –falsely- think that they can ge away with analyzing real-world phenomena without any (commitment to) theory. But –data never speaks for itself. Without a prior statistical set-up, there actually are no data at all to process. And – using a machine learning algorithm will only produce what you are looking for.

Clever data mining tricks are never enough to answer important scientific questions. Theory matters.

So, when one combines the fact that in many domains we have, at best, very weak theory, and that nowadays we are flooded with cheap available data the temptation to go hyper statistical can be overwhelming.

Let me put this another way. As AG notes, successful inquiry needs strong theory and careful measurement. Not the ‘and.’ Many read the ‘and’ as an ‘or’ and allow that strong theory can substitute for paucity of data or that tons of statistically curated data can substitute for virtual absence of significant theory. But this is a mistake. But a very tempting one if the alternative is having nothing much of interest or relevance to say at all. And this is what AG underplays: a central problem with stats is that it often tries to sell itself as allowing one to bypass the theory half of the conjunction. Further, because it “looks” technical and impressive (i.e. has a mathematical sheen) it leads to cargo cult science, scientific practice that looks "science" rather than being scientific. 

Note, this is not bad faith or corrupt practice (though there can be this as well). This stems from the desire to be, what AG dubs, a scientific “hero,” a disinterested searcher for the truth. The problem is not with the ambition, but the added supposition that any problem will yield to scientific inquiry if pursued conscientiously. Nope. Sorry. There are times when there is no obvious way to proceed because we have no idea how to proceed. And in these domains no matter how careful we are we are likely to find ourselves getting nowhere.

I think that there is a third source of the problem that resides in the complexity of the problems being studied. In particular, the fact that many phenomena we are interested in arise from the interaction of many causal sub-systems. When this happens there is bound to be a lot of sensitivity to the particular conditions of the experimental set up and so lots of opportunities for forking paths (i.e. p-hacking) stats (unintentional) abuse. 

Now, every domain of inquiry has this problem and needs to manage it. In the physical sciences this is done by (as Diogo once put it to me) “controlling the shit out of the experimental set up.” Physicists control for interaction effects by removing many (most) of the interfering factors. A good experiment requires creating a non-natural artificial environment in which problematic factors are managed via elimination. Diogo convinced me that one of the nice features of linguistic inquiry is that it is possible to “control the shit” out of the stimuli thereby vastly reducing noise generated by an experimental subject. At any rate, one way of getting around interaction effects problem is to manage the noise by simplifying the experimental set up and isolating the relevant causal sub-systems.

But often this cannot be done, among other reasons because we have no idea what the interacting subsystems are or how they function (think, for example, pragmatics).  Then we cannot simplify the set up and we will find that our experiments are often task dependent and very noisy. Stats offers a possible way out. In place of controlling the design of the set up the aim is to statistically manage (partial out) the noise. What seems to have been discovered (IMO, not surprisingly) is that this is very hard to do in the absence of relevant theory. You cannot control for the noise if you have no idea where it comes from or what is causing it. There is no such thing as a theory free lunch (or at least not a nutritious one). The revolution AG discusses, I believe, has rediscovered this bit of wisdom.

Let me end with an observation special to linguistics. There are parts of linguistics (syntax, large parts of phonology and morphology) where we are lucky in that the signal from the underlying mechanisms are remarkably strong in that they withstand all manner of secondary effects. Such data are, relatively speaking, very robust. So, for example, ECP or island or binding violations show few context effects. This does not mean to say that there are no effects at all of context wrt acceptability (Sprouse and Co. have shown that these do exist). But the main effect is usually easy to discern. We are lucky. Other domains of linguistic inquiry are far noisier (I mentioned pragmatics, but even large parts of semantics strike me as similar (maybe because it is hard to know where semantics ends and pragmatics begins)). I suspect that a good part of the success of linguistics can be traced to the fact that FL is largely insulated from the effects of the other cognitive subsystems it interacts with. As Jerry Fodor once observed (in his discussion of modularity), the degree to which a psych system is modular to that degree it is comprehensible. Some linguists have lucked out. But as we more and more study the interaction effects wrt language we will run into the same problems. If we are lucky, linguistic theory will help us avoid many of the pitfalls AG has noted and categorized. But there are no guarantees, sadly.



[1]I apologize for not being able to link to the original. It seems that in the post where I discussed it, I failed to link to the original and now cannot find it. It should have appeared in roughly June 2017, but I have not managed to track it down. Sorry.

Tuesday, August 21, 2018

Language and cognition and evolang

Just back from vacation and here is a starter post on, of all things, Evolang (once again).  Frans de Waal has written a short and useful piece relevant to the continuity thesis (see here). It is useful for it makes two obvious points, and it is important because de Waal is the one making them. The two points are the following:

1.     Humans are the only linguistic species.
2.     Language is not the medium of thought for there are non-verbal organisms that think.

Let me say a few words about each.

de Waal is quite categorical about the each point. He is worth quoting here so that next time you hear someone droning on about how we are just like other animals just a little more so you can whip out this quote and flog the interlocutor with it mercilessly. 

You won’t often hear me say something like this, but I consider humans the only linguistic species. We honestly have no evidence for symbolic communication, equally rich and multifunctional as ours, outside our species. (3)

That’s right: nothing does language like humans do language, not even sorta kinda. This is just a fact and those that know something about apes (and de Waal knows everything about apes) are the first to understand this. And if one is interested in Evolang then this fact must form a boundary condition on whatever speculations are on offer. Or, to put this more crudely: assuming the continuity thesis disqualifies one from participating intelligently in the Evolang discussion. Period. End of story. And, sadly, given the current state of play, this is a point worth emphasizing again and again and again and… So thanks to de Waal for making it so plainly.

This said, De Waal goes on to make a second important point: that even if it is the case that no other animals have our linguistic capacities even sorta kinda, it does not mean that some of the capacities underlying language might not be shared with other animals. In other words, the distinction between a faculty for language in the narrow versus a faculty for language in the broad sense is a very useful one (I cannot recall right now who first proposed such a distinction, but whoever it was thx!). This, of course, cheers a modern minimalist’s heart cockles, and should be another important boundary condition on any Evolang account. 

That said, de Waal’s two specific linguistic examples are of limited use, at least wrt Evolang. The first is bees and monkeys who de Waal claims use “sequences that resemble a rudimentary syntax” (3). The second “most intriguing parallel” is the “referential signaling” of vervet monkey alarm calls. I doubt that these analogous capacities will shed much light on our peculiar linguistic capacities precisely because the properties of natural language words and natural language syntax are where humans are so distinctive. Human syntax is completely unlike bee or monkey syntax and it seems pretty clear that referential signaling, though it is one use to which we put language, is not a particularly deep property of our basic words/atoms (and yes I know that words are not atoms but, well, you know…). In fact, as Chomsky has persuasively argued IMO, Referentialism (the doctrine (see here)) does a piss poor job of describing how words actually function semantically within natural language. If this is right, then the fact that we and monkeys can both engage in referential signaling will not be of much use in understanding how words came to have the basic odd properties they seem to have. 

This, of course, does not detract from de Waal’s correct two observations above. We certainly do share capacities with other animals that contribute to how FL functions and we certainly are unique in our linguistic capacities. The two cases of similarity that de Waal cites, given that they are nothing like what we do, endorses the second point in spades (which, given the ethos of the times, is always worth doing).

Onto point deux. Cognition is possible without a natural language. FoLers are already familiar with Gallistel’s countless discussions of dead reckoning, foraging, and caching behavior in various animals. This is really amazing stuff and demands cognitive powers that dwarf ours (or at least mine: e.g. I can hardly remember where I put my keys, let alone where I might have hidden 500 different delicacies time stamped, location stamped, nutrition stamped and surveillance stamped). And they seem to do this without a natural language. Indeed, the de Waal piece has the nice feature of demonstrating that smart people with strong views can agree even if they have entirely different interests. De Waal cites none other than Jerry Fodor to second his correct observation that cognition is possible without natural language. Here’s Jerry from the Language of Thought:

‘The obvious (and I should have thought sufficient) refutation of the claim that natural languages are the medium of thought is that there are non-verbal organisms that think.’ (3)

Jerry never avoided kicking a stone when doing so was all a philosophical argument needed. At any rate, here Fodor and de Waal agree. 

But I suspect that there would be more fundamental disagreements down the road. Fodor, contra de Waal, was not that enthusiastic of the idea that we can think in pictures, or at least not think in pictures fundamentally. The reason is that pictures have little propositional structure and thinking, especially any degree of fancy thinking, requires propositional structures to get going. The old Kosslyn-Pylyshyn debate over imagery went over all of this, but the main line can be summed up by one of Lila Gleitman’s bon mots: a picture is worth a thousand words, and that is the problem. Pictures may be useful aids to thinking, but only if supplied with captions to guide the thinking. In and of themselves, pictures depict too much and hence are not good vehicles for logical linkage. And if this is so (and it is) then where there is cognition there may not be natural language, but there must be a language of thought (LOT) (i.e. something with propositional structure that licenses the inferences that is characteristic of cognitive expansiveness in a given domain). 

Again this is something that warms a minimalist heart (cockles and all). Recall the problem: find the minimal syntax to link the CI system and the AP system. CI is where language of thought lives. So, like de Waal, minimalists assume that there is quite a lot of cognition independent of natural language, which is why a syntax that links to it is doing something interesting.

Truth be told, we know relatively little about LOT and it properties, a lot less that we know about the properties of natural language syntax IMO. But regardless, de Waal and Fodor are right to insist that we not mix up the two. So don’t.

Ok, that’s enough for an inaugural post vaca post. I hope your last two weeks were as enjoyable as mine and that you are ready for the exciting pedagogical times ahead.

Sunday, March 1, 2015

Fodor on PoS arguments

One of the perks of academic life is the opportunity to learn from your colleagues. This semester I am sitting in on Jeff Lidz’s graduate intro to acquisition course. And what’s the first thing I learn? That Jerry Fodor wrote a paper in 1966 on how to exploit Poverty of Stimulus reasoning that is as good as anything I’ve ever seen (including my own wonderful stuff). The paper is called “How to learn to talk: some simple ways” and it appears here. I cannot recommend it highly enough. I want to review some of the main attractions, but really you should go and get a copy. It’s really a good read.

Fodor starts with three obvious points:

1.     Speakers have information about the structural relations “within and among the sentences of that language” (105).
2.     Some of the speaker’s information must be learned (105).
3.     The child must bring to the task of language learning  “some amount of intrinsic structure” (106).[1]

None of these starting points can be controversial. (1) just says that speakers have a grammar, (2) that some aspects of this internalized grammar are acquired on the basis of exposure to instances of that grammar and (3) that “any organism that extrapolates from experience does so on the basis of principles that are not themselves supplied by its experience” (106). Of course, whether these principles are general or language-specific is an empirical question.

Given this troika of truisms, Fodor then asks how we can start investigating the “psychological process involved in the assimilation of the syntax of a first language” (106). He notes that this process requires at least three variables (106):

4.     The observations (i.e. verbalizations in its vicinity) that the child is exposed to that it uses
5.     The learning principles that the child uses to “organize and extrapolate these observations”
6.     The body of linguistic information that are the output of the application of the principles to the data that the child will subsequently use in speaking and understanding

We can, following standard convention, call (4) the PLD (Primary Linguistic Data), (5) UG, and (6) the G acquired. Thus, what we have here is the standard description of the problem as finding the right UG that can mediate PLD and G; UG(PLDL) = GL. Or as Fodor puts it (107): “the child’s data plus his intrinsic structure must jointly determine the linguistic information at which he arrives.”

And this, of course, has an obvious consequence:

…it is a conclusive disproof of any theory that about the child’s intrinsic structure to demonstrate that a device having the structure could not learn the syntax of the language on the basis of the kind of data that the child’s verbal environment provides (107).

This just is the Poverty of Stimulus argument (PoS): a tool for investigating the properties of UG by comparing the PLD to the features of the acquired G. In Fodor’s words:

…a comparison of the child’s data with a formulation of the linguistic information necessary to speak the language the child learns permits us to estimate the nature of the complexity of the child’s intrinsic structure. If the information in the child’s data approximates the linguistic information he must master, we may assume that the role of intrinsic structure is relatively insignificant. Conversely, if the linguistic information at which the child arrives is only indirectly or abstractly related to the data provided by the child’s exposure to adult speech, we shall to suppose that the child’s intrinsic structure is correspondingly complex.

I like Fodor’s use of the term ‘estimate.’ Note, so framed, the question is not whether there is intrinsic structure in the learner, but how “significant” it is (i.e. how much work it does in accounting for the acquired knowledge). And the measure of significance is the distance between the information the PLD provides and the G acquired.

Though Fodor doesn’t emphasize this, it means that all proposals regarding language learning must advert to Gs (i.e. to rules). After all, these are the end points of the process as Fodor describes it. And emphasizing this can have consequences. So, for example, observing that there are statistical regularities in the PLD is perhaps useful, but not enough. A specification of how the observed regularities lead to the rules acquired must be provided. In other words, as regularities are not themselves rules of G (though they can be the basis on which the learner infers/decides what rules G contains) accounts that stop at adumbrating the statistical regularities in a corpus are accounts that stop much too soon. In other words, pointing to such regularities (e.g. noting that a statistical learner can divide some set of linguistic data into two pieces) cannot be the final step in describing what the learner has learned. There must be a specification of the rule.

This really is an important point. Much stuff on statistical properties of corpora seem to take for granted that identifying a stats regularity in and of itself explains something. It doesn’t, at least in the case of language. Why? Because we know that there lurks rules in them there corpora. So though finding a regularity may in fact be very helpful (identifying that which the LAD looks for to help determine the properties of the relevant rules) it is not by itself enough. A specification of the route from the regularity to the rule is required or we have not described what is going on in the LAD.

Fodor then proceeds to flesh this tri-partite picture out. The first step is to describe the PLD. It consists of “a sample of the kinds of utterances fluent speakers of his language typically produce” (108). If entirely random (indeed, even if not), this “corpus” will be pretty noisy (i.e. slips of the tongue, utterances in different registers, false starts, etc.). The first task the child faces is to “discover regularities in these data (109).” This means ignoring some portions of the corpus and highlighting and grouping others. In other words, in language acquisition, the data is not given. It must be constructed.

This is not a small point. It is possible that some pre-sorting of the data can be done in the absence of the relevant grammatical categories and rules that are the ultimate targets of acquisition, but it is doubtful (at least to me) that these will get the child very far. If this is correct, then simply regimenting the data in a way useful to G acquisition will already require a specification of given (i.e. intrinsic) grammatical possibilities. Here’s Fodor on this (109):

His problem is to discover regularities in these data that, at the very least, can be relied upon to hold however much additional data is added. Characteristically the extrapolation takes the form of a construction of a theory that simultaneously marks the systematic similarities among the data at various levels of abstraction, permits the rejection of some of the observational data as unsystematic, and automatically provides a general characterization of the possible future observations. In the case of learning language, this theory is precisely the linguistic information at which the child arrives by applying his intrinsic information to the analysis of the corpus. In particular, this linguistic information is at the very least required to provide an abstract account of syntactic structure in terms of which systematically relevant features of the observed utterance can be discarded as violating the formation rules of the dialect, and in terms of which the notion “possible sentence of the language” can be defined (my emphasis, NH).

Nor is this enough. In addition we need principles that bridge from the provided data to the principles. What sorts of bridges? Fodor provides some suggestions.

…many of the assertions the child hears must be true, many of the things he hears referred to must exist, many of the questions he hears must be answerable, and many of the commands he receives must be performable. Clearly the child could not learn to talk if adults talked at random. (109-110).

Thus, in addition to a specification of the given grammatical possibilities, part of the acquisition process involves correlating linguistic input with available semantic information.

This is a very complicated process, as Fodor notes, and involves the identification of a-linguistic predicates that can be put into service absent knowledge of a G (those enjoying “epistemologically priority” in Chomsky’s sense). Things like UTAH can play a role here. For example, if we assume that the capacity for parsing a situation into agents and patients and recipients and themes and experiencers and…is a-linguistic, then this information can be used to map words (once identified) onto structures. In particular, if one assumes a mapping principle that says agents are external arguments and themes are internal arguments (so we can identify agenthood and themehood for at least a core number of initially available predicates) then hearing “Fido is biting Max” allows one to build a representation of this sentence such as [Fido [biting [Max]]].[2] This representation when coupled with what UG requires of Gs (e.g. case on DPs, agreement of unvalued phi features etc.) should allow this initial seeding to generate structures something like [TP Fido is [VP Fido [V’ biting Max]]]]. As Fodor notes, even if “Fido is biting Max” is observable, [TP Fido is [VP Fido [V’ biting Max]]]] is not. The PoS problem then is how to go from things like the first (utterances of sentences) to things like the second (phrase markers of sentences), and for this, a specification of the allowable grammatical options seems unavoidable, with these options (not themselves given in the data) being necessary for the child to organize the data in a usable form.

One of the most useful discussions in the paper begins around p. 113. Here Fodor distinguishes two problems: (i) specification of a device that given a corpus provides the correct G for that input, and (ii) specification of a device that given a corpus only attempts “to describe its input in terms of the kinds of relations that are known to be relevant to systematic linguistic description.” (i) is the project of explaining how particular Gs are acquired. (ii) is the project of adumbrating the class of possible Gs. Fodor describes (ii) as an “intermediate problem” (113) on the way to (i) (see here for some discussion of the distinction and some current ways of exploring the bridge). Why “intermediate”? Because it is reasonable to believe that restricting the class of possible Gs, (what Fodor describes as “characterizing a device that produces only non-“phony” extrapolations of corpuses” (114)) will contribute to understanding how a child settles on its specific G. As Fodor notes, there are “indefinitely many absurd hypothesis” that a given corpus is consistent with. And “whatever intrinsic structure the child brings to the language-learning situation must a least be sufficient to preclude the necessity of running through” them all.

So, one way to start addressing (i) is via (ii). Fodor also suggests another useful step (114-5):

It is worth considering the possibility that the child may bring to the language-learning situation a set of rules that takes him from the recognition of specified formal relations within and among strings in his data to specific putative characterizations of underlying structures for strings of those types. Such rules would implicitly define the space of hypotheses through which the child must search in order to arrive at the precisely correct syntactic analysis of his corpus.

So, Fodor in effect suggests two intermediate steps: a characterization of the range of possible extensions of a corpus (roughly a theory of possible Gs) and a specification of “learning rules” that specify trajectories through this space of possible Gs. This still leaves the hardest problem, how to globally order the Gs themselves (i.e. the evaluation metric). Here’s Fodor again (115):

Presumably the rules [the learning rules, NH] would have to be so formulated as to assume (1) that the number of possible analyses assigned to a given corpus is fairly small; (2) that the correct analysis (or at least any even a best analysis) is among these; (3) that the rules project no analysis that describes the corpus in terms of the sorts of phony properties already discussed, but that all the analyses exploit only relations of types that sometimes figure in adequate syntactic theories.

Fodor helpfully provides an illustration of such a learning rule (116-7). It maps non-contiguous string dependencies into underlying G rules that related these surface dependencies via a movement operation. The learning rule is roughly (1):

(1)  Given a string ‘I X J’ where the forms of I and J swing together (think be/ing in is kissing’ over a variable X (i.e. ‘kiss’), the learner assumes that this comes from a rule like (I,J) X à I X J.

This is but one example. Fodor notes as well that we might also find suggestions for such rules in the “techniques of substitution and classification traditionally employed in attempts to formulate linguistic discovery procedure[s].” As Fodor notes, this is not to endorse attempts to find discovery procedures. This project failed. Rather in the current more restricted context, these techniques might prove useful precisely because what we are expecting of them is more limited:

I am proposing…that the child may employ such relations as substitutability-in-frames to arrive at tentative classifications of elements and sequences of elements in his corpus and hence at tentative domains for the application of intrinsic rules for inducing base structures. Whether such a classification is retained or discarded would be contingent upon the relative simplicity of the entire system of which it forms a part.

In other words, these procedures might prove useful for finding surface statistical regularities in strings and mapping them to underlying G rules (drawn from the class of UG possible G rules) that generate these strings.

Fodor notes two important virtues of this way of seeing things. First, it allows the learner to take advantage of “distributional regularities” in the input to guide him to “tentative analyses that are required if he is to employ rules that project putative descriptions of underlying structure” (118). Second, these learning procedures need not be perfect, and more often than not might be wrong. The idea is that their frailties will be adequately compensated for by the intrinsic features of the learner (i.e. UG). In other words, he provides a nice vision of how techniques of statistical analysis of the corpus can be combined with principles of UG to (partially (remember, we still need an evaluation metric for a full story)) explain G acquisition.

There is lots more in the paper. It is even imbued with a certain two fold modesty.

First, Fodor’s outline starts by distinguishing two questions; a hard one (that he suggests we put aside for the moment) and an approachable one (that he outlines). The hard question is (in his words) “What sort of device would project a unique correct grammar on the basis of exposure to the data?” The approachable one is “What sort of device would project candidate grammars that are reasonably sensitive to the contents of the corpus and that operate only with the sorts of relations that are known to figure in linguistic descriptions?” (120). IMO, Fodor makes an excellent case for thinking that solving the approachable problem would be a good step towards answering the hard one. PoS arguments fit into this schema in that they allow us to plumb UG, which serves to specify the class of “non-phony” generalizations.  Add that to rules taking you from surface regularities to potential G analyses and you have the outlines of a project aimed at addressing the second question.

Fodor’s second modest moment comes when he acknowledges that his proposal “is surely incorrect as stated” (120). Here he is, IMO, the modesty is misplaced. Fodor’s proposal may be wrong in detail, but it lays out the various kinds of questions that need addressing and some mechanisms for how to do so lucidly and concisely.

As I noted, it’s great to sit in on other’s classes. It’s great to discover “unknown-to-you” classics. So, take a busman’s holiday. It’s fun.






[1] Fodor like the term ‘intrinsic’ rather than ‘innate’ for he allows for the possibility that some of these principles may themselves be learned. I would also add that ‘innate’ seems to raise red flags for some reason. As Fodor notes (as has Chomsky repeatedly) there cannot reasonably be an “innateness hypothesis.” Where there is generalization from data, there must be principles licensing this generalization. The question is not whether these given principles exist, but what they are. In this sense, everyone is a nativist.
[2] Of course, if one has a rich enough UG then this will also allow a derivation of the utterance wherein case and agreement have been discharged, but this is getting ahead of ourselves. Right now, what is relevant is that some semantic information can be very useful for acquiring syntactic structure even if, as Fodor notes, syntax is not reducible to semantics.

Monday, February 25, 2013

Fodor on Concepts


There has been a bit of a kerfuffle in the thread to What’s Chomsky Thinking Now concerning Fodor’s claim that all of our concepts are innate.  Unfortunately, with the exception of Alex Drummond, most who have participated in the discussion appear unacquainted with Fodor’s argument.  It has two parts, both of which are interesting. To help focus the indignation of his critics, I will outline them below as a public service.  Before starting however, let me share with you my own personal rule of thumb in these matters, one that I learned from Kuhn’s discussions of Aristotelian physics: when someone very smart says something that you think is obviously dumb then go back and reread it for it may be that you have thoroughly misunderstood the point. I am acquainted with Jerry Fodor. Let me assure you he is very smart. So let’s start.

As noted Fodor has a two pronged argument.  The first part (an excellent very short form of which can be found here 143ff) is an observation about learning as a form of inductive logic. Fodor distinguishes between theories of concept acquisition and theories of belief fixation. The latter is what theories of learning are about. Learning theories have nothing general to say about concept acquisition because, being inductive theories, they presuppose the availability of a set of basic concepts without which the inductive learning mechanism cannot get off the ground.  If this all sounds implausible, considering Fodor’s example will make his intent clear.

Consider someone learning a new word miv in a classical learning context.  The subject is shown cards, some of which are miv and some non-miv. The subject is given a cookie whenever s/he correctly identifies the miv cards and is hit by a bolt of lightening when s/he fails (we want the reinforcement here to be very unambiguous).  What does the subject do, according to any classical learning theory, s/he considers a hypothesis of the form “X is miv iff X is…”, the blank being filled in with a specification of the features that are criterial for being a miv.  The data is then used to assess the truth of the hypotheses with various values of “…”.  So if miv means “red and round” then the data will tend to confirm  “X is miv iff X is red and round” and disconfirm everything else. This much Fodor takes to be obvious. If learning is a form of inductive inference (and, as he notes, there is no other theory of learning), then it takes the indicated form. 

Fodor then asks where do the hypotheses that are tested come from? In other words, where do the fillers of “…” come from?  They are GIVEN. Inductive theories presuppose that the set of alternatives that the data filter are provided up front. Given a hypothesis space, the data (environmental input) can be used to assign a number (a probability) of how well that hypothesis fits the data.  What inductive theories don’t do is provide the hypothesis space.  Another way of making the same point is that what inductive logics (i.e. learning theories) do is explain how given some input the user of that logic should/does navigate the hypothesis space: where’s the best place to be given that the data has been such and so.  However, if this is what inductive logics do (and, I cannot repeat this enough, all learning theories are species of inductive logics), then the field of concepts used by the inductive logic cannot themselves be fixed by the inductive logic.  Or as Fodor puts it (147):

You have to be nativistic about the conceptual resources of the organism because the inductive theory of learning simply doesn’t tell you anything about that – it presupposes it – and the inductive theory of learning is the only one we’ve got.

So, Fodor’s argument amounts to pointing out what everybody should be nodding in agreement with: no induction without a hypothesis space.  If the inductive theory is a theory of learning, then the hypothesis space must be innate and that means that all the concepts used to define it must be innate as well.  As I said, this part of the argument is apodictic, cannot be gainsaid and, in fact, never has been. Even Quine, a rather extreme associationist, agreed that everyone is a nativist to some degree for without some nativism (enough to define the hypothesis space) there can be no induction and hence no learning. Fodor’s point is to emphasize this point and use it against theories that suggest that one can bootstrap one’s way form less conceptually complex systems of “knowledge” to more complex ones.  If this means that one can expand one’s hypothesis space by learning and ‘learning’ means induction then this is impossible.[1] 

None of this should be surprising or controversial. Controversy arises with respect to Fodor’s second prong of the argument. He takes the concepts words tag to be effectively atomic. Another way of making this point in the domain of language is that there is no lexical decomposition, or at least very very little. Why is this assumption so important? Because the relation between the input and the atomic features of the hypothesis space is causal, not inductive. You see a red thing and +red lights up. Pure transduction.  Induction proceeds given this first step: count how many of the lit features are red+round vs red+not-round, vs green+round etc.  So, for atomic features/concepts the relation between their “lighting up” and the environment is not one of learning (one doesn’t learn to light them up) it’s just a brute fact (they light up).  So, and this is an important point, to the degree that most of our words denote atomic concepts (i.e. to the degree that there is no lexical decomposition) to that degree there is no interesting inductive theory of concept acquisition. Note, this does not preclude their being a possibly interesting causal theory, e.g. maybe being exposed to a miv is causally responsible for triggering the concept miv or maybe being exposed to a dax is causally responsible, or maybe being exposed to a miv right after birth is or while being snuggled by your mother etc. The causal triggers might conceivably be very complex and finding them may be very difficult. However, with resepct to atomic features, one can only discover brute causal connections, not inductive ones. Fodor’s point is that we should not confuse them as they are very different. Recently Fodor has speculated that prototypes are causally implicated in causally triggering concepts, but he insists, rightly given his strong atomicity, that this relation is not inductive (See here).

To recap, the logic of the first argument is that primitive concepts cannot be “learned” as they are presupposed for learning to take place. This allows the possibility that one “learns” to combine these primitives in various ways and that’s what concept acquisition is.  Concept acquisition is just learning to form complex concepts. Fodor is conceptually happy with this possibility. It is logically possible that concept “acquisition” amounts to defining new concepts in terms of the primitive ones. As applied to words (which I am assuming denote concepts), it is logically possible that most/many words are complex definitions. Logically possible? Yes. Actually the case? No, or that’s what Fodor has been arguing for a very long time.

His arguments are almost always of the same form: someone proposes some complex definition for a term and he shows that it doesn’t work.  Indeed, very few linguists, psychologists or philosopher have managed to provide any but a handful of purported definitions. ‘Bachelor’ may mean unmarried man, but as Putnam noted a long time ago, there are not many words like it.

Fodor is actually in a good position to understand this point for he along with Katz once investigated reducing meanings to feature trees. David Lewis derided this “markerese” approach to semantics (another instance of be careful what you hurl as it may boomerang back at you (see Paul on Harman on Lewis here)), but what really killed it was the realization that virtually all words bottomed out in terminal faetures referring to the very concept that the featural semantics was intended to explicate. So, e.g. the markerese representation for ‘cat’ ended up having a terminal CAT. This clearly did not move explanation forward, as Fodor realized.

So is Fodor right about definitions?  Well, I am slightly less skeptical than he is about the virtues of decomposition, however, this said, I cannot find good examples showing him to be wrong. As the first part of his argument is unassailable, then those that don’t like the conclusion that ‘carburetor’ is innate (i.e. a primitive of our conceptual hypothesis space) had better start looking for ways of defining these words in terms of the available primitives.  If past history is any guide, they will fail.  Definitions in terms of sense data have come and (happily) gone and cluster concepts, once considered seriously, have long been abandoned. There is a little industry in linguistics working on argument structure in the Hale-Keyser (HK) framework, but, at least from where I sit, Fodor has drawn significant blood in his debates with HK aficionados. Suffice it for now to repeat, that this is where the action must be if Fodor is to be proven incorrect and the ball is clearly not in his court.  It is easy to show that he is wrong, viz. show that most/many words denote complex concepts.  How to show Fodor is wrong is easy. Showing that he is has proven to be far more challenging.[2]

So that’s the argument. The first step is clearly correct. All the action concerns the second.  One further point: there has been a lot of discussion in the thread that Fodor is advocating a nutty kind of nativism that eschews learning from the environment. As should be clear, this is simply false. If word learning is belief fixation then it can be as inductivist as you like. However, if word learning is concept acquisition then the question entirely revolves around the nature of the primitives concerning which everyone must take as innate and hence not acquired. Fodor’s bottom line is that hypothesis spaces are not acquired but presupposed and that as a matter of fact there is far less definition one might have supposed. That’s the argument; fire away!


[1] Alex Clark mentioned Sue Carey’s recent book that appeared to consider this bootstrapping possibility. Gallistel reviewed her book making effectively this point that induction/learning cannot expand a hypothesis space (here). To repeat, all that such theories show is how to most effectively navigate this space given certain data.
[2] One interesting avenue that Paul has been exploring revolves around Frege’s notion of definition.  For Frege definition changed a person’s cognitive powers. This is really interesting. Paul’s work starts from Jeff Horty’s discussion of Frege’s notion (here and considers how to extend it to theories of meaning more generally (c.f. here and here).