Comments

Sunday, June 15, 2014

Method and the logically possible; a reply to Alex C

Alex C says the following in a reply of mine (here). The interested should look at the whole thread. I lift the comment here for discussion as I think it highlights an important point: that if you ignore what GG has wrought over the last 60 years then your work does not address the question: what is the structure of FL?

Maybe yes and maybe no is all that I am claiming, and all that the dialectic needs. You need to have a definite no to sustain your argument.



(I think ECP effects are a good example of something that might be a reflex of some structural property of the class of grammars-- so I think this is the easiest one for you to make your argument. If you can't do it for this, then you can't do it for any).

You mistake the dialectical lay of the land. You are defending a position that both you and I consider unassailable: viz. that it is logically possible that you are barking up the right tree. Of course it is. It is also logically possible that the Lochness Monster exists, that there is no human induced global warming, that the earth is flat, and many other wonderful possibilities.  So my argument is not and never has been that it may not be possible that your methods will yield results relevant to FL. I had a much more modest claim, in fact two: (i) that as of yet it has NOT yielded any that address what I take to be the core phenomena pertinent to describing the basic architecture of FL and (ii) that there is little reason to think that they ever will lead in such directions. Let me comment on each.

Re (i), I think that here we are in agreement. I have asked you to pick out any of the effects discovered over the last 60 years and show how any of them could be explained in ways fundamentally different from those that GGers have pursued. You cite the ECP as possibly a tough nut for you to crack and as my best case.  I don’t see it this way. I think all of the examples of effects that I have cited are interesting cases. And as I never like to take advantage of others (well, sometimes, but not always), that’s why I offered you a discussion of any of the others.  But alas, till now, you have remained mum. I take this to mean that you have nothing on offer. You have claimed in past comments to have different fish to fry. I agree. And they are clearly not my fish or anything like my fish. Thus, the point stands: you have nothing on offer for the GG discoveries.  From what I can tell, however, you do offer votive candles: you hold out the promissory HOPE that some day, one day, you will get around to these questions. You counsel patience and forbearance. I grant you both, but I don’t grant you my attention or interest. Call me when you get somewhere.

Re (ii), I have no reason to think that you will get anywhere. Why do I think this? Not because it is logically impossible that you will. But because you haven’t come close yet! You see, in empirical domains being logically coherent is a very low bar, and a position that even our illustrious flat earthers and climate deniers can hurdle with ease. What one wants is not only this, but some reason for thinking that this position should be taken seriously. A position that we should allot time to developing.  This is a much higher hurdle. Here’s what it requires: a result or two bearing on the questions of interest. I think that I have been quite straightforward about what these questions of interest are. I’ve listed concrete cases galore. I’ve asked you to discuss just a handful to make your case that doing things your way gets us anywhere with these. To repeat, to date you’ve refused. Fine. So your position is logically coherent, but to date it is sterile (at least as regards what I (and GG) take to be the questions of interest). Note, this leaves little room for interesting “dialectic” as you put it. None actually. Or, the dialectic will look a lot like this. Entertaining? Yes, very. Productive? Not really.

There is a third reason that I can add to the two above: I believe that there are better ways of developing concrete models of real time learning than the ones you put on offer. I’ve cited examples of these also. People like Berwick, Wexler, Fodor, Yang, Lidz, Dresher, etc. are doing just that. Their work is responsive to the data that linguists have uncovered over the last 60 years. Their work is not merely logically coherent, it is even relevant! In short, there are better models on offer, ones that offer concrete reasons for optimism in that they address the questions of interest.


One last point: There is a kind of argument that seems to appeal to philosophers and mathematicians. It goes something like this: if you cannot show that a position is logically incoherent then it’s worth taking seriously. I never liked this argument, at least when advanced by philosophers (see here for a good example of this move). In the non-mathematical sciences, you need a lot more than this. To defend a line of research you need to demonstrate that it will get you somewhere with respect to the questions you are interested in.  In lecture 2 (which I am in the middle of) Chomsky noted that programs are not true or false. But they can be premature or sterile.  Part of the game is to show how the program you are recommending will lead towards answers to questions you are interested in. I believe that I have been very clear about what sorts of questions interest me and that emerge from the history of GG. These are questions that reflect the history of GG and its many discoveries. They suggest a research program that builds on (rather than ignores) these very non-trivial discoveries. I believe that any approach to the study of FL that does otherwise will get us nowhere. Indeed, failure to do this means that you don’t think much of these results regardless however much you deny this. Like in much else, it’s the walk not the talk that counts. So, as I keep saying, obviously raising the hackles of many, there can is no compromise on this. You need to choose!!. Either you take these results of the last 60 years to be directly relevant to current work, indeed the empirical foundations on which further work will build, or you don’t. If you don’t, then, from where I sit, there is little else for us to discuss. We are just doing different things. We may seem to be talking about the same things, but we aren’t. It’s logically possible that I am wrong about this. But I wouldn’t bet on it, and bets are what research is all about.

Thursday, June 12, 2014

Two comments

Here are two unconnected comments about previous posts. The first is a little rantish as it relates to a shibboleth I believed I had chopped to the ground.  The second is an obvious point about Chomsky’s first lecture that I stupidly forgot to remark on but is really very important.  Here they are:

1.     Ewan makes the following comment (here):

I try not to give up hope that some day linguists and psychologists will deconflate indispensable generic notions like abstractness and similarity from issues about being domain-general or domain-specific. Completely orthogonal and there isn't much hope for the field until that distinction is made. The exchange between Norbert and Alex in the first comment thread gives me a bit of optimism (because Norbert came around) but it can get tiring to have to say it again and again and again.

There is nothing to object to in the content of this remark. Who wouldn’t want these notions “deconflated”?  However, I find a lot to object to in its connotation, and in two ways.

The first is the suggestion that linguists (like moi) have finally “come around” to recognizing that these dimensions are orthogonal. This suggests that there was a time when there was widespread confusion among linguists on this matter. If so, I don’t know who was so confused and when it was.  We have long known that both domain general and domain specific acquisition mechanisms need notions of abstractness and similarity to get them to go beyond the input. There is no “learning” without biases and there are no biases without a specification of dimensions of similarity, some of them “concrete” and some of them “abstract.” Abstract, in such contexts, means going beyond dimensions afforded by a purely perceptual quality space (see Reflections on Language for a long discussion of these themes in the context of a critique of Quine).  Empiricists have long argued that all we need are such perceptual dimensions and they have been more than happy to assume that coding these dimensions is an innate part of the mind. Part of what Rationalists have argued is that this is insufficient and that more “abstract” dimensions of generalization are required. 

This noted, there is a second question: what’s the nature of the abstractness required?  Is the abstractness that is required peculiar to some domain (e.g. peculiar to visual computations or linguistic computations or auditory computations alone), or are these all aspects of one and the same kind of abstract computation.[1]  This is where the domain general/domain specific meets the abstract/non-abstract dimension. Let me explain.

Domain general explanations are, ceteris paribus, preferable to domain specific ones. Why? Because they are more general. That means that they potentially apply to a wider range of data than domain specific accounts apply to precisely because they are domain general.  And, all things being equal, this means that were they empirically adequate they would have more kinds of evidence in their favor than domain specific accounts would have. So, imagine a world (not this one from what we can tell at present) where we could unify vision and language in one set of common principles. And say that these principles could explain things like Island Effects or the ECP or Binding Effects etc. as well as the Muller-Lyer illusion or Common Fate. This would be a GREAT theory and clearly better than a purely linguistic specific explanation of Islands, the ECP etc. Moreover, it’s obvious why, but let me say it anyhow: it would be better because it explains more than the purely domain specific linguistic account does.  For the record, UG has nothing to say about the Muller-Lyer illusion. So, to repeat, were there such accounts linguists like me would rejoice and pop open the champagne, and nominate the relevant scientists for Nobel Prizes. All agree that such domain general accounts would be very nice to have, and there is a sense in which Minimalists are betting that some might be available. So as far as our druthers are concerned, we are all singing from the same hymnal.

So, given this huge agreement among all good thinking people what’s the fighting all about. Well, it’s NOT about this! Rather the above noted aspirations have never come close to realization. The problem linguists like me have with domain general explanations of linguistic phenomena of interest is that they do not currently exist!!! The reason I resist all the hype surrounding a very attractive prospect is that it seems to have relieved advocates of concretely coming up with the goods.  The reason I respect domain specific, i.e. UGish, accounts is that they are currently the only game in town (actually, I think that there are a few more general accounts of a minimalist variety, but they are rightly controversial).  In other words, I respect domain specific accounts for empirical reasons. Moreover, IMO, the big difference between those inclined to dismiss domain specific accounts and those that don’t is that the former respect the discoveries GG has made over the last 60 years and those that don’t, really don’t.  Alex C and I had a long involved “debate” about this (here) and I think that it is fair to say that we agreed that we considered different things to be key data currently in need of explanation. I take the results to 60 years of work of GG to give us a whole slew of effects that should be the targets of linguistic explanation. He is “interested in a different problem.”  And, as of this moment in time he has nothing to say about these effects and so the domain specific accounts are the only ones available. As I’ve also suggested, I think that ignoring these effects is bolstered by insisting that they do not really exist. This leads different interests to couple with a kind of skepticism about what GG has found. It is psychologically, if not logically, comforting to ignore GG results if one denies the validity of these results (e.g. see (here).

If this diagnosis is correct (and, of course, I believe it is) then the domain general vs domain specific issues have never been confused, nor have they ever hindered fruitful dialogue. The gulf between linguists and some cognitive researchers has to do with what the worthwhile problems are taken to be. GGers refuse to be diverted from their discoveries by promises of a potential theory that might one day in ways we do not yet begin to comprehend solve our problems.  As I’ve state repeatedly, give us some concrete domain general accounts and they will be carefully considered. I personally hope they exist and are produced soon. But as I learned when I was about 7, wishing things to be so doesn’t make them so. 

So, contra Ewan the problem is not that GGers like me don’t “get” the distinctions he considers vital to internalize. We get them alright. We just don’t see how they are currently relevant. Moreover, I would argue that pointing to the “failure” Ewan identifies, serves simply to throw sand in GGers eyes. It tells us to discard or demote results that we have labored hard to gain in favor of theories that don’t yet exist (not even in the faintest outline) and that address problems in ways that we think overly simplistic.  It’s cloud cuckoo land advice and serves only to mislead the interested parties as to what the real fight is about. It’s not that some prize domain general accounts and others prize domain specific ones. It’s that what GGers want to explain currently have only domain specific accounts and that those with domain general stories to peddle are trying to convince us that we should stop worrying about the facts we have found.  No thanks.

2.     On Lecture 1

One point I should have made about lecture 1 but did not was that it’s amazing that this is Chomsky’s first lecture on a series of very technical topics in minimalist grammar. It’s clear that he thinks that what follows is interesting because it is situated in a long tradition of questions about minds, brains, evolution, learning and more.  Put another way, the technical discussion that follows should be understood as addressing these far more general questions.

This is Chomsky’s standard modus operandi, and it is what makes GG so exciting IMO.  GG has always been part of the Rationalist tradition (as the lecture makes clear). Thus, the success of the GG enterprise is philosophically very telling. It is rare that large philosophical issues can be related to concrete empirical issues, albeit abstract ones, but this is one such case.  What Chomsky has always been very good at showing is how very abstract philosophical concerns are reflected in detailed empirical worries and how empirical problems commit hostages to large-scale philosophical positions. What the first lecture makes clear is that Chomsky is a modern incarnation of the 17th and 18th century natural philosophers. Big empirical issues have philosophical roots and each has implications for the other.  That’s an important insight, and nobody delivers it better than Chomsky.



[1] Gallistel and King, for example, doubt the coherence of a generalized “sensing” mechanism (as opposed to visual or olfactory sensing) and they seem to doubt the coherence of a purely domain general notion of sensing as an abstraction.

Monday, June 9, 2014

POS and PLD

What do linguists study? The obvious answer is language. GGers beg to differ. The object of study, at least if you buy into the Chomsky program in GG (which all right thinking linguists do), is not language but the faculty of language (FL). GGers treat FL realistically. It is part of the mind/brain, a chunk of psychology/biology.  IMO, perhaps Chomsky’s biggest achievement lies in having found a way of probing the fine structure of FL by studying its outputs and thinking backwards to the properties that a mind/brain with such outputs would have to have. This mode of reasoning has a name: the Poverty of Stimulus argument (POS).

The POS licenses inferences about the structure of FL from the structure of the Gs that native speakers of a natural language (NL) acquire.  As I’ve before illustrated how the POS can be pressed into service for this end, I will refrain from doing so again here. Instead, I’d like to recommend a good paper by Jeff Lidz and Annie Gagliardi (L&G) (here) that goes over these issues again with novel illustrations. Here’s some of what’s in it.

First, the paper illustrates something that’s often obscured: how POS reasoning focuses attention onto the properties of the Primary Linguistic Data (PLD). As usually structured, the argument relies on there being a gap between what can be gleaned from the data the LAD can exploit and the structure of the knowledge attained. It’s through this gap that structure of FL can be discerned. L&G nicely walk us through the logic of the gap and the inferential procedure underlying the POS.  They observe, correctly in my view, that there is nothing antithetical between GGers and a commitment to UG like principles and statistical learning procedures (a point Charles has made in his recent posts as well).  The question is not whether stats are relevant, but how they are. Here’s what L&G says.

L&G notes that there are effectively two views of how stats can enter into acquisition discussions. The first view, what L&G dubs the “input driven view,” (IDV) conceives of acquisition as a process of “track[ing] patterns of co-occurrence in the environment and stor[ing] them in a summarized format” (3). Generalizing beyond actual experience is licensed by this “summary representation of experience” (4). IDVs tend to take this learning procedure to be domain general, rather than linguistically specific.

There is a second approach, the “knowledge-drive view” (KDV). For KDV, experience functions to “select” a knowledge representation from a pre-specified (aka: innate) domain of options.

There is a long tradition distinguishing “instructive” versus “selective” theories of “learning.” This has even played a role in studies of the immune system where several people got Nobels (Jerne, Edelman) for showing that earlier instructive theories of antibody formation were wrong and that antibody formation is actually selection against a given set of possible options.

Nowadays, given the rise of Bayes, selection theories of learning are all the rage. However, there is a sense in which any statistical theory must be selective for probabilities presuppose an antecedent demarcation of possibilities over which the probabilities are assigned. In other words, we need a given hypothesis space/algebra in order to assign, e.g. probability densities. No space, no probabilities.[1] The probable selects from the possible.

One of the nice features of the L&G discussion is that it notes that the instruction/selection question is orthogonal to the empiricism/rationalism issue. The relevant question for the latter is the nature of the relation between the input and what is “selected.” If selection requires matching inputs (i.e. if acquisition is essentially a bottom up affair in which the generalizations are simply distillations of regularities found in the input) then the relevant mechanism is Empiricist as, using traditional terminology, the generalization “resembles” the input. If one allows for a “distance” between the attained generalization and the input, then one is approaching the issue from a Rationalist perspective. Rationalists are necessarily selectionists, but Empiricists need not be instructionists. The deep question lies not on this dimension, but on how close the input must “resemble” the generalized output. As L&D explains, the important question is whether the statistically massaged PLD “shapes” or “triggers” the attained generalization.[2] For the interested, this is further discussed (here).

Here’s how L&G contrasts IDV and KDV:

1.     For IDV the learning process goes from “specific to general” as the main operation is a process that “generalizes over specific cases” (4). For KDV, learners do not need to “arrive at abstract representations via a process of generalization” as these are built in. Rather input is sampled to “identify the realization of those abstract representations” in the PLD (5).
2.     For IDV the end product of learning is “a recapitulation of the inputs to the learner, [t]he acquired representation [being] a compressed memory representation of the regularities found in the input” (5). For KDV, the attained representations (or many features thereof) are innately provided so the “representation may bear no obvious relation to the input that triggered it” (5). There need be, in other words, no similarity between the attained generalizations and the specific structure of the PLD. Or, another way of putting this (which L&G does) is that whereas for IDVs the input shapes the attained representation, for KDVs it merely triggers it. 

Thus, though both IDVs and KDVs recognize that experience is critical in G attainment (i.e. to learn English you need English PLD), how this PLD functions is different in the two cases. Forming the Gish generalizations in the case of IDVs and triggering them in the case of KDVs. As such, both approaches require inferential mechanisms that take one from PLD to Gs. 

Indeed, because KDVs cannot rely on “similarity” to bridge the gap from data to generalizations as IDVs do,[3] KDVs need learning theories that show how FL/UG supplies “predictions about what the learner should expect to find in the environment” to guide acquisition which, for KVDs consists in “compar[ing] those [FL/UG provided (NH)] predicted features against the perceptual intake…[which] drives an inference about the grammatical features responsible for the sentence under consideration” (10-11).  L&G provide a nice little schema for this interactive process in their table 1.

Given these distinctions one can now begin to investigate to what degree language acquisition resembles IDVs vs KDVs.  L&G provide several interesting case studies. Let me mention two.

The first is based on Gagliardi’s thesis work on noun class acquisition in Tsez. Here’s a surprise: there are lots of them and it’s complicated.  L&G notes that classification (at least in part) depends on a noun’s semantic and phonological features. What L&G notes is that despite being sensitive to both kinds of info, LADs used them “out of proportion with their statistical reliability, even ignoring highly predictive semantic features” (24-5). L&G asks why. The discussion relies on distinguishing the input to the LAD from its intake; the former being the information available in the ambient linguistic environment, the latter being the PLD that the child actually uses.

This is an important distinction and it has been put to use elsewhere to great effect.  One of my favorites is the explanation of how Old English, which was a good Germanic OV NL, changed to VO.  David Lightfoot (here, here) asked a very interesting question about this transition: how could the change have happened? Here’s the problem: while the data in main clauses concerning the OV status of Old English (OE) was obscure and not particularly conclusive, the data that OE was OV was very clear in embedded clauses. He reasoned that if LADs had (robust) access to embedded clause PLD then there would have been no change from OV to VO in English as there was plenty of very reliable data showing that the language was OV in embedded clauses. Conclusion: the LAD did not have use embedded clause information. Given the reasonable conclusion that OE kids had the same FLs as anyone else, this means that in language acquisition kids do not “intake” embedded clause data (i.e. the PLD is degree 0+).

Lisa Pearl in her thesis work and later publications (here, here) addresses this question in terms analogous to L&G (viz. intake vs input of PLD to the OE kids). Pearl asks how much embedded data would have been enough to prevent the shift to VO order. She does this to pin down whether the LAD actively ignores embedded data (an intake restriction) or whether embedded data info is absent because complex clauses with embedding are just rare in the PLD (an input restriction). Lisa was able to “quantify” the question and the data she harvested from CHILDES (current English data not OE, of course) suggested that kids must actively ignore available data if we are to account for the English shift from OV to VO. Note, there could have been another answer to Lightfoot’s question: degree 1 sentences are too rare in the input for their information to be effectively usable, i.e. the restricted intake is due to restricted input. However, both Lisa and L&G show apparent cases where this answer is insufficient. It appears that the LAD might be deliberately myopic. And this raises another interesting question: WHY?

L&G discuss a second case that I would like to bring to your attention; acquisition of Double Object Constructions (DOC) in Kannada. The facts are once again subtle, though recognizably similar to what we find in languages like English and Spanish. At any rate, it is very complex, with binding allowed between pronouns in the accusative and dative DPs in some conditions but not in others. L&G notes that the data present a standard POS puzzle and show how it might be addressed. What was most interesting to me is the discussion of possible triggers and how exactly FL/UG might focus attention on some kinds of available information (spoiler alert: it has to do with animacy in possession constructions and their relation to DOCs), which would function as triggers. L&G notes that this kind of info is available with some statistical regularity if you are primed to look for it. In other words, the abstract POS considerations immediately generate a search for triggers that FL/UG might make salient and thus be in-takeable (i.e perceptually prominent data that would permit the triggering).  This is a very nice illustration of the fecundity of POS considerations in considering how real time acquisition works. It also nicely illustrates how the gaps that POS arguments identify can be bridged via data that is in no way “similar” to the representations attained.

Let me end. L&G is a fun paper. And there’s lots of other good stuff (I really liked the discussion of Korean and how POS allows a kind of data free parameter setting). It gives a good overview of the intuitions behind the competing acquisition traditions, makes some nice distinctions concerning how to think of PLD in the wild, and provides a budget of nice novel illustrations of POS arguments. A colleague (Colin Phillips) keeps insisting that he is tired of the focus on Y/N questions in English as the “hard” case of POS reasoning. I am not nearly as tired as he is, but we agree on one thing: POS arguments are thick on the ground once one looks at any even slightly complex bit of grammatical competence. There is nothing special about Chomsky’s original illustration, except perhaps its extreme accessibility. L&G provides some good new grist for those interested in this kind of milling and shows how useful such reasoning is in attacking the question of how real kids (not just ideal speaker-hearer LADs) manage to acquire Gs.



[1] Perfors makes essentially this point (here).
[2] For the interested, this is further discussed (here).

[3] This is not to agree that ‘similarity’ is a serviceable notion. It’s not, as philosopher’s (e.g. Nelson Goodman) have repeatedly shown.

Sunday, June 8, 2014

Lecture 1: Comments

Here are some comments on lecture 1 (here). I’ll try to comment on the others as I get through them sometime in the next couple of weeks.

The aim of the first lecture is to locate the Generative enterprise conceptually. Chomsky notes there have been two lasting themes since the inception: (i) that the central fact about Natural Language is that it involves the generation of an infinite array of hierarchically structured objects that link to systems of thought and systems of externalization; The aim of GG being to describe this I-language in detail. (ii) Language is a biological system, not a social construct. It should be studied the way other such biological systems are studied; the aim being to figure out the underlying structure of its phenotypical properties.

Given these two themes there are several obvious projects: (a) study the structure of I-langauge, (b) study how I-language is acquired (c) study how I-language emerged in the species.  None of this is new or exciting to readers of this blog, I would hope. What is fun to see is how Chomsky understands these projects in a wider philosophical and historical context.  Chomsky is very very good at giving a broad sweep history of earlier influences and differences that the modern perspective has adopted. He’s very good at comparing his views to those of these important precursors.  It’s actually amazing how many giant’s shoulders are trotted out for perching on (Darwin, Descartes, Newton, Leibniz, Galileo, Locke, Hume, von Humboldt, Russell, Turing, Church, Kleene) and how many views are dumped on (Dummett, Quine, Lewis, Tomasello, Construction Grammarians, Churchlands). 

I especially enjoyed Chomsky’s discussion of the history of the calculus and the early history of Chemistry. He makes a point that he has made before but is worth repeating in the current cultural climate. The point is that methodological pronouncements often lead us astray if the history of science is any indication. Newton’s calculus had real foundational problems, ones that Berkeley identified. It seems that parts of the system were based on equivocations, which vitiated many of the proofs. British mathematicians took Berkeley’s arguments very seriously with the result that they contributed almost nothing to the next steps in the development of the calculus. Continentals basically ignored these problems and made fundamental contributions to its development.  When were these problems resolved? At the end of the 19th century when solving them really were required for the problem with infinitesimals started impeding mathematical progress.  So are fuzzy equivocal concepts always bad and resolving them always good? Not if history of science is the guide. Sometimes the fuzziness should be tolerated for too much conceptual fussiness has its costs.

I cannot help but think that this has relevance for some issues that we have debated on this blog concerning the utility of formalization in current linguistics. There are some that find the concepts too fuzzy to be born. Others think them clear enough, while conceding problems that will be cleared up when the need arises. We all know who we are. What Chomsky notes is that history on these matters doesn’t always (or even usually) come down on the side of methodological hygiene. Or, really, being careful matters more at some times than at others and one needs to show that a “confusion” is impeding progress before one insists on stopping research that ignores it in its tracks.

I also loved Chomsky’s discussion of the history of the “reduction” of Chemistry to Physics. There was none. The prestige science, physics, never succeeded in explaining chemistry. Rather physics had to change radically before it had anything to say about chemistry, whose methods remained effectively unchanged.  Chomsky, notes that there is every reason to think that this is the same now in domains of relevance to linguists. Think the reduction of the mental to the neural. There is a presupposition that the soft mental sciences must adjust their findings to fit in with the hard brain sciences. The history of chemistry suggests, however, that this reading, even if correct, is tendentious.  Error can come anywhere and when there is a problem it is never obvious what theory requires adjustment.  Moreover, at least in the case of minds and brains, Chomsky notes that there are reasons for thinking that the brain people are barking up the wrong trees. He notes Gallistel&King’s work suggesting the neuro types have gotten hold of the wrong end of the stick. Readers of this blog will recognize that I could not agree more.

There are many other excellent riffs in this two hour segment. Chomsky does an excellent job of identifying intellectual precursors and outlining where he thinks they got things right and where wrong.  He arranges his discussion of different conceptions of language around Darwin, Descartes and von Humboldt, noting how each could be seen on focusing attention in I-language as the proper object of study. Darwin by understanding that language was a biological system instrumental in undergirding the distinctive nature of human (vs animal) cognition, Descartes by seeing the creativity of human language as a completely distinctive kind of natural phenomenon calling for a novel kind of scientific approach, von Humboldt by zeroing in on the recursive property of language as the thing that needed explaining. Chomsky also notes that their conceptions required some cleaning up to get to the ones that we now take as fundamental: that what is studyable is linguistic competence, not linguistic behavior, that language really is qualitatively different form what we see in other parts of the biological and physical world, and that there are some aspects of language (and cognition) that may forever stay shrouded in mystery.

Chomsky also sets the stage for his coming minimalist disquisition. He does this in two ways. First, he reviews the logic that drives the search for a few very simple principles that allow for the emergence of FL. He notes that the capacity for language is very recent, and has been stable since its emergence. This implies that whatever prompted its emergence is very simple and very few in number; hopefully a single addition. Second, he notes that the target of explanation is the recursive systems that produces an infinite array of hierarchical expressions that ling to systems of externalization and meaning, the latter being more basic (hour 1;14). An aside: one interesting point Chomsky makes given recent discussions in the comments section here, is his observation that we know next to nothing about the system of thought and its objects (hour 1;15). That would appear to make the role of Bare Output Conditions marginal in what follows, but that’s just a hunch so stay tuned.



There’s lots more: comments on the roles of reference as a semantic primitive (nope!), on how far we should expect to be able to understand the world and our theories of them (the first, not at all, the second up to a point), the importance of being puzzled, the hollowness of most current discussion on the evolution of language and much more. The lecture is a little like some of Dylan’s concert albums with the Band. The songs, though familiar, are all played slightly differently, so they are familiar without being boring and part of the fun is humming along. Sort of high class comfort food. Chomsky has a knack for making the big picture accessible in a way that nobody else can. This is a fun 2 hours. I hope that once he gets technical, it stays as scintillating. I suspect however that come lecture 2 it’ll be time to fasten one’s seatbelt and put all trays into an upright and locked position.

Addendum:

Chomsky makes one point that is relevant to some prior discussions on this blog. He points out that assuming that FL is a real biological object licenses using a very wide array of data to triangulate on its properties. Dat from Japanese bears on the structure of English Gs, and data from acquisition bears on the structure of English Gs and data from processing bears on the structure of English Gs and…In other words, once one sees our theories as theories OF FL then there is (at least in principle) no way of a priori delimiting what data will count as useful in probing its structure. This, Chomsky notes, contrasts with earlier structuralist conceptions wherein the aim was simply a way of effectively organizing a set of data or corpus. Here, anything beyond such data is irrelevant. The same holds if one considers the object of study instrumentally or Platonistically. Such views serve to a priori blinker investigation by dismissing empirical considerations that go  beyond the particular <s,m> pairs before one's nose. In other words, understanding theories to be about FL/UG realistically construed widens the scope of one's inquiries.