Even though I have recently moved on from the bed of nails that is the current job market to a cushy tenure track job, I still find myself reading the job announcements on LinguistList on a daily basis. There's of course all kinds of professional reasons for doing so, but the actual driving force behind this minor obsession of mine is more twisted. For you see, I have an existentialist streak that allows me to derive perverse amounts of joy from things that should cause me grief, worry, pain, and outbreaks of homicidal rage. And job searches for computational linguists got all of that aplenty.
Sunday, November 10, 2013
Friday, November 8, 2013
Celebrity MOOCs
Brad Larson sent me this from Slate. It's site amusing actually. As we all know, teaching is part theater, so why not follow this to its obvious conclusion? I should add that once we are going in this direction, there is no reason that just any actors should be recruited. We know looks matter (as a peak at news anchors or leading people in films suggests) as do other things (background music, type of dress (or not), sexiness factor, etc.). I am beginning to like the way that this is going. Once we stop dropping academic categories and start thinking really creatively I bet we can soon make all of education much more fun. No more onerous homework, no more incomprehensible classes, no more ugly profs. And this year, the nominees for best Graduate Syntax class are...
Monday, November 4, 2013
Help
I know that there are many out there that know how Blogger works. So I ask for some help. I embedded a graph of a learning curve in the last post. I copied this from the R&A paper and put it in the post. I then checked to see if it stuck in previewing the post and it did. But I go there now and the insert is gone, replaced by a cute blue box with a '?' or nothing at all. So, can anyone walk me through how to insert a graph that I've copied from someplace else and insert it so that it can be viewed in the blog? Thx.
Learning is to cognition what phlogiston is to chemistry
Last week John Trueswell gave a colloquium talk at UMD that
I unfortunately could not attend. I was in Montreal at a workshop in honor of
an old prof of mine, Jim McGilvray. The
Montreal gig was great and it was a pleasure to be able to fete Jim in person
(he supervised one of my earliest linguistics projects, (my undergrad thesis on
a Reichenbachian theory of tense), but I confess that I would have loved to
have been at John’s talk as well. I await the day when some clever physicist
figures out how to allow someone to be in two places at once. Until that happy
time arrives, I thought it appropriate to do some penance for my physical
failings by re-reading a
terrific paper by Roediger and Arnold (R&A) on the history of one-trial
learning experiments. Why the R&A paper? Because, John’s recent work on
lexical acquisition is in the one-trial learning tradition whose history they
review. I’ve discussed this work already in a couple of places (here
and here)
but I wanted to bring your attention to the R&A paper again for it
highlights some of the more interesting implications of this line of research
for topics near and dear to my intellectual prejudices: despite the common
conviction that Empiricism has (at least) something going for it, there is a
remarkable absence of evidence
supporting this very weak view.[1]
Here’s what I mean.
Behind every theory there is an inspirational picture. In
the mental sciences, the two grand traditions, Rationalism (R) and Empiricism
(E), are animated by two contrasting conceptions of the underlying mechanisms
of mental life. For Es, the afflatus is the blank wax tablet, learning
consisting of imprinting by experience on this tablet and the clarity and distinctness
of the resultant concept/idea being a function of the number of repeated imprints.
The more the experience, the deeper and clearer the resulting acquired concept/idea. Here’s Ebbinghause’s
version (quoted in R&A: 129):
These relations [between repetition
and performance] can be described figuratively by speaking of the series as
being more or less deeply engraved on some mental substratum. To carry out this
figure: as the number of repetitions increases, the series are engraved more
and more deeply and indelibly; if the number of repetitions is small, the
inscription is but surface deep and only fleeting glimpses of the tracery can
be caught…
Note that on this conception, repeated experience forms the
concepts in the mind (e.g. makes the grooves). Repetition is critical for the
mind’s main character is its receptivity to external formative forces, the mind itself being structurally rudimentary. On
this view, to understand acquisition requires analyzing the fine structure of
the input for what minds/brains do in forming mental constructs (ideas,
concepts, etc.) is sift and manipulate these input experiences. It is not
surprising, that this view focuses on minds’ significant statistical capacities
for these are obvious candidate mechanisms for organizing the inputs and
separating the significant wheat from the non-significant chaff.
This contrasts with Rish proposals. For these, the mind is
very articulated. There is lots of given pre-experiential structure. Thus, the
role of experience is not to construct the relevant concepts attained but to
kick start them into activation. Experience on this view is a trigger, not an artificer. Not
surprisingly, this conception focuses on discovering the natively provided
mental structures that experience serves (importantly but modestly) to activate.
On considering these two different raw philosophical pictures,
one can understand the intrinsic interest in one-trial learning (OTL). The
existence of OTL would be a problem for E but not for R. The empirical question
then is whether OTL exists and how common it is. Investigating this requires
translating the philosophical pictures into testable theories, and this leads
to learning curves.
R&A observe that one of the biggest pieces of evidence
for the E view of the world is the classical learning curve; you know the one
that rises from low left to high right decelerating as it goes (as below
reproduced from R&A p. 128).
R&A note that this curve perfectly embodies the E
conception that the Ebbinghause quote poetically describes. R&A point out
two important features of learning curves consonant with the leading E idea.[2]
First, “[t]he fact that the learning curve shows a gradual increase in
performance is a reflection of the underlying mechanism- the build up of
strength- which is itself also gradual.” And second, that this curve is “the
same across astonishingly different experimental situations and dependent
measures, as well as across species from slugs to humans,” strongly suggesting
general “underlying mechanisms” and general “laws of learning” (129). In a
word, this curve, it is argued, puts paid to the R idea of triggering and its
concomitant conception of a highly structured mind. If learning curves describe
the mechanics of learning, then E beats R. [3]
Unless
this curve is actually an artifact of, e.g. how experimental data is crunched, rather
than a description of an underlying mental mechanism. And that’s where the story
that R&A tell gets really interesting.
In the late 1950s and early 1960s Irvin Rock and William
Estes (these two were big psych shots, look them up) did a series of
experiments that showed (at the very least) that this interpretation of the
curve as describing an underlying mechanism that is similarly smooth and
incremental is premature, and (at the most) that it was false. They showed two
things: (i) that these curves were both consistent with an underlying OTL
mechanism (i.e. “although the learning curves were continuous, the underlying
processes were anything but continuous” (p. 129)) and (ii) that there was very
good evidence that OTL is the norm. Let’s discuss each point separately.
Here’s R&A quoting Rock (that’s the “p.186” below) and
then commenting (p. 130) wrt (i):
“Another possibility is that
repetition is essential because only a limited number of associations can be
formed in one trial, and improvement with repetition is only an artifact of
working with long lists of items.” (p.186 ). … That subset is learned
perfectly, but all the rest of the associations that were presented are not learned at all….The “artifact” Rock
referred to is essentially that of averaging across many subjects learning many
lists on many trials: despite the
all-or-none nature of the underlying process, the learning curve will be smooth
when performance is averaged over these several parameters. (my
emphasis;NH)
This is a very important conceptual point for it divorces
the big E conclusion that the mechanisms of learning are gradual and driven by
repeated environmental inputs (viz. repeated engravings by experience on a
mental substratum) from the fact that learning curves have the shape they do
and can be found quite generally across tasks and species. Put more pointedly,
if this is correct, then the smooth shape of the learning curve implies nothing at all about the smoothness and
gradualness of the underlying mechanism.
Rock and Estes not only made this important observation but also
then went on to show that in classical cases of “learning”[4]
(i.e. acquiring paired associates), there is good evidence against the
classical picture. The basic set up was to have two groups, one that learned
listed pairs by going through one list again and again (the control group) and
a second that learns a list that removes the non-learned pairs so that they are
not encountered again. The prediction if E is correct is that the control group
will do better than the second group. I will not review Rock’s and Este’s experiments
here as that’s what the R&A paper does so well. Suffice it to say, that, as
R&A put it, their papers showed that “there was no hint for the
continuity/incremental hypothesis in the data” (p.131).
Note that if (ii) is correct, an account that uses an
incremental mechanism to derive
acquisition data correctly described
by the classical learning curve is incorrect.
Let me beat this horse good and dead: Point (i) shows that a classical
learning curve is consistent with OTL mechanisms. Point (ii) argues that OTL
mechanisms are in fact what we find. Hence,
if correct, theories that deploy incremental mechanisms even if they can derive
classical learning curves are wrong. This should not be surprising: such
curves graph a correlation between trials and responses. The aim of a theory is
not to “model” the data but to “model” the mechanisms that generate the data.
The E mistake, is to wrongly infer that smooth incremental data implies smooth
incremental mechanisms. It doesn’t, though thinking it does is an Eish
diathesis. [5]
It goes without saying that both Rock’s and Estes’ results
were contested. Methodological problems were purportedly found that confounded
the conclusions. However, and this is interesting, no good evidence for the classical theory appeared to be
forthcoming. Rather, the critics seemed satisfied with a draw, viz. showing
that the Rock/Estes results need not be interpreted as debunking the classical
E view. One particularly cute study that R&A report seemed satisfied with
the conclusion “that the incremental theory is untestable” or, quoting the
critics directly (Underwood and Keppel): “certain theories are not capable of
disproof. Certain aspects of the incremental theory seem to be of this nature”
(p. 135). If this be a vindication of the classical E view imagine what a
refutation would look like.
John’s current work on lexical acquisition develops the
Rock/Estes conception. This is what makes it so interesting for people like me.
If they are correct, then classical E conceptions of learning don’t exist, or,
more modestly, there is precious little evidence in its favor. To me, this has
enormous implications for standard approaches to modeling acquisition that
assume some form of gradual process taking place, some form of gradual hill
climbing or gradual strengthening of connections. Indeed, in some moods (e.g.
now), I think that this work, if correct, implies that Gallistel and Matzel are
correct and there is no general theory of learning to be had, as there are no mechanisms, mental or neural, that
correspond to what the E picture took learning to be.[6]
However, for now I would be happy with more modest conclusions: (i) that there
is precious little evidence in favor of the E conception, (ii) that there is
little evidence in favor of the view that mental mechanisms are gradual and
continuous, and (iii) that there is pretty good evidence that we have mechanisms
that enable what amounts to one trial “learning” and that this kind of
acquisition requires something very much like the classical R conception of the
mind.[7]
There is no good reason, in other words, for taking the E conception to be the
default and every reason to think that the problem of acquisition is largely
one of getting the pre-packaged representational formats correct.
Let me end with a request and an exhortation. First the
request: does anyone have a poster case of learning not susceptible to the
Rock/Estes critique from the psych literature. It would be nice to have one.
Second read the R&A papers and the recent papers developing these ideas by
John and Lila and Charles and Jesse and their students. If their insights are
internalized, we may finally be able to break the grip of E conceptions of
mental mechanisms as the default position. One, at least, can always hope;
after all we got rid of phlogiston, didn’t we?
[1]
I think it was Lila Gleitman who first advanced the following PoS argument:
Empiricism must be innate for what else could explain the widespread conviction
that it is true despite the dearth of evidence in its favor. Lila also was kind
enough to bring the R&A paper to my attention. I should also add that I
doubt that Lila would endorse my interpretation of this work as outlined below.
In fact, I am sure she wouldn’t given this.
[2]
Gallistel and Matzel (see here)
note that the LTP view of brains has been taken to similarly support an E
picture. G&M argue that the E picture is widely accepted in the neurosciences,
despite there being little to recommend it (and a lot to disavow it). R&A’s
discussion merges well with G&M’s and supports a similar conclusion.
[3]
R&A identify, in passing, an attraction of E views of learning to the
formally inclined that comes from the mathematical tractability of learning
curves. I quote: “Learning curves (like forgetting curves) are smooth and
beautiful, and psychologists with a mathematical bent can have a field day
fitting equations to them” (p.128). One
should never underestimate the attraction of a conclusion that fits snugly with
your available technology.
[4]
Note the scare quotes: if Rock and Estes (and Gleitman and Trueswell and
Company are right) then learning is a hypothesis
about the mechanisms of acquisition,
not a neutral description of an observed phenomenon. What we observe is change
over time given environmental inputs. This change may be due to learning,
maturation, growth, or whatever. The cognitive question concerns the mechanism
and learning is a proposal for one such.
[5]
This is the kind of mistake that modeling which takes the name of the game to
be getting the input/output relations right is particularly susceptible to. E
conceptions are prone to this kind of misconception given their picture that
mental structure mirrors the structure of the input. However, modeling I/O
relations confuses the data to be explained for the mechanism that does the
explanation. For a related point in a Bayesian context see Glymour on
“Osiander’s Psychology” in the comments to Jones & Love’s discussion of
modern Bayesianism here
and some blogish discussion by me (here).
J&L note a similarity between earlier behaviorist conceptions and some
modern Bayesian analyses. The above suggests how two programs that appear so
different on the surface might nonetheless lead to the same conceptual place
via a shared partiality to associationism and/or a misunderstanding of what modeling
is supposed to do.
[6]
After all, if (roughly) one trial “learning” is the norm then there is little
for E like mechanisms to do. Acquisition on this view is more akin to transduction
than to mental computation. I would probably be satisfied if it turned out that
there was “a little” learning, but this is a topic for another discussion.
[7]
After all if we don’t acquire knowledge by carefully sifting through the input
it’s because such sifting is not necessary and it wouldn’t be necessary if the
knowledge is basically, already, all there.
Thursday, October 31, 2013
Lifted from the comments on 'Mother knows best.'
There has been an intense discussion of the APL paper in the comments to my very negative review of this paper (here). I would like to highlight one interchange that I think points to ways that UG like approaches have combined with (distributional) learners to provide explicit analyses of child language data. In other words, there exists concrete proposals dealing with specific cases that fruitfully combine UG with explicit learning models. APL deals with none of this kind of material, despite its obvious relevance to its central claims and despite citing papers that deal with such work in its bibliography. It makes one (e.g. me) think that maybe the authors didn't read (or understand) the papers APL cites.
At any rate, here is a remark by Alex C that generated this informative reply (i.e. citations for relevant work) by Jeff Lidz.
Alex Clark:
So consider this quote: (not from the paper under discussion)
"It is standardly held that having a highly restricted hypothesis space makes
it possible for such a learning mechanism to successfully acquire a grammar that is compatible with the learner’s experience and that without such restrictions, learning would be impossible (Chomsky 1975, Pinker 1984, Jackendoff 2002). In many respects, however, it has remained a promissory note to show how having a well-defined initial hypothesis space makes grammar induction possible in a way that not having an initial hypothesis space does not (see Wexler 1990 and Hyams 1994 for highly relevant discussion).
The failure to cash in this promissory note has led, in my view, to broad
skepticism outside of generative linguistics of the benefit of a constrained initial hypothesis space."
This seems a reasonable point to me, and more or less the same one that is made in this paper: namely that the proposed UG don't actually solve the learnability problem.
Jeff Lidz:
Alex C (Oct 25, 3am ([i.e. above] NH])) gives a quote from a different paper to say that APL have identified a real problem and that UG doesn't solve learnability problems.
The odd thing, however, is that this quote comes from a paper that attempts to cash in on that promissory note, showing in specific cases what the benefit of UG would be. Here are some relevant examples.
Sneed's 2007 dissertation examines the acquisition of bare plurals in English. Bare plural subjects in English are ambiguous between a generic and an existential interpretation. However, in speech to children they are uniformly generic. Nonetheless, Sneed shows that by age 4, English learners can access both interpretations. She argues that if something like Diesing's analysis of how these interpretation arise is both true and innate, then the learner's task is simply to identify which DPs are Heim-style indefinites and the rest will follow. She then provides a distributional analysis of speech to children does just that. The critical thing is that the link between the distributional evidence that a DP is indefinite and the availability of existential interpretations in subject position can be established only if there is an innate link between these two facts. The data themselves simply do not provide that link. Hence, this work successfully combines a UG theory with distributional analysis to show how learners acquire properties of their language that are not evident in their environment.
Viau and Lidz (2011, which appeared in Language and oddly enough is cited by APL for something else) argues that UG provides two types of ditransitive construction, but that the surface evidence for which is which is highly variable cross-linguistically. Consequently, there is no simple surface trigger which can tell the learner which strings go with which structures. Moreover, they show that 4-year-olds have knowledge of complex binding facts which follow from this analysis, despite the relevant sentences never occurring in their input. However, they also show what kind of distributional analysis would allow learners to assign strings to the appropriate category, from which the binding facts would follow. Here again, there is a UG account of children's knowledge paired with an analysis of how UG makes the input informative.
Takahashi's 2008 UMd dissertation shows that 18month old infants can use surface distributional cues to phrase structure to acquire basic constituent structure in an artificial language. She shows also that having learned this constituent structure, the infants also know that constituents can move but nonconstituents cannot move, even if there was no movement in the familiarization language. Hence, if one consequence of UG is that only constituents can move, these facts are explained. Distributional analysis by itself can't do this.
Misha Becker has a series of papers on the acquisition of raising/control, showing that a distributional analysis over the kinds of subjects that can occur with verbs taking infinitival complements could successfully partition the verbs into two classes. However, the full range of facts that distinguish raising/control do not follow from the existence of two classes. For this, you need UG to provide a distinction.
In all of these cases, UG makes the input informative by allowing the learner to know what evidence to look for in trying to identify abstract structure. In all of the cases mentioned here, the distributional evidence is informative only insofar as it is paired with a theory of what that evidence is informative about. Without that, the evidence could not license the complex knowledge that children have.
It is true that APL is a piece of shoddy scholarship and shoddly linguistics. But, it is right to bring to the fore the question of how UG makes contact with data to drive learning. And you don't have to hate UG to think that this is valuable question to ask.
One more point: Jeff observes that "it is right to bring to the fore the question of how UG makes contact with data to drive learning," and that APL are right to raise this question. I would agree that the question is worth raising and worth investigating. What I would deny is that APL contributes to advancing this question in any way whatsoever. There is a tendency in academia to think that all work should be treated with respect and politesse. I disagree. All researchers should be so treated, not their work. Junk exists (APL is my existence proof were one needed) and identifying junk as junk is an important part of the critical/evaluative process. Trying to find the grain of trivial truth in a morass of bad argument and incoherent thinking retards progress. Perhaps the only positive end that APL might serve is to be a useful compendium of junk work that neophytes can go to to practice their critical skills. I plan to use it with my students for just this purpose in the future.
At any rate, here is a remark by Alex C that generated this informative reply (i.e. citations for relevant work) by Jeff Lidz.
Alex Clark:
So consider this quote: (not from the paper under discussion)
"It is standardly held that having a highly restricted hypothesis space makes
it possible for such a learning mechanism to successfully acquire a grammar that is compatible with the learner’s experience and that without such restrictions, learning would be impossible (Chomsky 1975, Pinker 1984, Jackendoff 2002). In many respects, however, it has remained a promissory note to show how having a well-defined initial hypothesis space makes grammar induction possible in a way that not having an initial hypothesis space does not (see Wexler 1990 and Hyams 1994 for highly relevant discussion).
The failure to cash in this promissory note has led, in my view, to broad
skepticism outside of generative linguistics of the benefit of a constrained initial hypothesis space."
This seems a reasonable point to me, and more or less the same one that is made in this paper: namely that the proposed UG don't actually solve the learnability problem.
Jeff Lidz:
Alex C (Oct 25, 3am ([i.e. above] NH])) gives a quote from a different paper to say that APL have identified a real problem and that UG doesn't solve learnability problems.
The odd thing, however, is that this quote comes from a paper that attempts to cash in on that promissory note, showing in specific cases what the benefit of UG would be. Here are some relevant examples.
Sneed's 2007 dissertation examines the acquisition of bare plurals in English. Bare plural subjects in English are ambiguous between a generic and an existential interpretation. However, in speech to children they are uniformly generic. Nonetheless, Sneed shows that by age 4, English learners can access both interpretations. She argues that if something like Diesing's analysis of how these interpretation arise is both true and innate, then the learner's task is simply to identify which DPs are Heim-style indefinites and the rest will follow. She then provides a distributional analysis of speech to children does just that. The critical thing is that the link between the distributional evidence that a DP is indefinite and the availability of existential interpretations in subject position can be established only if there is an innate link between these two facts. The data themselves simply do not provide that link. Hence, this work successfully combines a UG theory with distributional analysis to show how learners acquire properties of their language that are not evident in their environment.
Viau and Lidz (2011, which appeared in Language and oddly enough is cited by APL for something else) argues that UG provides two types of ditransitive construction, but that the surface evidence for which is which is highly variable cross-linguistically. Consequently, there is no simple surface trigger which can tell the learner which strings go with which structures. Moreover, they show that 4-year-olds have knowledge of complex binding facts which follow from this analysis, despite the relevant sentences never occurring in their input. However, they also show what kind of distributional analysis would allow learners to assign strings to the appropriate category, from which the binding facts would follow. Here again, there is a UG account of children's knowledge paired with an analysis of how UG makes the input informative.
Takahashi's 2008 UMd dissertation shows that 18month old infants can use surface distributional cues to phrase structure to acquire basic constituent structure in an artificial language. She shows also that having learned this constituent structure, the infants also know that constituents can move but nonconstituents cannot move, even if there was no movement in the familiarization language. Hence, if one consequence of UG is that only constituents can move, these facts are explained. Distributional analysis by itself can't do this.
Misha Becker has a series of papers on the acquisition of raising/control, showing that a distributional analysis over the kinds of subjects that can occur with verbs taking infinitival complements could successfully partition the verbs into two classes. However, the full range of facts that distinguish raising/control do not follow from the existence of two classes. For this, you need UG to provide a distinction.
In all of these cases, UG makes the input informative by allowing the learner to know what evidence to look for in trying to identify abstract structure. In all of the cases mentioned here, the distributional evidence is informative only insofar as it is paired with a theory of what that evidence is informative about. Without that, the evidence could not license the complex knowledge that children have.
It is true that APL is a piece of shoddy scholarship and shoddly linguistics. But, it is right to bring to the fore the question of how UG makes contact with data to drive learning. And you don't have to hate UG to think that this is valuable question to ask.
Wednesday, October 30, 2013
Some AI history
This recent piece (by James Somers) on Douglas Hofstadter (DH) has a brief review of AI from it's heady days (when the aim was to understand human intelligence) to its lucrative days (when the goal shifted to cashing out big time). I have not personally been a big fan of DH's early stuff for I thought (and wrote here) that early AI, the one with cognitive ambitions, had problems identifying the right problems for analysis and that it massively oversold what it could do. However, in retrospect, I am sorry that it faded from the scene, for though there was a lot of hype, the ambitions were commendable and scientifically interesting. Indeed, lots of good work came out of this tradition. Marr and Ullman were members of the AI lab at MIT, as were Marcus and Berwick. At any rate, Somers gives a short history of the decline of this tradition.
The big drop in prestige occurred, Somers notes, in about the early 1980s. By then "AI…started to…mutate…into a subfield of software engineering, driven by applications…[the]mainstream had embraced a new imperative: to make machines perform in any possible, with little regard for psychological plausibility (p. 3)." The turn from cognition was ensconced in the conviction that "AI started working when it ditched humans as a model, because it ditched them (p. 4)." Machine translation became the poster child for how AI should be conducted. Somers gives a fascinating thumb nail sketch of the early system (called 'Candide' and developed by IBM) whose claim to fame was that it found a way to "avoid grappling with the brain's complexity" when it came to translation. The secret sauce according to Somers? Machine translation! This process deliberately avoids worrying about anything like the structures that languages deploy or the competence that humans must have to deploy it. It builds on the discovery that "almost doesn't work: a machine…that randomly spits out French words for English words" can be tweaked "using millions of pairs of sentences…[to] gradually calibrate your machine, to the point where you'll be able to enter a sentence whose translation you don't know and get a reasonable result…[all without] ever need[ing] to know why the nobs should be twisted this way or that (p. 10-11).
For this all to work requires "data, data, data" (as Norvig is quoted as saying). Take …"simple machine learning algorithms" plus 10 billion training examples [and] it all starts to work. Data trumps everything" Josh Estelle at Google is quoted as noting (p. 11).
According to Somers, these machine-learning techniques are valued precisely because they allow serviceable applications to be built by abstracting away from the hard problems of human cognition and neuro-computattion. Moreover, the partitioners of the art, know this. These are not taken to be theories of thinking or cognition. And, if this is so, there is little reason to criticize the approach. Engineering is a worthy endeavor and if we can make life easier for ourselves in this way, who could object. What is odd is that these same techniques are now often recommended for their potential insight into human cognition. In other words, a technique that was adopted precisely because it could abstract from cognitive details is now being heralded as a way of gaining insight into how minds and brains function. However, the techniques here described will seem insightful only if you take minds/brains to gain their structure largely via environmental contact. Thinking from this perspective is just "data, data, data"plus the simple systems that process it.
As you may have guessed, I very much doubt that this will get us anywhere. Empiricism is the problem, not the solution. Interestingly, if Somers is right, AI's pioneers, the people that moved away from its initial goals and deliberately moved it in a more lucrative engineering direction knew this very well. It seem that it has taken a few generations to loose this insight.
The big drop in prestige occurred, Somers notes, in about the early 1980s. By then "AI…started to…mutate…into a subfield of software engineering, driven by applications…[the]mainstream had embraced a new imperative: to make machines perform in any possible, with little regard for psychological plausibility (p. 3)." The turn from cognition was ensconced in the conviction that "AI started working when it ditched humans as a model, because it ditched them (p. 4)." Machine translation became the poster child for how AI should be conducted. Somers gives a fascinating thumb nail sketch of the early system (called 'Candide' and developed by IBM) whose claim to fame was that it found a way to "avoid grappling with the brain's complexity" when it came to translation. The secret sauce according to Somers? Machine translation! This process deliberately avoids worrying about anything like the structures that languages deploy or the competence that humans must have to deploy it. It builds on the discovery that "almost doesn't work: a machine…that randomly spits out French words for English words" can be tweaked "using millions of pairs of sentences…[to] gradually calibrate your machine, to the point where you'll be able to enter a sentence whose translation you don't know and get a reasonable result…[all without] ever need[ing] to know why the nobs should be twisted this way or that (p. 10-11).
For this all to work requires "data, data, data" (as Norvig is quoted as saying). Take …"simple machine learning algorithms" plus 10 billion training examples [and] it all starts to work. Data trumps everything" Josh Estelle at Google is quoted as noting (p. 11).
According to Somers, these machine-learning techniques are valued precisely because they allow serviceable applications to be built by abstracting away from the hard problems of human cognition and neuro-computattion. Moreover, the partitioners of the art, know this. These are not taken to be theories of thinking or cognition. And, if this is so, there is little reason to criticize the approach. Engineering is a worthy endeavor and if we can make life easier for ourselves in this way, who could object. What is odd is that these same techniques are now often recommended for their potential insight into human cognition. In other words, a technique that was adopted precisely because it could abstract from cognitive details is now being heralded as a way of gaining insight into how minds and brains function. However, the techniques here described will seem insightful only if you take minds/brains to gain their structure largely via environmental contact. Thinking from this perspective is just "data, data, data"plus the simple systems that process it.
As you may have guessed, I very much doubt that this will get us anywhere. Empiricism is the problem, not the solution. Interestingly, if Somers is right, AI's pioneers, the people that moved away from its initial goals and deliberately moved it in a more lucrative engineering direction knew this very well. It seem that it has taken a few generations to loose this insight.
Had we known; Gerald would have been Chomsky's perfect foil
It seems that Coco is not the only articulate ape. Bob Berwick sent me this film of Gerald. I wish he had been available back when Amy (Penny) brought me (Coco) to debate Chomsky (Elan).
Subscribe to:
Posts (Atom)
