Here are some papers on the current hype over AI. They are moderately skeptical about the current state of the art. Not so much about whether there are tech breakthroughs to be had. All agree that these are forthcoming. The skepticism concerns the implications of this. Let me a say a word or two about this.
There is a lot of PR concerning how Big Data will revolutionize our conceptions of how the mind works and how science should be conducted. Big Data is Empiricism on steroids. It is made possible because of hardware breakthroughs in memory and speed of computation. We can do more of what we have always done faster and this can make a difference. I doubt that this tells us much about human cognition. Or more accurately, what it does tell us is likely wrong. Big Data is often coupled with Deep Learning. And linguists have every reason to believe that Deep Learning is an incorrect model of human cognition. Why? Because it is a modern version of the old discovery procedure. Level1 generalizations are generalized again at level 2 and level two generalizations are generalized agains t level 3 and so on. As a model of cognition, this tells us that higher levels are just generalizations over lower ones (e.g. from phonemes we get morphemes and from morphemes we get we get phrase structure and from phrase structure we get...). GG started form the demonstration that this is an incorrect understanding of linguistic organization. Levels exist, but they are in no sense reducible to the ones lower down. Indeed, whether it makes sense to speak of 'higher' and 'lower' is quite dubious. The levels interact but don't reduce. And any theory of learning that supposes otherwise is wrong if this is right (and there is no reason to think that it is not right, or at least no argument has been presented arguing against level independence). So, Deep Learning and Big Data are. IMO, dead theories walking. We will see this very soon.
The interview with Gary Marcus (here) discusses these issues and notes that historically what we have here is more of the same that has in the past proven to be wildly oversold. He thinks (and I agree) that we are getting another snow job this time around too. The interview rambles somewhat (bad editing) but there is lots in here to provoke thought.
A second paper on a similar theme is here in Aeon. The fact that it is in Aeon should not be immediately held against it. True, there is reason to be suspicious given the track record, but the paper was not bad, IMO. It argues that there is no "coming of the machines."
Here is a third piece on programming and advice about how not to do it (the Masaya). It interestingly argues for a Marrian conception of programming. Understand the computational problem before you write code. Seems like reasonable advice.
Last point: I mentioned above that Big Data is not only influencing how we conceive of Minds and Brains but also on how we should do science. The idea seems to be that that with enough data, the search for basic causal architecture becomes quaint and unnecessary. We can just vacuum up the data and the generalizations of scientific utility will pop out. On this view, theories (and not only minds) are just compendia of data generalizations and given that we can now construct these compendia more efficiently and accurately by analyzing more and more data, theory construction becomes a quaint pastime. The only real problem with Empiricism is that we did not gather enough data fast enough. And now, Big Data can fix this. The dream of understanding without thinking is finally here. Oy vey!
Showing posts with label Big Data. Show all posts
Showing posts with label Big Data. Show all posts
Wednesday, May 25, 2016
Saturday, April 13, 2013
Gary Marcus on Big Data
Gary Marcus has a nice little comment on Big Data (here). He makes three important points. First that it's not all that clear what 'Big Data' means, though it seems to involve doing something with very very large data sets and that it's currently the really BIG thing. Second, that there are some things that for some problems, looking for patterns in the data can be quite successful. To quote Gary: "Big Data can be especially helpful in systems that are consistent over time, with straightforward and well-characterized properties, little unpredictable variation, and relatively little underlying complexity." And third, that we've seen this before. As Gary notes, this is a reprise of the overhype that sunk strong AI (a fad that in its hay day was as immodest and on the make as Big Data is now). Or as Gary says it: "In fact, one could see the entire field of artificial intelligence as an inadvertent referendum on Big Data, because nowadays virtually every problem that has ever been addressed in A.I. -- from machine vision to natural language understanding -- has also been attacked from a data perspective. Yet most of the problems are unsolved, Big Data or no."
Scientists like to think of themselves as immune to fads. Nope. Big Data is the current Big Thing. Gary's piece makes many of its (unjustified) pretensions evident. It's a good read.
Scientists like to think of themselves as immune to fads. Nope. Big Data is the current Big Thing. Gary's piece makes many of its (unjustified) pretensions evident. It's a good read.
Wednesday, March 27, 2013
Big Data
Probably since forever philosophers and mathematicians have
dreamed of mechanizing thought, of removing judgment from thinking. The newest
aspirant in this millennial quest is Big Data, and not surprising there is an
eponymous book (excerpted here) with the following provided summary:
This revelatory exploration of big
data, which refers to our newfound ability to crunch vast amounts of
information, analyze it instantly and draw profound and surprising conclusions
from it, discusses how it will change our lives and what we can do to protect
ourselves from its hazards.
Big Data (BD) is the new New Thing, the method by which diligence
can substitute for thought. The idea actually has a certain charm as it
reverberates with our sense of justice. Collecting data is hard work, but it is
generally the kind of work for which effort is rewarded. Work hard and you will
do well. Put in the hours and the data will pile up. It’s an activity that rewards virtue.
In this it is entirely unlike
coming up with a plausible analysis, aka thinking. This activity is totally
unfair. Lazy people can have excellent
ideas. Sloth is no bar to insight and profligacy no guarantee of intellectual
stagnation. Here even the wicked,
sloppy, and lazy can prosper. How unfair.
In a just world, virtue would be rewarded. In a just world
hard work would guarantee enlightenment. We don’t live in a just world. Big
Data is the unfounded belief that this can be remedied. The hard work of data
gathering can substitute for the caprice of thought. It cannot be, and,
unfortunately, believing it can is likely to deform scientific practice. To see this, consider the following quote:
The era of big data challenges the
way we live and interact with the world. Most strikingly, society will need to
shed some of its obsession for causality in exchange for simple correlations:
not knowing why but only what. This overturns centuries of
established practices and challenges our most basic understanding of how to
make decisions and comprehend reality (10).
And that is precisely the problem. Big Data is part of an enterprise aimed at
reforming scientific practice. Dump why
aim for what. However, contrary to the prevailing
conception, without a model/theory it is not clear what it means to just “look
for correlations.” Data do not speak
for themselves. So gathering lots of data will not result in eloquent models
that understand the whats that
matter. Big data sets cannot pull themselves up by the bootstraps (nothing can pull itself up by the
bootstraps!) thereby yielding useful models. So, without explicit thoughtful models
that guide the enterprise, we will be saddled with implicit models that obscure
(and trivialize) what we are doing (as noted here without good models it is
even difficult to separate good data from bad).
None of this would be worth mentioning were it not for the
mesmerizing powers of Big Data. We have seen this before (here, and here for
example). Big Data is the modern avatar
to classical empiricist methodology. It’s appeal is its promises to provide
insight without intellectual sweat. This time, however, Empiricism has found a slogan
attached to a technology, Google being the all-powerful mantra. Not
surprisingly, money-making slogans can be very enticing and Google
intellectuals (e.g. Peter Norvig) can gain powerful platforms. And though I am
quite sure that like all other (empiricist) attempts to circumvent thought,
this too will ultimately fail, it’s demise may not come soon enough to prevent
serious damage. So when you hear the siren calls of Big Data I suggest the
following prophylactic procedure; repeat Kant’s dictum to yourself, viz. data
without theory is blind, data without theory is blind, data without theory is
blind…and hope it soon goes away.
Friday, March 8, 2013
How to Study Brains
A new comment in Scientific
American (here) discusses the challenges facing the newly hyped brain
mapping project that seems ready to get scads of funding. The writer, Partha Mitra, is the Crick-Clay
professor at the Cold Spring Harbor lab, a pretty fancy place for brain
research I am told. The comment is
interesting for several reasons and I urge you take a look. It’s pretty short
so it won’t cut into your weekend relaxation time much. Here are three things that caught my
attention.
First, he offers a shout out to the (often maligned)
competence-performance distinction within linguistics. As Mitra notes “from a
scientific perspective we want to know what the brain is capable of doing in principle, not what it actually does in a specific instance. In other words, we want to
understand the laws of brain
dynamics, not the details of brain
dynmaics.” I couldn’t agree more. Same
within linguistics. The aim of studying what people actually do linguistically
is in service of figuring out the underlying capacity. As noted here, this is the essence of the
rationalist perspective: what something is does not reduce to what something
does.
Second, Mitra points out that getting more and more data,
more and more careful measurements of more and more neurons is likely a
pointless exercise as this project suffers from conceptual incoherence: “the
proposal does not stand up to theoretical scrutiny.” There is a disease that Big Data Science is
susceptible to: viz. if its collectible, collect it! As Mitra points out, before vacuuming up
every stray data point “we need a theoretical
question of what we gain by recording every neuron.” Amen.
There is no substitute for thought and big data by itself cannot substitute
for theory. With theory in place data
becomes comprehensible. Without it, more data is as likely to mislead as to
enlighten. As Mitra poetically puts it: without theory “we risk the fate of the
naturalist Stapelton as he rushed across the great Grimpen Mire at the
conclusion of the Hound of the Baskervilles: even with all his knowledge, he
stepped into a bog, and was not heard of again.”
Third, Mitra makes a very useful analogy between earlier
work in physics and current work in neuroscience. Mitra first observes is that
in many domains of physics the “phenomena were first discovered at the
macroscopic level, studied at the macroscopic level, and even the theoretical
framework was established at the macroscopic level; the microscopic
measurements and statistical mechanical theories entered at a later stage to
refine the understanding already established.”
He then unpacks the fable to reveal the following moral: “Animal
behavior provides a close analog of the macroscopic behaviors of physical
systems…The study of psychological
pehnomena in terms of constructs like memory, attention, language and affect
also get at the macroscopic properties of nervous system dynamics…”(my
emphasis, NH). A least for Mitra, the
biolinguistic program is in no need of justification. It provides the necessary
Marrian “ “computationalist” perspective” that biology using “engineering
principles and evolution” can aim to explain. In other words, generative
theories of linguistic competence are targets of neuroscientific explanation,
the payoff being a deeper understanding of how evolution and good design “apply
to brains.”
All of this should sound pretty familiar, but it is nice to
hear the standard biolinguistic viewpoint endorsed with so little methodological
angst. Mitra takes it as obvious that cognitive theories are ultimately
biological theories and that they are necessary parts of the neuroscience
program of understanding how brains work. Indeed, for him linguistic theory is
to brain science what thermodynamics is to statistical mechanics. I’m perfectly happy with this company. If
this attitude spreads we may one day be spared the captious remarks of the
philosophically sophisticated. One can only hope.
Subscribe to:
Posts (Atom)