In case you do not know, the collection of essays that Angel Gallego and Dennis Ott put together to celebrate the publication of Aspects 50 years ago is now out and available both from MITWPL and online here. I'd be interested in people's reactions to these. For what it is worth, I think that the papers reflect (surprise surprise) two different conceptions (or at least emphases) in the field. Oddly, (and need I add that this is my opinion) the Chomsky vision expressed so eloquently and provocatively in Chapter 1 seems absent (at least in any explicit form) from many of the papers in the volume. Aspects is one of the most eloquent expressions of the cognitive (and, hence the biological) conception of linguistics. And if it is true (which it is) that God is in the details, then it is equally true that details that fail to bear on the central cognitive questions (descriptive and explanatory adequacy) are not details of obvious relevance to the Aspects program. In other words, we should always be asking ourselves what a particular analysis contributes to our understanding of FL. We may not always know, but we should always be asking.
I hope to write a little more about this in coming posts. But please feel free to let say how this volume strikes you in the comments section and thanks to Angel and Dennis for putting in all that work. It is a good snapshot of where the filed currently is, I believe.
Tuesday, September 15, 2015
The ling curriculum
Here's a post that Colin out on the diversity of grad curricula. He talks about the UMD experience and how it reflects a changing view of the aim of ling research. I think that it accurately reflects a consensus of the faculty at UMD about what the field should be about. It emphasizes the use of a broad variety of techniques to investigate linguistic phenomena and it works hard not to privilege some kinds of work over others. I also agree that this approach has worked well enough and it correctly notes the kinds of trade-offs that a modern grad education in ling presents.
One thing that I would emphasize that Colin did not: this consensus has been possible because there is a broad agreement among the faculty of what the object of study is, viz. the Faculty of Language. In other words, we all agree that we are studying the same object and where we differ is in the kinds of probes we use to study it. Without this common conception I think that garnering consensus would have been very difficult. I also believe that the view of the UMD faculty is, sadly, not widely shared. Many core area linguists do not believe (or do not think it important and think it maybe even a little BSy) to interpret syntax or phonology or semantics or morphology from a broadly cognitive perspective. But unless one does, then it is hard to see why the conjunctive areas of interest (psycho, computational, neuro) are not from the perspective of linguistics secondary areas. From this (non UMD) perspective the core areas provide neutral tools to study the mind with but don't substantially prejudice this study empirically. Sort of like arithmetic and physics. IMO, if this is your view, then there really are first and second class citizens and it is irresponsible not to lead with the core areas.
Last point. Colin notes, but I would like to emphasize, that disengaging from the cognitive/neuro conception of language is likely to be very dangerous for the health of the field. There is a reason that linguistics has lost the prestige it once had among the psycho sciences and part of the reason is that linguists have lost interest in these wider issues. If we don't return to this conception we will loose out. So, not only is Colin's position intellectually reasonable, I suspect that the future health of the field depends on it.
One thing that I would emphasize that Colin did not: this consensus has been possible because there is a broad agreement among the faculty of what the object of study is, viz. the Faculty of Language. In other words, we all agree that we are studying the same object and where we differ is in the kinds of probes we use to study it. Without this common conception I think that garnering consensus would have been very difficult. I also believe that the view of the UMD faculty is, sadly, not widely shared. Many core area linguists do not believe (or do not think it important and think it maybe even a little BSy) to interpret syntax or phonology or semantics or morphology from a broadly cognitive perspective. But unless one does, then it is hard to see why the conjunctive areas of interest (psycho, computational, neuro) are not from the perspective of linguistics secondary areas. From this (non UMD) perspective the core areas provide neutral tools to study the mind with but don't substantially prejudice this study empirically. Sort of like arithmetic and physics. IMO, if this is your view, then there really are first and second class citizens and it is irresponsible not to lead with the core areas.
Last point. Colin notes, but I would like to emphasize, that disengaging from the cognitive/neuro conception of language is likely to be very dangerous for the health of the field. There is a reason that linguistics has lost the prestige it once had among the psycho sciences and part of the reason is that linguists have lost interest in these wider issues. If we don't return to this conception we will loose out. So, not only is Colin's position intellectually reasonable, I suspect that the future health of the field depends on it.
Sunday, September 6, 2015
Two shortish perusables
Here are a couple of things that crossed my desk this week
that you might find interesting.
The first (here)
is from Christoph and lifted from the comments section of this.
The piece is a short comment by Hilary Putnam (who, btw, was the head of my
thesis committee) on the “innateness hypothesis.” He reiterates why he rejects
the hypothesis. It comes down to the mis-assertion that Chomsky’s version of
FL/UG requires that it “provide for all
possible meanings,” which Putnam
takes to be ridiculous because Chomsky specifies no “mechanism [that] could
have endowed the brains of primitive men and women with such ‘particular
meanings’ as “quantum potential” and “macroeconomic” or with terms by means of
which they could be defined, if indeed, there are more elementary terms in
which this could be done.” The absence of a proposed mechanism makes it clear
to Putnam that the ‘innateness hypothesis’ must be wrong. That’s the argument.
Unfortunately, it’s really very bad.
Philosophers love this kind of argument from incredulity.
Putnam’s main current contribution to the “debate” is to point to concepts like
macroeconomic rather than carburetor as the ones that are
particularly problematic. But the claim has the familiar
surely-you-don’t-mean-to-say form so beloved of philosophers. The unfortunate
thing is that Putnam, I believe, has simply misunderstood both Chomsky’s and,
more relevantly, Fodor’s positions on these matters. Let me explain.
First a small terminological point. As Chomsky has often
remarked, it is unclear what the ‘innateness hypothesis’ is supposed to be. It
cannot be that anyone doubts that minds come equipped with innate structure. Everyone
assumes that the mind/brain has given structures and operations that guide/bias
learning/acquisition. Truly blank slates stay blank. The question has never
been whether minds/brains have innate
structure but what is innate, what
kinds of generalizations are minds/brains predisposed to make so that when
confronted with input they generalize beyond it? Everyone thinks that there is
something. The question is what. Chomsky’s simple point is and always has been
that there is every reason to think that in the domain of language, the
mind/brain has methods of generalization specific to linguistic forms and that
any kind of simple associationism built around mere sensory input has not, will
not and cannot work. If there is an innateness hypothesis worth discussing, it
is the specific suggestion that language competence relies on language specific
mental capacities and cannot be reduced entirely to other cognitive capacities.
And this requires discussing details, something that your average friendly
famous philosopher of language has rarely (never?) done.
Second, if this is what Chomsky intended, then it’s clear
that Putnam’s observations don’t bear on it. Specifically, so far as I know,
Chomsky has had very little to say about where concepts or lexical meanings
come from. In fact, so far as I know, nobody
(including Putnam) has any idea of how concepts arise in minds. Chomsky has
repeatedly said as much, pointing to the human capacity for lexical acquisition
as being a mystery. So, whatever Putnam
is saying here, it bears less on Chomsky’s views (which have largely been
confined to claims about syntactic
structure (catchy phrase huh? Maybe a good book title?). Indeed, since the very
earliest days of GG, Chomsky has been very very very circumspect in his
discussion of lexical meaning and how much we understand about it. My reading
of his discussions of these matters is that Chomsky’s main conclusion is a
negative one; that whatever lexical meaning is, it’s not just a matter of
referential dependency (see here for
some discussion).
But, third, maybe the target of Putnam’s comments is not
Chomsky but Fodor. Fodor has made an argument that family resembles the one
that Putnam summarily dismisses. But, even as regards Fodor, I think that
Putnam mistakes the claims. Very
briefly, what Fodor argued is that the only theories of induction we have are
“selection” theories. What I mean is that they understand learning as the
inductive fixation of belief given a space of possible hypotheses. Induction
works by moving one around this space of possibilities and, if there is enough
data of the right kind, settling in one part of the space or another (see here for
discussion). Thus, what we have are theories of belief fixation given a space
of possible concepts/beliefs. Given that this is what we have, there is a sense
in which you cannot possibly acquire anything that you don’t already have the
wherewithal to conceptually represent. That was Fodor’s first point.
He coupled it with a second series of arguments that denied
that most lexical meanings were decompositional (something that the Putnam
quote above seems to agree with). So, if most word meanings cannot decompose
and fixation requires representation of the concepts that are fixed then there
is a sense in which the meaning of ‘carburetor’ is there in the mind from the
get-go. This is Fodor’s argument.
Now Putnam clearly dislikes the conclusion. Say he is right,
that it is a reductio whose premise must be false. What does that tell us. Well
it implies that there must be some other theory of learning besides the
inductive ones that we all know and love. Recall that Fodor’s argument is that
inductive learning theories imply all the fixable concepts are (in one sense)
innate. So if you don’t like this conclusion you must show that either
inductive learning theories do not presuppose hypothesis spaces (or analogues
thereof) contrary to what Fodor noted, or that there are other theories of
learning that are non-inductive that explain how we acquire words and meanings.
In other words, if you don’t like the conclusion then you either need to show
where Fodor’s description of induction fails or suggest that induction is not
the only way to learn and to provide an outline of the other kinds. Putnam does
neither.
Curiously, I think that Fodor might agree with the second
option. In the Modularity of Mind, if
I recall correctly, Fodor suggested that only modular cognitive systems are
amenable to current investigation. One way of reading this is that only
informationally restricted modular domains are ones where the hypothesis
space/inductive procedure story can be made to work. Moreover, Fodor is on
record opposing the conception of the mind as massively modular (i.e. made up
of endless numbers of small modules) and thinks that something else (he knows
not what) is going on in central system cognition. It is consistent with
Fodor’s views that lexical acquisition is not
inductive and so there is some other way that concepts are acquired. But, and
this is key, he does not have the remotest inkling as to what this other
process might be and, so far as I can tell, neither does Putnam.
Putnam also throws in some comments on Chomsky and Fodor’s
views about evolution but it is entirely unclear what any of the (misnamed) innateness
controversy has to do with evolution. As regards lexical acquisition, for
example, Fodor is not against our conceptual repertoire changing over time, he
is against it changing by induction/learning over time.
Sorry to have gone on so long. Putnam’s remarks are not new,
as he himself points out. It seems, however, that his current views miss the
mark as much today as they did when first advanced them. The more things change…
Here
is a second paper on referentially ready minds. The paper is in Nature and two of the authors should be
well known to linguists. It makes, to my mind, the modest point that kids are
ready to take language as an indicator of referential intent when accompanied
by other behavior (e.g. eye gaze). It seems that even very young kinds show
indications of thinking that language use goes hand in hand with referential
intent.
One question I had is what I am supposed to take away form
this? What does it tell us about acquisition? Is the suggestion that
establishing referential value is an important factor in language acquisition?
If so, how big a factor? Bigger than distributional analysis? Is being
reference ready a critical pre-condition for language acquisition? Is the
supposition that reference is what meaning consists in (if so, see Chomsky’s
relevant remarks on this topic linked to above). I am not sure. So, let me ask
you: what’s the take home message here and why is what the paper argued for important?
BTW, this is a sincere question: what’s the overall take home message? That
language can be used referentially and that kids come natively equipped to
believe this? Or is there something more going on here?
Thursday, September 3, 2015
Brains with language
Humans speak and other animals don’t, or at least don’t like
we do it. The qualitative difference is so obvious that only sophisticates
could talk themselves into thinking that the linguistic capacities we find in
humans are even roughly continuous with those found in other animal
communication systems. If you are not a dualist, this suggests that there must
be something different about human brains upon which this unique linguistic
capacity supervenes. Everyone would love
to know what this difference consists in. In fact this desire makes even
skeptical scientists (yes, this is ironic) trigger happy, for the moment anyone
finds any difference between human brains and those of our “nearest” cousins
this news makes a big splashy headline announcing the discovery of yet another
breakthrough explaining how we got to be so verbose. One example nice of this is a recent paper by
Wang, Uhrig, Jarraya and Dehaene (WUJD) here.
Whatever the value of the results (I’ll discuss more details
in a minute) the PR associated with it is quite breathless. Here is one
example from Phys.org. The headline is “Breakthrough in understanding the
origins of human language.” The result, if it bears on this question at all, apparently
does so in an entirely modal way. What I mean by this, the most that one can
conclude from this study (as the carfule wording quoted below shows) is that maybe the thing that WUJD finds relates
to language in some way. But then
again, many things might relate to
language in some way. My own view, on
a quick reading of WUJD, is that it is quite unclear which features of language
(if any) the differences WUJD discovers between human and macaque brains
explain. And if, per chance, we take recursion to be the most distinctive
property of human language, then it is likely that the WUJD results tell us
nothing at all about this most distinctive property. Let me elaborate.
The paper shows that humans treat sound sequences differently
form macaques as follows. Both groups can track patterns (e.g. AAAB vs ABAB)
and both can track different number of sounds (e.g. AAAB vs AAAAB). However,
whereas in macaques this information is segregated in different parts of the
brain, human brains have areas that are sensitive to both these parameters at
once (i.e. there are areas in the human brain which integrate these two kinds
of information). Moreover, it appears that the locus of this integration
coincides with areas of the brain implicated in language processing (let’s hear
it for Broca’s area!!)[1].
The big conclusion is “while some abstract properties of auditory sequences are
available to non-human primates, a recently evolved circuit may
endow humans with a unique ability for representing linguistic and
non-linguistic sequences in a unified manner “
(my emphasis, NH, p. 1966).[2]
Sure it may, but then again it may not. What neither the paper nor the reports
discuss is what specific aspects of language this novel brain property is
responsible for. In other words, granted that humans can do these things and
that there are dedicated brain areas tasked with keeping track of this kind of
joint information, which features of language does this newly enabled capacity underwrite.
I have no idea.
And this is not good. As I’ve said many times before, we
know a lot about the formal properties of human language. The paper is
suggesting that being able to integrate two kinds of abstract information is
useful in doing something linguistic. Ok, which linguistic thing is it useful
for? I am pretty sure that it says nothing at all about the biggie Chomsky
talks about (i.e. recursion) for the stimuli discussed involve simple
unembedded templates. These are abstract, but not nearly as abstract as the
recursive structures that natural languages generate. We can track these
templates in the data for they create
regular patterns there (see here).
The problem with linguistic patterns is that there aren’t any in the absence of
the generative procedures that give rise to them.[3]
If by patterns you mean abstract templates, then much of language structure is
without pattern. Which part? Well, at least the recursive part.
Of course, WUJD might be pointing to something else. Maybe
the relevant patterns are not like out syntactic ones but more like phonological
ones. Maybe. But then it would have been useful were WUJD to indicate how the
kind of abstract information integration WUJD notes suffices for/is necessary
for/ relates to the formal patterns we find in the sound structures or
morphological structure or whatever structure found in natural languages.
Absent this, there is little reason to think that the identified integrative
capacity has anything much to do with
language.
Let me say this another way. One contribution linguists have
made to the cog-neuro of language has been to provide pretty well worked out formal
descriptions of the kinds of rules we find in natural languages. These rules
specify an endless variety of different kinds of possible patterns. Language
competent brains must be able to compute these patterns using these kinds of
rules. A brain based “breakthrough in understanding the origins of language”
must discuss how brains compute these very well described structures. It is not
enough to generically note that these structures are “abstract” or “algebraic”
or “symbolic.” Sure they are, but so are
many other things as there are many (infinitely many?) ways of being
“abstract,” “algebraic,” and “symbolic.” And given that we actually know
something about the very specific ways that human languages are “abstract” and
“algebraic” and “symbolic” it behooves those wishing to address the etiology of
our linguistic capacities or their brain bases to relate their findings to
these specific properties. It is simply not enough to note that our brains
differ from “their” brains in some way (a typical example: area A is tied to B
in us but not in them or our brains are round and theirs are oblong therefore…)
and think that this tells us something about language. It really doesn’t. It
certainly does not merit the kind of hype noted above. And this is from
scientists, even ones that I admire.
[1]
To be sure, it seems to be common wisdom that Broca’s area lights up in almost
any fMRI experiment. In fact, I recall David Poeppel once snarking that one
prerequisite of any well-designed brain experiment is that it light up Broca’s
area.
[3]
This very important point is made in Jackendoff’s mis-named intro text Patterns in the Mind.
Tuesday, September 1, 2015
How I spent my summer vacation
Like all avid GGers, I spend part of my summer vacation
rereading some of the greatest hits.
This year, this included rereading Current
Issues in Linguistic Theory (CILT). If you haven’t done so recently, you
should go out and read it (again) now,
before the BBC series comes out on public TV. It is a fantastic little book and
very timely for it lays out more clearly than anything else I know what the
original GG enterprise took the central questions of interest to be. It is
worth knowing what these were (and still are) for it helps prevent
energetically chasing off in the wrong direction in pursuit of answers to
questions of dubious utility. In other
words, knowing what you are interested in, what the research questions are,
helps you to avoid wasting time. Remember: those things not worth doing are not
worth doing well. And, there are many too many things that are really not worth
doing. I will mention one such enterprise below that seems to have recently stirred
the tea post tempestuously once again.
This said, back to CILT. The book starts with a short
important chapter on the goals of ling theory: what are the things we want to
explain? Chomsky points to two central questions: (i) Linguistic Creativity: how
do competent speakers go (in various kinds of uses (e.g. production,
comprehension)) from utterances to structural descriptions of those utterances
and (ii) The Logical Problem of Language Acquisition: how do kids/acquirers go
from primary linguistic data to their acquired generative grammars. These are the two questions we want to answer
and this involves limning the fine structure of Gs and FL/UG. More abstractly
CILT provides the following two important mnemonic diagrams.
(1) utterances
à
A à
structural description
(2) PLD à B à
generative grammar
A is a place-holder for (at least) a particular G and B for
the theory of FL/UG. The aim of inquiry
is to describe the innards of A and B. Or
as Chomsky puts it:
The perceptual model A is a device
that assigns a full structural description D to a presented utterance U,
utilizing in the process its internalized generative grammar G, where G
generates a phonetic representation R of U with the structural description
D…The learning model B is a device which constructs a theory G (a generative
grammar G of a certain langue) as its
output on the basis of primary linguistic data (e.g. specimens of parole) as input…We can think of general
linguistic theory as an attempt to specify the character of device B. We can
regard a particular grammar as, in part, an attempt to specify the information
available in principle (i.e. apart from limitations of attention, memory, etc.)
to A that makes it capable of understanding an arbitrary utterance, to the
highly non-trivial extent that understanding is determined by the structural
description provided by the generative grammar. (26)
Thus G and FL are seen as causally relevant factors in
explaining various kinds of performances; normal discourse and acquisition. The
aim of linguistics is to describe these two mechanisms, which causally
contribute to these two kinds of “behavior,” (i.e. talking and language
acquisition).
What is the empirical criterion of adequacy for the two
cases? The relevant measure of evaluation for the first is that it “correctly
describes the linguistic intuition of the speaker” (26). Note the singular intuition! What we call linguistic
intuitions reflect a speakers
grammatical intuition (i.e. the sense of his/her language). That’s why they are
important. But the thing we want our theory of G to match is the singular, the
plural being interesting to the degree that it reveals this. We return to this
anon.
The relevant measure for evaluating proposals about (2) are
that the Gs B selects correspond to “the speakers’ linguistic intuition, in the
case of particular languages” (27). Thus, the adequacy of B is judged in
relation to how good the Gs it selects are in describing a native speaker’s
actual grammatical intuition (i.e. the mental structures that underlie
linguistic facility).
The aim, then, is to describe the features of real cognitive
objects, either Gs that speakers actually have and procedures for constructing
these Gs that speaker’s come equipped with. These are the objects of inquiry
and what linguistic theories should aim to model.
Why does Chomsky take these as the two central problems?
Because of two basic very big and very obvious facts. The first, and the one
that he hammers again and again in CILT, is the fact of linguistic creativity,
by which Chomsky intends the following:
…a mature native speaker can
produce a new sentence of his language on the appropriate occasion, and other
speakers can understand it immediately, though it is equally new to them. Most
of our linguistic experience, both as speakers and hearers, is with new sentences;
once we have mastered a language, the class of sentences with which we can
operate fluently is so vast that for all practical purposes (and, obviously,
for all theoretical purposes), we may regard it as infinite. (7)
So, the fact of linguistic creativity implicates mastery of
a recursive procedure (aka a G) that is
used by native speakers in understanding and producing utterances. And the fact
that Gs are required to explain this creativity means that a native speaker
must acquire such a G in order to be fluent. So, the aim of linguistics is to
describe these Gs and explain how they are acquired.
Chomsky also notes a second interesting capacity that G
knowledge endows a native speaker with: the “ability to identify deviant
sentences and, on occasion, to impose an interpretation on them” (7).
As we all know, this second capacity (aka: native speaker
linguistic intuitions) has proven to
be an excellent window into the structure of a native speakers language
specific capacity. The generative enterprise has relied on this capacity to
probe the structure of G and FL. It is what licenses GGs reliance on linguistic
intuitions (note the ‘s’ here) as
guides to the structure of linguistic intuition (note the absence of an ‘s’).
So, Gs explain (in part) how linguistic creativity is possible and their
structure can be probed by querying native speakers’ evaluations of the
acceptability and interpretation of products of these Gs.
It is worth noting that the second capacity does not follow
from the fact that speaker’s possess Gs. This could have been true without it
being true that speakers could usefully reflect on G products. Speakers could have
used Gs to speak and understand without having reliable linguistic intuitions useful
for probing this capacity.
Chomsky discusses skepticism regarding such judgment data in
chapter 3. Not surprisingly he concludes that it’s the best thing we’ve got and
that “[w]e neglect such data at the cost of destroying the subject” (56). However,
as Chomsky noted in CILT such judgments are not “sacrosanct and beyond any
conceivable doubt” (56). Some such data might be bad (as any data in any area
might be). We can firm it up by looking for “consistency among speakers of
similar backgrounds” as well as “for a particular speaker on different
occasions” (56). In other words, such data as a class are fine, though
particular instances are reasonably challenged.
Chomsky notes a second important check on the reliability of
such data. I call it the proof-of-the-pudding test: “The possibility of
constructing a systematic and general theory” also matters. Theory tests data
just as much as data tests theory. With 60 years of hindsight we can conclude
that such data has been very useful and reliable precisely because the theories
built on it have proven to be remarkably insightful.
So, CILT picks out two central questions for linguistic
investigation and explains why they should be cynosures of further inquiry.
Moreover, he outlines the kind of data that is relevant in pursuing these
questions. Moreover, and most famously, in chapter 2 he outlines what he takes
to be the relevant measures of theoretical adequacy; observational, descriptive
and explanatory. Let’s turn to this
next.
Chomsky identifies three levels of adequacy for grammatical
description: (i) observational adequacy, (ii) descriptive adequacy and (iii)
explanatory adequacy.
Observational adequacy is “the lowest level” and is achieved
“if the grammar presents the observed primary data correctly” (29). As Chomsky
is quick to point out (see his note 1) what constitutes the relevant observable data is not at all
straightforward. One measure of relevance involves the “possibility for a
systematic theory.” Moreover, in an important sense, what linguists are looking
for are data that bear on linguistic structure and so good data is that which
are sensitive to these structures. Sadly, however, linguistic structure is not itself
observable and is only accessible to a speaker only via an utterance that
embodies it. Some utterances are good windows into these structures and so
judgments based on these are generally useful (that’s why quite often unacceptability is a better window into
G than acceptability). But some are not. Useful data allows one to infer
grammatical structure from its effects in the visible utterance, and what data
does this is not always obvious. As Chomsky puts it:
The problem of determining what
data is valuable and to the point is not an easy one. What is observed is often
neither relevant nor significant, and what is relevant and significant is often
very difficult to observe, in linguistics no less than…anywhere in science.
What’s a descriptively adequate description? It’s one that
“gives a correct account of the linguistic intuition of the native speaker, and
specifies the observed data (in particular) in terms of significant
generalizations that express the underlying regularities in the language” (28).
So, a descriptively adequate description will enumerate the properties of that
G that the speaker has internalized. Given that Gs are recursive rules systems,
they will (implicitly) embody regularities characteristic of the language they
generate.
Last of all we get to explanatory adequacy. Theories of
grammar achieve this level if they provide
“a general basis for selecting a grammar that achieves he second level of
success over other grammars consistent with the relevant observed data that do
not achieve this level of success” (28).
Explanatory adequacy is a predicate of theories of FL. Descriptive
adequacy is a predicate of Gs. Explanatorily adequate FLs are those that derive
descriptively adequate Gs relative to some specification of PLD. As is clear,
issues of descriptive and explanatory adequacy are intimately intertwined with
considerations of both bearing on the adequacy of each. Like it or not, claims
about descriptive adequacy commit hostages to explanatory adequacy no less than
do claims about the latter for the former. Given the close connection between explanatory
adequacy and Plato’s Problem and the PoS issues that surround it, Chomsky’s
vision of linguistics demands that these concerns be at the center of every
linguist’s attention (sad to say, IMO, this is hardly the case nowadays).
Chapter 2 does a very nice job operationalizing these
notions in the context of linguistic theory circa the early to mid 60s. There
is still lots to learn by reading these discussions (especially, IMO, the
section on levels of adequacy in semantics).
It is also worth carefully re-reading section 2.4 where Chomsky sums up
the discussion of the importance of the measures. Here is his blunt assessment
(52):
…three levels of adequacy have been
sketched…Of these, only the levels of descriptive and explanatory adequacy (and
ultimately on the latter) are of sufficient interest to justify further
discussion.
This makes perfect sense given the two questions CILT
highlights. If you are interested in
human language, then the name of the game is ultimately to describe the
properties of FL/UG. All else is interesting to the degree that it contributes
to this end. Unfortunately, much current work on language seems to assume that
discussions of PL/UG are at best premature and quite often little more than cow
pie. Many are happy to limit their
interests to “coverage the data,” aiming primarily for observational adequacy,
with some pretensions to descriptive adequacy. Chomsky has some choice remarks
about this. Here are two:
It is important to bear in mind
that a grammar that assigns correctly the mass of structural descriptions
(remote as this is from present hopes) would still be of no particular
linguistic interest unless it also were to provide some insight onto those
formal properties that distinguish a natural language from arbitrary,
enumerable sets of structural descriptions. At best, such a grammar would help
to clarify the subject matter for linguistic theory, just as a fourteenth
century clock depicting the positions of the heavenly bodies merely posed, but
did not even suggest an answer to the questions to which classical physics
addressed itself. (52-3)
In other words, work that fails to at least suggest
something about the structure of FL/UG is of very dubious value given the
central questions of GG. A corollary suggests itself: it is always worth
explicitly asking what light some piece of work tells us about FL/UG. If the
answer is unclear, then this is very much worth knowing.
Here’s the second quote:
Comprehensiveness of [data, NH]
coverage does not seem to me to be a serious or significant goal in the present
stage of linguistic science. Gross coverage of data can be achieved in many
ways, by grammars of very different forms. Consequently, we learn little about
the nature of linguitc structure from the study of grammars that merely
accomplish this…[I]t is only by studying the properties of gramamrs that
achieve higher levels of adequacy and by gradually increasing the scope of
description without sacrificing depth of analysis that we can hope to sharpen
and extend our understanding of the nature of linguistic structure. (53)
I see no reason to think that we have finally reached a
stage where big data work will shed much light on the structure of grammar. In
fact, I would go further. I doubt that data coverage in the big data/corpus
linguistic sense will ever be of linguistic interest. Why? Because it is seldom
driven by the impulse of uncovering the basic properties of linguistic
structure. An explanatory theory aims to uncover the basic operations and
principles that descriptively adequate Gs deploy. There is no doubt that these
principles interact with many other non-linguistic factors in every day speech.
But if your interest is in these principles and operations, then it needs a
good argument to conclude that looking at speech in the wild or lots of it will
reveal what these principles are. It’s not how things proceed in the “real”
sciences, so why think that this is the right way of doing linguistics? Beats
me.
One last bon mot from Chomsky: He makes an important distinction
between exceptions and counter-examples. The latter are important, the former
not so much, or not obviously much. A counter example is interesting because it
contradicts a principle and principles are what FL/UG is all about. An
exception need not. It may simply show what is already conceded, that our
theories do not aim towards broad data coverage, i.e. text fidelity. Here’s Chomsky on this:
Examples that lie beyond the scope
of a grammar are quite innocuous unless they show the superiority of some
alternative grammar. They do not show that the grammar as already formulated is
incorrect. Examples that contradict the principles formulated in some general
theory show that, to at least this extent, the theory is incorrect and needs
revision. (55)
What we are interested in is the failure of principles, not
in the failure of coverage.
CILT should be required reading for all GGers. From where I
sit, its theoretical and methodological observations are as relevant today as
they were when first written. In fact, they may be more relevant today. As a
field becomes technically more sophisticated it can loose its bearings.
Technique substitutes for insight. Keeping ones eyes on the central questions
of interest is a useful prophylactic against this. CILT has a very clear
research agenda. It has a clear target of explanation and outlines relevant
criteria of success. It has the virtue of being clear about these things. If
you too are interested in these fundamental questions, then nightly chanting
from the pages of CILT will serve you well.
Monday, August 31, 2015
A useful addendum to the previous post
I ran across this nice little discussion today (here) that others might find interesting. It identifies a plausible cost behind getting too methodologically stringent. There is a tradeoff between generating false positives and false negatives. What we want are methods that get to the truth when it is there (sensitivity) and tell us when what we think is so is not (specificity). There is, not surprisingly, a tradeoff here and tightening our "standards" can have unfortunate effects. This is especially true in domains where we know little. When you know a lot, then maybe generating exciting new ideas is less important than not being mislead experimentally. But when we know little, then maybe a little laxity is just what the scientists ordered. At any rate, I found the framing of the issues useful so I pass the piece onto you.
Friday, August 28, 2015
Stats and the perils of psych results
As you no doubt all know, there is a report today in the NYT about a study in Science that appears to question the reliability of many reported psych experiments. The hyperventilated money quote from the article is the following:
Here are some random thoughts, but I leave it to others who know more about these things than I do to weigh in.
First, it is not clear to me what should be made of the fact that "only" 39% of the studies could be replicated (the number comes from here). Is that a big number or a small one? What's the base line? If I told you that over 1/3 of my guesses concerning the future value of stocks were reliable then you would be nuts not to use this information to lay some very big bets and make lots of money. If I were able to hit 40% of the time I came up to bat I would be a shoe-in inductee at Cooperstown. So is this success rate good or bad? Clearly the headline makes it look bad, but who knows.
Second, is this surprising? Well, some of it is not. The studies looked at articles from the best journals. But these venues probably publish the cleanest work in the field. Thus, by simple regression to the mean, one would expect replications to not be as clean. In fact, one of the main findings is that even among studies that did replicate, the effects sizes shrank. Well, we should expect this given the biased sample chosen from.
Third, To my mind it's amazing that any results at all replicated given some of the questions that the NYT reports being asked. Experiments on "free will" and emotional closeness? These are very general kinds of questions to be investigating and I am pretty sure that these phenomena are the the results of the combined effects of very many different kinds of causes that are hard to pin down and likely subject to tremendous contextual variation due to unknown factors. One gets clean results in the real sciences when causes can be relatively isolated and interaction effects controlled for. It looks like many of the experiments reported were problematic not because of their replicability but because they were not looking for the right sorts of things to begin with. It's the questions stupid!
Fourth, in my shadenfreudeness I cannot help but delight in the fact that the core data in linguistics gathered in the very informal ways that it is is a lot more reliable (see Sprouse and Almeida and Schutze stuff on this). Ha!!! This is not because of our methodological cleverness, but because what we are looking for, grammatical effects, are pretty easy to spot much of the time. This, does not mean, of course, that there aren't cases where things can get hairy. But over a large domain, we can and do construct very reliable data sets using very informal methods (e.g. can anyone really think that it's up for grabs whether 'John hugged Mary' can mean 'Mary hugged John'?). The implication of this is clear, at least to me: frame the question correctly and finding effects becomes easier. IMO, many psych papers act as if all you need to do is mind your p-values and keep your methodological snot clean and out will pop interesting results no matter what data you throw in. The limiting case of this is the Big Data craze. This is false, as anyone with half a brain knows. One can go further, what much work in linguistics shows is that if you get the basic question right, damn methodology. It really doesn't much matter. This is not to say that methodological considerations are NEVER important. Only that they are only important in a given context of inquiry and cannot stand on their own.
Fifth, these sorts of results can be politically dangerous even though our data are not particularly flighty. Why? Well, too many will conclude that this is a problem with psychological or cognitive work in general and that nothing there is scientifically grounded. This would be a terrible conclusion and would affect linguistic support adversely.
There are certainly more conclusions/ thoughts this report prompts. Let me reiterate what I take to be an important conclusion. What these studies shows is that stats is a tool and that method is useful in context. Stats don't substitute for thought. They are neither necessary nor sufficient for insight, though on some occasions they can usefully bring into focus things that are obscure. It should not be surprising that this process often fails. In fact, it should be surprising that it succeeds on occasion and that some areas (e.g. linguistics) have found pretty reliable methods for unearthing causal structure. We should expect this to be hard. The NYT piece makes it sound like we should be surprised that reported data are often wrong and it suggests that it is possible to do something about this by being yet more careful and methodologically astute, doing our stats more diligently. This, I believe, is precisely wrong. There is always room for improvement in one's methods. But methods are not what drive science. There is no method. There are occasional insights and when we gain some it provides traction for further investigation. Careful stats and methods are not science, though the reporting suggests that this is what many think it is, including otherwise thoughtful scientists.
...a painstaking yearslong effort to reproduce 100 studies published in three leading psychology journals has found that more than half of the findings did not hold up when retested.The clear suggestion is that this is a problem as over half the reported "results" are not. Are not what? Well, not reliable, which means to say that they may or may not replicate. This, importantly, does not mean that these results are "false" or that the people who reported them did something shady, or that we learned nothing from these papers. All it means is that they did not replicate. Is this a big deal?
Here are some random thoughts, but I leave it to others who know more about these things than I do to weigh in.
First, it is not clear to me what should be made of the fact that "only" 39% of the studies could be replicated (the number comes from here). Is that a big number or a small one? What's the base line? If I told you that over 1/3 of my guesses concerning the future value of stocks were reliable then you would be nuts not to use this information to lay some very big bets and make lots of money. If I were able to hit 40% of the time I came up to bat I would be a shoe-in inductee at Cooperstown. So is this success rate good or bad? Clearly the headline makes it look bad, but who knows.
Second, is this surprising? Well, some of it is not. The studies looked at articles from the best journals. But these venues probably publish the cleanest work in the field. Thus, by simple regression to the mean, one would expect replications to not be as clean. In fact, one of the main findings is that even among studies that did replicate, the effects sizes shrank. Well, we should expect this given the biased sample chosen from.
Third, To my mind it's amazing that any results at all replicated given some of the questions that the NYT reports being asked. Experiments on "free will" and emotional closeness? These are very general kinds of questions to be investigating and I am pretty sure that these phenomena are the the results of the combined effects of very many different kinds of causes that are hard to pin down and likely subject to tremendous contextual variation due to unknown factors. One gets clean results in the real sciences when causes can be relatively isolated and interaction effects controlled for. It looks like many of the experiments reported were problematic not because of their replicability but because they were not looking for the right sorts of things to begin with. It's the questions stupid!
Fourth, in my shadenfreudeness I cannot help but delight in the fact that the core data in linguistics gathered in the very informal ways that it is is a lot more reliable (see Sprouse and Almeida and Schutze stuff on this). Ha!!! This is not because of our methodological cleverness, but because what we are looking for, grammatical effects, are pretty easy to spot much of the time. This, does not mean, of course, that there aren't cases where things can get hairy. But over a large domain, we can and do construct very reliable data sets using very informal methods (e.g. can anyone really think that it's up for grabs whether 'John hugged Mary' can mean 'Mary hugged John'?). The implication of this is clear, at least to me: frame the question correctly and finding effects becomes easier. IMO, many psych papers act as if all you need to do is mind your p-values and keep your methodological snot clean and out will pop interesting results no matter what data you throw in. The limiting case of this is the Big Data craze. This is false, as anyone with half a brain knows. One can go further, what much work in linguistics shows is that if you get the basic question right, damn methodology. It really doesn't much matter. This is not to say that methodological considerations are NEVER important. Only that they are only important in a given context of inquiry and cannot stand on their own.
Fifth, these sorts of results can be politically dangerous even though our data are not particularly flighty. Why? Well, too many will conclude that this is a problem with psychological or cognitive work in general and that nothing there is scientifically grounded. This would be a terrible conclusion and would affect linguistic support adversely.
There are certainly more conclusions/ thoughts this report prompts. Let me reiterate what I take to be an important conclusion. What these studies shows is that stats is a tool and that method is useful in context. Stats don't substitute for thought. They are neither necessary nor sufficient for insight, though on some occasions they can usefully bring into focus things that are obscure. It should not be surprising that this process often fails. In fact, it should be surprising that it succeeds on occasion and that some areas (e.g. linguistics) have found pretty reliable methods for unearthing causal structure. We should expect this to be hard. The NYT piece makes it sound like we should be surprised that reported data are often wrong and it suggests that it is possible to do something about this by being yet more careful and methodologically astute, doing our stats more diligently. This, I believe, is precisely wrong. There is always room for improvement in one's methods. But methods are not what drive science. There is no method. There are occasional insights and when we gain some it provides traction for further investigation. Careful stats and methods are not science, though the reporting suggests that this is what many think it is, including otherwise thoughtful scientists.
Subscribe to:
Posts (Atom)