Colin Phillips has a very interesting discussion of his three year experiment with an open access journal that he, Matt Wagers and Claudia Felser edited in the Frontiers series (here). The online journal (here) has been quite successful and Colin does something very important and timely; he reflects on what went well, what went less well and WHY. In other words, he brings some first hand empirical experience to bear on a topic that we have discussed at FoL. The whole discussion is very interesting and I strongly recommend it for those interested in the topic.
A key point that Colin makes is that there are virtues other than cost in assessing how a journal is contributing to inquiry; readership, time from submission to publication, Impact Factor all matter, in addition to cost. Moreover, he notes that, interestingly, these are all factors that can be managed better or worse depending on what seem to be easily implementable procedures.
One that seems particularly noteworthy is the apparent fact that most of the submissions are accepted. This is not entirely true for there is a per-submission vetting process that insures that most of the submissions are of a kind to be accepted. However, the fact remains that a large proportion of the papers submitted get into print. The main reason seems to be because "[a]rticles are judged only for soundness, not for impact, i.e. if your study is sound but has minimal novelty or importance it can still be accepted." This makes the reviewing process less contentious and the goal posts easier to identify and hence cuts down on reviewing gamesmanship (though that is not how Colin puts it).
Let me make one observation and again encourage you to read the whole post because it is very good. One of the downsides of the current review process, IMO, is that it penalizes originality. How so? Well, original work is by its nature contentious and less well-formed than less original work. This is especially so when it comes to theory, where making things clear is very hard and then it is harder still to accommodate all the myriad problems that novelty will face. One way of finessing this problem is to publish most everything. This is what the above journal does (well, depending on how 'soundness' is evaluated: what makes a paper sound?). Another way is to to go the way that the old Cognition did: place your trust in editors (Mehler and Bever) with good taste and let them exercise it. As I have mentioned before, some of the greatest journals were run in this way (Keynes ran a journal as did Planck and both were outstanding). Now, I am not acquainted with Ms Felser, but I do know Colin and Matt quite well and I know them to have excellent taste in questions. So, I suspect that one reason for the success of their journal is this X factor. We need more of this. Not to replace the mainline journals, but to allow idiosyncrasy to sometimes get a hearing. We need editors that evaluate papers in terms of novelty and importance of the ideas. This is not quite impact (at least in the terms that Colin notes are relevant to this currently important metric), but it is a big deal, IMO. It can take time for a new idea to gain a foothold, but when it does, well, you can finish this sentence as well as I can.
So, look at the post. It's really thought provoking. I'd be interested to see comments on this in Colin's blog.
Showing posts with label Colin Phillips. Show all posts
Showing posts with label Colin Phillips. Show all posts
Monday, January 4, 2016
Thursday, April 17, 2014
The SMT again
Two revisions below thx to Ewan spring a thinko, a slip of the mind.
I have recently urged that we adopt a particular understanding of the Strong Minimalist Thesis (SMT) (here). The version that I favor treats the SMT as a thesis about systems that use grammars and suggests that central features of the grammatical representations that they use will be crucial to explaining why they are efficient. If this proves to be doable, then it is reasonable to describe FL and the grammars it makes available as “well designed” and “computationally efficient.” Stealing from Bob Berwick (here), I will take parsing efficiency to mean real time parsing and (real time) acquisition to mean easy acquisition given the PLD. Put this all together and the SMT is the conjecture that the grammatical format of Gs and UG is critical to allowing parsers, acquirers, producers, etc. to be very good at what they do (i.e. to be well-designed). On this view, grammars are “well designed” or “computationally efficient” in virtue of having properties that allow their users to be good at what they do when such grammars are embedded transparently within these systems.
I have recently urged that we adopt a particular understanding of the Strong Minimalist Thesis (SMT) (here). The version that I favor treats the SMT as a thesis about systems that use grammars and suggests that central features of the grammatical representations that they use will be crucial to explaining why they are efficient. If this proves to be doable, then it is reasonable to describe FL and the grammars it makes available as “well designed” and “computationally efficient.” Stealing from Bob Berwick (here), I will take parsing efficiency to mean real time parsing and (real time) acquisition to mean easy acquisition given the PLD. Put this all together and the SMT is the conjecture that the grammatical format of Gs and UG is critical to allowing parsers, acquirers, producers, etc. to be very good at what they do (i.e. to be well-designed). On this view, grammars are “well designed” or “computationally efficient” in virtue of having properties that allow their users to be good at what they do when such grammars are embedded transparently within these systems.
One particularly attractive virtue of this interpretation
(for me) is that I understand how I could go about empirically investigating
it. I confess that this is not true for
other versions of the SMT that talk about neat fits between grammar and the CI
interface, for example. So far as I can tell, we know rather little about the
CI interface and so the question of fit is, at best, premature. On the other
hand we do know a bit about how parsing works and how acquisition proceeds so
we have something to fit the grammar to.[1]
So how to proceed? In two steps I believe. The first is to
see if use systems (e.g. parsers) actually deploy grammars in real time, i.e.
as they parse. Thus, if it is true that the basic features of grammatical
representations are responsible for how (e.g.) parsers manage to efficiently do
what they do then we should find real time evidence implicating these representations
in real time parsing. Second, we should look for how exactly the implicated
features manage to make things so efficient. Thus, we should look for
theoretical reasons for why parsers that transparently embody, say, Subjacency
like principles, would be efficient. Let
me discuss each of these points in turn.
There is increasing evidence from psycho-ling research indicating
that real time parsing respects grammatical distinctions, even very subtle
ones. Colin Phillips is a leader in this
kind of work and he and his (ex) students (e.g. Brian Dillon, Matt Wagers,
Ellen Lau, Dave Kush, Masaya Yoshida) have produced a body of work that
demonstrates how very closely parsers respect grammatical conditions like
islands, c-command, and local binding domains. And by closely I mean very closely. So, for example, Colin shows (here)
that online parsing respects the grammatical conditions that license parasitic
gaps. So, not only do parsers respect islands, but they even treat
configurations where island effects are amnestied as if they were not islands.
Thus, parsers respect both the general conditions that grammars lay down
regarding islands and the exceptions
to these general conditions that grammars allow. This is what I mean by
‘close.’
There is a recent excellent demonstration of this from
Masaya Yoshida, Lauren Ackerman, Morgan Purier and Rebekah Ward (YLPW) (here
are slides from a recent CUNY talk).[2]
YLPW analyzes the processing of backward sluicing constructions like (1):
(1) I
don’t recall which writer, but the editor notified a writer about a new project
There is an ellipsis “gap” right after which writer that is redeemed by anchoring it to a writer in the following sentence. What
YLPW is looking to determine is whether the elided
gap site is sensitive to online
parsing effects. YLPW uses a plausibility effect as probe as follows.
First, it is well known that a wh in CP triggers an active search for a verb/gap that will give it
an interpretation. ‘Active’ here means that the parser uses a top down
predictive process and is eagerly looking to link the wh to a predicate without first consulting bottom information that
would indicate the link to be ill-advised. YLPW show that the eagerness to
“fill a gap” is as true for implicit gaps within ellipsis sites as it is for
“real” gaps in regular wh sentences. YLPW shows this by demonstrating a
plausibility effect slowdown in
sentences like (2a) parallel to the ones found in (2b):
(2) a. I
don’t remember which writer/which book,
but the editor notified a writer about a new book
b.
I don’t remember which writer/which book
the editor notified GAP about a new book
When the wh is which book then there is a significant
pause at notified in both sentences
in (2), as contrasted with the same sentences where which writer is the antecedent of the gap. This is because parsers, we know, greedily
try and relate the wh to the first
syntactically available position encountered and in the case of which book the wh is not a plausible filler of the gap and the attempted filling
results in a little lingering about the verb (*notify this book about…). If the antecedent is which writer no such pause occurs, for obvious reasons. The plausibility effect, then, is just a
version of the well-known filled gap effect, with a little semantic kicker to
add some frisson. At any rate, the first important discovery is that we find
the same plausibility effect in both (2a) with the gap inside a sluiced
ellipsis site, and (2b) where the gap is “overt.”
The next step is to see if this plausibility/filled gap
effect slowdown occurs when the relevant antecedent for the sluiced ellipsis
site is inside an island. It is well known that ellipsis is not subject to island restrictions.
Thus, if the parser cleaves tightly
to the distinctions the grammar makes (as the SMT would lead us to expect) then
we should find plausibility slowdowns except
when the gap is inside an ellipsis site for the latter are not subject to
island restrictions [and so should induce filled gap/plausibility effects (added: thx Ewan)]. And that’s exactly
what YLPW find. Though plausibility effects are not found at notified in cases like (3) they are found in cases like (4) where the
“gap” is inside a sluice sight.
(3) I
don’t remember which book [the editor
who notified the publisher about
some science book] had recommended to me
(4) I
don’t remember which book, but [the
editor who notified the publisher
about some science book] recommended a new book to me
This is just what we expect from a parser that transparently
embeds a UG like grammar that treats movement but not ellipsis as a product of
(long) movement.
The conclusion: it seems that parsers make just the
distinctions that grammars make when they parse in real time, just as the SMT
would lead us to expect.
So, there is growing evidence that parsers transparently
embed UG like grammars. This readies us
for the second step. Why should they do so?
Here, there is less current research
that bears on the issue. However, there is work from the 80s by Mitch Marcus,
Bob Berwick and Amy Weinberg that showed that a Marcus style parser that
incorporated grammatical features like Subjacency (and, interestingly, Extension)
could parse sentences efficiently (effectively, in real time). This is just what the doctor ordered. It goes
without saying (though I will say it) that this work needs updating to bear
more directly on the SMT and minimalist accounts of FL. However, it provides a
useful paradigm of how one might go
about connecting the discoveries concerning online parsing with computational questions
of parsing efficiency and their relationship to central architectural features
of FL/UG.
The SMT is a bold conjecture. Indeed, it is likely false, at
least in fine detail. This does not, however, detract from its programmatic
utility. The fact is that there is
currently lots of research that can be understood as bearing on its accuracy
and that fruitfully brings together work in syntax, psycholinguistics and
computational linguistics. The SMT, in
other words, is a terrific hypothesis that will generate fascinating work
regardless of its ultimate empirical fate.
That’s what we want from a research program and that’s something that
the Strong Minimalist Thesis is ready to deliver. Were this all that the
Minimalist Program provided, it would have been enough (dayenu!). There is more, but for the nonce, this is more than
enough. Yay, for the Minimalist Program!!!
[1]
Let me modulate this: we know something about some other parts, see here
for discussion of magnitude estimation in the visual domain. Note that this
discussion fits well with the version of the SMT deployed here precisely
because we know something about how this part of the visual system works. We
cannot say as much about most of the other parts of CI. Indeed, we don’t really
know how many “parts” CI has.
[2]
They are running some more experiments, so this work is not yet finished.
Nonetheless, it illustrates the relevant point well, and it is really fun
stuff.
Monday, February 24, 2014
DTC redux
Syntacticians have effectively used just one kind of probe
to investigate the structure of FL, viz. acceptability judgments. These come in
two varieties: (i) simple “sounds good/sounds bad” ratings, with possible
gradations of each (effectively a 6ish point scale ok, ?, ??, ?*, *, **), and
(ii) “sounds good/sounds bad under this interpretation” ratings (again with
possible gradations). This rather crude empirical instrument has proven to be
very effective as the non-trivial nature of our theoretical accounts indicates.[1]
Nowadays, this method has been partially systematized under the name
“experimental syntax.” But, IMO, with a few important conspicuous exceptions,
these more refined rating methods have effectively endorsed what we knew
before. In short, the precision has been useful, but not revolutionary.[2]
In the early heady days of Generative Grammar (GG), there
was an attempt to find other ways of probing grammatical structure.
Psychologists (following the lead that Chomsky and Miller (1963) (C&M)
suggested) took grammatical models and tried to correlate them with measures
involving things like parsing complexity or rate of acquisition. The idea was a
simple and appealing one: more complex grammatical structures should be more
difficult to use than less complex
ones and so measures involving language use (e.g. how long it takes to
parse/learn something) might tell us something about grammatical structure.
C&M contains the simplest version of this suggestion, the now infamous
Derivational Theory of Complexity (DTC). The idea was that there was a
transparent (i.e. at least a homomorphic) relation between the rules required
to generate a sentence and the rules used to parse it and so parsing complexity
could be used to probe grammatical structure.
Though appealing, this simple picture can (and many believed
did) go wrong in very many ways (see Berwick and Weinberg 1983 (BW) here
for a discussion of several).[3]
Most simply, even if it is correct that there is a tight relation between the
competence grammar and the one used for parsing (which there need not be, though
in practice there often is, e.g. the Marcus Parser) the effects of this algorithmic complexity need not show up
in the usual temporal measures of complexity, e.g. how long it takes to parse a
sentence. One important reason for this is that parsers need not apply their
operations serially and so the supposition that every algorithmic step takes
one time step is just one reasonable assumption among many. So, even if there is a strong transparency
between competence Gs and the Gs parsers actually deploy, no straightforward measureable
time prediction follows.
This said, there remains something very appealing about DTC
reasoning (after all, it’s always nice
to have different kinds of data converging on the same conclusion, i.e. Whewell’s
consilience) and though it’s true that the DTC need not be true, it might be worth looking for places where the
reasoning succeeds. In other words, though the failure of DTC style reasoning
need not in and of itself imply
defects in the competence theory used, a successful DTC style argument can tell
us a lot about FL. And because there are many ways for a DTC style explanation
to fail and only a few ways that it can succeed, successful stories if they exist can shed interesting light
on the basic structure of FL.
I mention this for two reasons. First, I have been reading
some reviews of the early DTC literature and have come to believe that its
demonstrated empirical “failures” were likely oversold. And second, it seems
that the simplicity of MP grammars has made it attractive to go back and look
for more cases of DTC phenomena. Let me elaborate on each point a bit.
First, the apparent demise of the DTC. Chapter 5 of Colin
Phillips’ thesis (here)
reviews the classical arguments against the DTC. Fodor, Bever and Garrett (in their 1974 text)
served as the three horsemen of the DTC apocalypse. They interned the DTC by
arguing that the evidence for it was inconclusive. There was also some
experimental evidence against it (BW note the particular importance of Slobin
(1966)). Colin’s review goes a very long way in challenging this pessimistic
conclusion. He sums up his in depth review as follows (p.266):
…the received view that the
initially corroborating experimental evidence for the DTC was subsequently discredited
is far from an accurate summary of what happened. It is true that some of the
experiments required reinterpretation, but this never amounted to a serious
challenge to the DTC, and sometimes even lent stronger support to the DTC than
the original authors claimed.
In sum, Colin’s review strongly implies that linguists
should not have abandoned the DTC so quickly.[4]
Why, after all, give up on an interesting hypothesis, just because of a few
counter-examples, especially ones that when considered carefully seem on the
weak side? In retrospect, it looks like the abandonment of the strong
hypothesis was less a matter of reasonable retreat in the face of overwhelming
evidence than a decision that disciplines occasionally make to leave one
another alone for self-interested reasons. With the demise of the DTC,
linguists could assure themselves that they could stick to their investigative
methods and didn’t have to learn much psychology and psychologists could
concentrate on their experimental methods and stay happily ignorant of any
linguistics. The DTC directly threatened this comfortable “live and let live”
world and perhaps this is why its demise was so quickly embraced
by all sides.
This state of comfortable isolation is now under threat,
happily. This is so for several reasons.
First, some kind of DTC reasoning is really the only game in town in cog-neuro.
Here’s
Alec Marantz’s take:
...the “derivational theory of
complexity” … is just the name for a standard methodology (perhaps the dominant
methodology) in cognitive neuroscience (431).
Alec rightly concludes that given the standard view within GG
that what linguists describe are real mental structures, there is no choice but
to accept some version of the DTC as the null hypothesis. Why? Because, ceteris paribus:
…the more complex a representation-
the longer and more complex the linguistic computations necessary to generate
the representation- the longer it should take for a subject to perform any task
involving the representation and the more activity should be observed in the
subject’s brain in areas associated with creating or accessing the
representation or performing the task (439).
This conclusion strikes me as both obviously true and
salutary, with one caveat. As BW has shown us, the ceteris paribus clause can in practice be quite important. Thus, the common indicators of complexity
(e.g. time measures) may be only indirectly related to algorithmic complexity.
This said, GG is (or should be) committed to the view that algorithmic
complexity reflects generative complexity and that we should be able to find
behavioral or neural correlates of this (e.g. Dehaene’s work (discussed here)
in which BOLD responses were seen to track phrasal complexity in pretty much a
linear fashion or Forster’s work finding temporal correlates mentioned in note
4).
Alec (439) makes an additional, IMO correct and important,
observation. Minimalism in particular, “in denying multiple routes to
linguistic representations,” is committed to some kind of DTC thinking.[5]
Furthermore, by emphasizing the centrality of interface conditions to the
investigation of FL, Minimalism has embraced the idea that how linguistic
knowledge is used should reveal a great deal about what it is. In fact, as I’ve
argued elsewhere, this is how I would like to understand the “strong minimalist
thesis,” (SMT) at least in part. I have suggested that we interpret the SMT as
committed to a strong “transparency hypothesis” (TH) (in the sense of Berwick
& Weinberg), a proposal that can only be systematically elaborated by how
linguistic knowledge is used.
Happily, IMO, paradigm examples of how to exploit “use” and
TH to probe the representational format of FL are now emerging. I’ve already
discussed how Pietroski, Hunter, Lidz and Halberda’s work relates to the SMT
(e.g. here
and here).
But there is other stuff too of obvious relevance: e.g. BW’s early work on
parsing and Subjacency (aka Phase Theory) and Colin’s work on how islands are
evident in incremental sentence processing. This work is the tip of an
increasingly impressive iceberg. For example, there is analogous work showing
that that parsing exploits binding restrictions incrementally during processing
(e.g. by Dillon, Sturt, Kush).
This latter work is interesting for two reasons. It validates
results that syntacticians have independently arrived at using other methods
(which, to re-emphasize, is always worth doing on methodological grounds). And,
perhaps even more importantly, it has started raising serious questions for
syntactic and semantic theory proper. This is not the place to discuss this in
detail (I’m planning another post dedicated to this point), but it is worth
noting that given certain reasonable assumptions about what memory is like in
humans and how it functions in, among other areas, incremental parsing, the
results on the online processing of binding noted above suggest that binding is
not stated in terms of c-command but some other notion that mimics its effects.
Let me say a touch more about the argument form, as it is
both subtle and interesting. It has the following structure: (i) we have
evidence of c-command effects in the domain of incremental binding, (ii) we
have evidence that the kind of memory we use in parsing cannot easily code a
c-command restriction, thus (iii) what the parsing Grammar (G) employs is not
c-command per se but another notion
compatible with this sort of memory architecture (e.g. clausemate or
phasemate). But, (iv) if we adopt a strong SMT/TH (as we should), (iii) implies
that c-command is absent from the competence G as well as the parsing G. In
short, the TH interpretation of SMT in this context argues in favor of a
revamped version of Binding Theory in which FL eschews c-command as a basic
relation. The interest of this kind of argument should be evident, biut let me
spell it out. We S-types are starting to face the very interesting prospect
that figuring out how grammatical information is used at the interfaces will help us choose among alternative competence theories by placing interface
constraints on the admissible primitives. In other words, here we see a
non-trivial consequence of Bare Output Conditions on the shape of the grammar.
Yessss!!!
We live in exciting times. The SMT (in the guise of TH)
conceptually moves DTC-like considerations to the center of theory evaluation.
Additionally, we now have some useful parade cases in which this kind of reasoning
has been insightfully deployed (and which, thereby, provide templates for
further mimicking). If so, we should expect that these kinds of considerations
and methods will soon become part of every good syntactician’s armamentarium.
[1]
The fact that such crude data can be used so effectively is itself quite remarkable.
This speaks to the robustness of the system being studied for such weak signals
should not be expected to be so useful otherwise.
[2]
Which is not to say that such more careful methods don’t have their place.
There are some cases where being more careful has proven useful. I think that
Jon Sprouse has given the most careful thought to these questions. Here
is an example of some work where I think that the extra care has proven to be
useful.
[3]
I have not been able to find a public version of the paper.
[4]
BW note that Forster provided evidence in favor of the DTC even as Fodor et.
al. were in the process of burying it. Forster effectively found temporal
measures of psychological complexity that tracked the grammatical complexity
the DTC identified by switching the experimental task a little (viz. he used an
RSVP presentation of the relevant data).
[5]
I believe that what Alec intends here is that in a theory where the only real
operation is merge then complexity is easy to measure and there are pretty
clear predictions of how this should impact algorithms that use this
information. It is worth noting that the heyday of the DTC was in a world where
complexity was largely a matter of how many transformations applied to derive a
surface form. We have returned to that world again, though with a vastly
simpler transformational component.
Friday, September 13, 2013
Acceptability Judgements
As many of you know, there has been a debate in the literature lately about the reliability of acceptability judgements as used by linguists. We (i.e. me) use rather informal methods in our data collection. It often amounts to little more than asking a dozen or so colleagues about a given sentence contrast, e.g. how does A sound compared to B, can A have interpretation A'? At any rate, the reliability of these kinds of data gathering methods has been questioned, with the suggestion that the whole Generative enterprise is an insubstantial house of cards built on sand without a leg to stand on. Last week in Potsdam, there was a workshop dedicated to the question of understanding acceptability judgements and Colin Phillips, one of the presenters, circulated his slides to the UMD department at large. The slides review a large number of different kinds of judgement studies and conclude that the dad from laymen and experts by and large support the conventional linguistic wisdom as regards these data in the vast majority of cases (roughly the conclusion that Sprouse and Almeida and Schutze have come to as well) to a very high degree of reliability. The compendium of results covers 1770 subjects in over 50 different kinds of experiments. Interestingly, the aim of this work was not to test the reliability of judgements but to norm materials for other kinds of psycho-linguistic investigations. The money slides are #31-#33, which consolidate the relevant findings. The rightmost column, a ordered pair of a number and a Yes/No, e.g. 35 Yes, indicate the size of the number of testees and whether the results coincide with the standard wisdom, for which, as indicated, there is a pretty good support.
Just as interesting are the cases where the support is not as robust. These tend to involve pronominal binding data (which, I have no trouble believing are more problematic based on my own judgements much of the time). Yet more interesting is that the data gets cleaner depending on the type of judgement demanded. It seems that forced choice (it's good vs it's bad) data yields the cleanest results (this is both in Colin's slides and was independently noted by Sprouse and Almeida in some of their work), cleaner than the magnitude estimation or 7 point rating scale measures that are all the rage now.
At any rate, Colin's slides are worth looking at and it is to be hoped that the other papers/materials that were presented in Potsdam can be soon made available to us all.
Just as interesting are the cases where the support is not as robust. These tend to involve pronominal binding data (which, I have no trouble believing are more problematic based on my own judgements much of the time). Yet more interesting is that the data gets cleaner depending on the type of judgement demanded. It seems that forced choice (it's good vs it's bad) data yields the cleanest results (this is both in Colin's slides and was independently noted by Sprouse and Almeida in some of their work), cleaner than the magnitude estimation or 7 point rating scale measures that are all the rage now.
At any rate, Colin's slides are worth looking at and it is to be hoped that the other papers/materials that were presented in Potsdam can be soon made available to us all.
Friday, April 19, 2013
One FInal Time into the Breach (I hope): More on the SMT
The discussion of the SMT posts has gotten more abstract
than I hoped. The aim of the first post discussing the results by Pietroski,
Lidz, Halberda and Hunter was to bring the SMT down to earth a little and
concretize its interpretation in the context of particular linguistic investigations. PLHH investigate the following: there are
many ways to represent the meaning of most,
all of which are truth functionally equivalent. Given this, are the representations empirically equivalent or are
there grounds for arguing choosing one representation over the others. PLHH
propose to get a handle on this by investigating how these representations are used by the ANS+visual system in
evaluating dot scenes wrt statements like most
of the dots are blue. They discover that the ANS+visual system always uses
one of three possible representations to evaluate these scenes even when use of the others would be both
doable and very effective in that context. When one further queries the
core computational predilections of the ANS+visual system it turns out that the
predicates that it computes easily coincide with those that the “correct”
representation makes available. The conclusion is that the one of the three
representations is actually superior to the others qua linguistic representation of the meaning of most, i.e. it is the linguistic meaning of most. This all fits rather well with the SMT. Why?
Because the SMT postulates that one way of empirically evaluating candidate
representations is with regard to their fit
with the interfaces (ANS+visual) that use it. In other words, the SMT bids us
look to how grammars fit with interfaces and, as PLHH show, if one understands
‘fit’ to mean ‘be transparent with’ then one meaning trumps the others when we
consider how the candidates interact with the ANS+visual system.
It is important to note that things need not have turned out
this way empirically. It could have been the case that despite core capacities
of the ANS+visual system the evaluation procedure the interface used when evaluating
most sentences was highly context
dependent, i.e. in some cases it used the one-to-one strategy, in others the
‘|dots ∩ blue| - |dots ∩ not-blue|’ strategy and sometimes the ‘|dots ∩ blue| - [|dots| - |dots ∩ blue|]’ strategy. But, and this is important,
this did not happen. In all cases the
interface exclusively used the third option, the one that fit very snugly with
the basic operations of the ANS+visual system. In other words, the
representation used is the one that the SMT (interpreted as the Interface
Transparency Thesis) implicates. Score one for the SMT.
Note that the argument puts together various strands: it
relies on specific knowledge on how the ANS+visual system functions. It relies
on specific proposals for the meaning of most
and given these it investigates what happens when we put them together. The kicker is that if we assume that the
relation between the linguistic representation and what the ANS+visual system
uses to evaluate dot scenes is “transparent” then we are able to predict[1]
which of the three candidate representations will in fact be used in a
linguistic+ANS+visual task (i.e. the task of evaluating a dot scene for a given
most sentence[2]).[3]
The upshot: we are able to use information from how the
interface behaves to determine a property of a linguistic representation. Read that again slowly: PLHH argue that
understanding how these tasks are accomplished provides evidence for what the
linguistic meanings are (viz. what the correct representations of the meanings
are). In other words, experiments like this bear on the nature of linguistic
representations and a crucial assumption in tying the whole beautiful package
together is the SMT interpreted along the lines of the ITT.
As I mentioned in the first post on the SMT and Minimalism
(here), this is not the only exemplar of the SMT/ITT in action. Consider one
more, this time concentrating on work by Colin Phillips (here). As previously
noted (here), there are methods for tracking the online activities of parsers.
So, for example, the Filled Gap Effect (FGE) tracks the time course of mapping
a string of words into structured representations. Question: what rules do parsers use in doing
this. The SMT/ITT answer is that parsers use the “competence” grammars that
linguists with their methods investigate. Colin tests this by considering a very complex instance: gaps within
complex subjects. Let’s review the argument.
First some background.
Crain and Fodor (1985) and Stowe (1986) discovered that the online
process of relating a “filler” to its “gap” (e.g. in trying to assign a Wh a
theta role by linking it to its theta assigning predicate) is very eager. Parsers try to shove wayward Whs into
positions even if filled by another DP.
This eagerness shows up behaviorally as slowdowns in reading times when
the parser discovers a DP already homesteading in the thematic position it
wants to shove the un-theta marked DP into. Thus in (1a) (in contrast to (1b),
there is a clear and measurable slowdown in reading times at Bill because it is a place that the who could have received a theta role.
(1) a.
Who did you tell Bill about
b.
Who did you tell about Bill
Thus, given the parser’s eagerness, the FGE becomes a probe
for detecting linguistic structure built online. A natural question is where do
FGEs appear? In other words, do they “respect” conditions that “competence”
grammars code? BTW, all I mean by
‘competence grammars’ are those things that linguists have proposed using their
typical methods (one’s that some Platonists seem to consider the only valid
windows into grammatical structure!)? The
answer appears to be they do. Colin reviews the literature and I refer you to
his discussion.[4] How do FGEs show that parsers respect
grammatical structure? Well, they seem not
to apply within islands! In other words, parsers do not attempt to related Whs to gaps within islands. Why? Well given
the SMT/ITT it is because Whs could not have moved from positions wihin islands
and so they are not potential theta
marking sites for the Whs that the parser is eagerly trying to theta mark. In
other words, given the SMT/ITT we expect parser eagerness (viz. the FGE) to be
sensitive to the structure of grammatical representations, and it seems that it
is.
Observe again, that this is not a logical necessity. There
is no a priori reason why the grammars
that parsers use should have the properties that linguists have postulated,
unless one adopts the SMT/ITT that is. But let’s go on discussing Colin’s paper
for it gets a whole lot more subtle than this. It’s not just gross properties
of grammars that parsers are sensitive to, as we shall presently see.
Colin consider gaps within two kinds of complex subjects.
Both prevent direct extraction of a Wh (2a/3a), however, sentences like (2b)
license parasitic gaps while those like (3b) do not:
(2) a.
*What1 did the attempt to repair t1 ultimately damage the
car
b.
What1 did the attempt to repair t1 ultimately damage t1
(3) a.
*What1 did the reporter that criticized t1 eventually
praise the war
b. *What did the reporter that criticized
t1 eventually praise t1
So the grammar allows gaps related to extracted Whs in (2b)
but not (3b), but only if this is a parasitic gap. This is a very subtle set of grammatical
facts. What is amazing (in my view
nothing short of unbelievable) is that the parser respects these parasitic gap licensing conditions. Thus, what Colin shows is that we find FGEs at
the italicized expressions in (4a) but not (4b):
(3) a.
What1 did the attempt to repair the
car ultimately …
b.
What1 did the reporter that
criticized the war eventually …
This is a case where the parser is really tightly cleaving to distinctions that the grammar makes. It
seems that the parser codes for the possibility of a parasitic gap while
processing the sentence in real time.
Again, this argues for a very transparent relation between the
“competence” grammar and the parsing grammar, just as the SMT/ITT would
require.
I urge the interested to read Colin’s article in full. What
I want to stress here is that this is another concrete illustration of the
SMT. If
grammatical representations are optimal realizations of interface conditions
then the parser should respect the distinctions that grammatical
representations make. Colin presents evidence that it does, and does so very
subtly. If linguistic representations are used
by interfaces, then we expect to find this kind of correlation. Again, it is
not clear to me why this should be true given certain widely bruited Platonic
conceptions. Unless it is precisely these
representations that are used by the parser, why should the parser respect its
dicta? There is no problem understanding
how this could be true given a standard mentalist conception of grammars. And
given the SMT/ITT we expect it to be true. That we find evidence in its favor
strengthens this package of assumptions.
There are other possible illustrations of the SMT/ITT. We should develop a sense of delight at
finding these kind of data. As Colin’s stuff shows, the data is very complex
and, in my view, quite surprising, just like PLHH’s stuff. In addition, they
can act as concrete illustrations of how to understand the SMT in terms of
Interface Transparency. An added bonus
is that they stand as a challenge to certain kinds of Platonist conceptions, I
believe. Bluntly: either these
representations are cognitively available or we cannot explain why the
ANS+visual system and the parser act as if they were. If Platonic
representations are cognitively (and neurally, see note 3) available, then they
are not different from what mentalists have taken to be the objects of study
all along. If from a Platonist perspective they are not cognitively (and
neurally) available then Platonists and mentalists are studying different
things and, if so, they are engaged in parallel rather than competing
investigations. In either case, mentalists need take heed of Platonist results
exactly to the degree that they can be reinterpreted mentalistically.
Fortunately, many (all?) of their results can be so interpreted. However, where this is not possible, they would be of absolutely no interest to the project of describing linguistic competence. Just metaphysical curiosities
for the ontologically besotted.
[1]
Recall, as discussed here, ‘predict’ does not mean ‘explain.’
[2]
Remember, absent the sentence and in
specialized circumstances the visual system has no problem using strategies
that call on powers underlying the other two non-exploited strategies. It’s
only when the visual system is combined with the ANS and with the linguistic most sentence probe that we get the
observed results.
[3]
Actually, I overstate things here: we are able to predict some of the properties of the right representation, e.g. that it
doesn’t exploit negatively specified predicates or disjunctions of predicates.
[4]
Actually, there are several kinds of studies reviewed, only some of which
involve FGEs. Colin also notes EEG studies that show P600 effects when one has
a theta-undischarged Wh and one crosses into an island. I won’t make a big deal
out of this, but there is not exactly a dearth of neuro evidence available for
tracking grammatical distinctions. They
are all over the place. What we don’t have are good accounts of how brains implement grammars. We have
tons of evidence that brain responses track grammatical distinctions, i.e. that
brains respond to grammatical structures. This is not very surprising if you
are not a dualist. After all we have endless amounts of behavioral evidence
(viz. acceptability judgments, FGEs, eye movement studies, etc.) and on the
assumption that human behavior supervenes on brain properties it would be
surprising if brains did not distinguish what human subjects distinguish
behaviorally. I mention this only to state the obvious: some kinds of Platonism
should find these kinds of correlations challenging. Why should brains track
grammatical structure if these live in Platonic heavens rather than
brains? Just asking.
Subscribe to:
Posts (Atom)