Sorry for the delay in getting David's lectures up and available. Here are the first two. The third does not exist due to technical problems. Or, to be more accurate, it does not exist in this world.
Interestingly, this third lecture cast aspersions on sidewards movement and though it is true that I have a fondness for this operation and would do almost anything to defend its viability and utility, you can rest assured that the absence of this third scurrilous lecture (the one attacking SWM) met technical difficulties for entirely innocent reasons.
Moreover, I am currently trying to recover this lecture. I have asked my team of semantics consultants to locate it for me. They tell me that it does exist in some possible worlds, though whether those worlds are accessible from ours is currently under intense technical investigation. They are hopeful (or some of their counterparts are). If anyone has suggestions of how to access the third lecture in those worlds and make it available in ours, please contact me asap. Until then, here and here are the first two.
Thx to Julia Buffinton for uploading the files and getting me the links.
Thursday, April 7, 2016
Charlie Jones
I just heard from an old student that Charlie Jones died this past week. He wrote a great thesis on purpose clauses (Mass) that subsequently became a book (here). He also was a gifted designer who delighted in producing accurate but completely unusable calendars. I got one very year in January and puzzled over them affectionately. Good guy. Good Linguist.
Yang on Bayes 2
Here
is part 2. I have included the previous paragraph so that you can get a running
start into the discussion. Let me remind you that Charles is culpable for the
content of what follows only insofar as reading his stuff stimulated me to
think as I did. In other words, very culpable.
It is
worth noting that Bayes makes these substantive assumptions for principled reasons. Bayes’ started out
as a normative theory of rationality. Bayes was developed as a formalization of
the notion “inference to the best explanation.” In this context the big
hypothesis space, full data, full updating, optimizing rule structure noted
above are reasonable for they
undergird a the following principle of rationality: choose that theory which
among all possible theories best matches the full range of possible data. This
makes sense as a normative principle of inductive rationality. The only caveat
is that Bayes seems to be too demanding for humans. There are considerable
computational costs to Bayesian theories (again CY notes these) and it is
unclear how good a principle of rationality is if humans cannot apply it.
However, whatever its virtues as a normative theory, it is of dubious value in
what L.J. Savage (one the founders of modern Bayesianism) termed “small world”
situations (SWS).
What
are SWSs? They are situations where the hypothesis space, the space of options
over which the probabilities are defined, is small. In SWSs the central problem
is to describe the structure of the “small world.” All else is secondary. So in SWSs useful idealizations will help
focus on these structures. And that is the problem with Bayes. As Pat Suppes notes
(CY quotes from this paper p. 47, my emphasis NH)):
…any
theory of complex problem solving cannot go far simply on the basis of Bayesian
decision notions of information processing. The core of the problem is that of
developing an adequate psychological theory to describe, analyze and predict
the structure imposed by organisms on the bewildering complexities of possible
alternatives facing them. The simple concept of an a priori distribution over
these alternatives is by no means sufficient and does little toward offering a solution to any complex problem.
…understanding
the structures actually used is important for an adequate descriptive theory of
behavior….As…inductive logic comes to grips with more realistic problems, the overwhelming combinatorial possibilities
that arise in any complex problem will make the need for higher-order
structural assumptions self-evident.
In
short, the hard problem is to find the structure of the SWSs (to describe the
restricted hypothesis space) and Bayes brings little (nothing?) of relevance to
the table for solving this problem. Bayes
does not prevent considerations of this problem, but it does not promote them
or make them central concerns either. And I would argue that by not recognizing
(or, worse, radically deemphasizing) the SWS character of most psychological
processes, Bayes deflects attention from the hard problem, the one that must be
solved if there is to be any progress, onto secondary concerns. That’s the
standard cost of a misidealization; it leads you to look in the wrong place.
Let me
put this another way, echoing Suppes and CY. The hard problem is “the
overwhelming combinatorial possibilities that arise in any complex problem.”
The Bayes idealization abstracts away from this at every step: its default
assumptions are big hypothesis spaces, large data, panoramic update and
optimization. Each of these causes problems. If the cognitive reality (at least
in language, but I suspect everywhere) is small worlds, very limited data,
myopic (and dumb) updating and satisficing decision then Bayes is just
misleading. Moreover, starting where Bayes does means that the bulk of the work
will not be done by the Bayes part of any account, but by the algorithms that
massage these assumptions out. And indeed, this is what we find when Bayesians
respond to their critics. Here’s an example of this move that CY discusses.
CY
notes that when confronted with problems Bayesians usually go Marrian. So, for
example, when it is observed that Bayes is inconsistent with probability
matching, the retort is that this is only so for the level 1 theory. The level
2 algorithms allow for probability matching when certain further constraints
are added to the optimizing function. However, if the critique of Bayes is that
the idealization is the wrong one, then the fact that one can patch up the
problem by saying more is not so much a sign of health but possible
confirmation of the initial wrong turn. Of course you can patch anything up
with further machinery. This is never in doubt (or almost never). The question
is not whether this is possible, but whether the patching up is consistent with
the spirit of the initial idealization. If it is, then no problem. If it isn’t,
then it argues against the idealization. CY argues that how Bayes handles probability matching goes against the Bayesian
spirit of the initial idealization (see the CY discussion of Steyvers et al p
12).
There
are many critiques of Bayes that go along the same lines. Many note that Bayes
stories have an ad hoc flavor to them (see here and references in linked paper). One critic, Clark
Glymour (who, BTW, is no schnook (see here)), has
described the work as “Ptolemaic” (see here) and not in a good way (is there a good way?). In the
cases discussed, the ad hoc flavor comes from specific assumptions added to the basic Bayes account to
accommodate the recalcitrant data, or so the critics argue. It is above my pay grade
to evaluate whether the critics are correct in each case. However, this is what
we expect if the basic Bayes account is based on a misidealization of the
relevant problems.
In fact,
this is precisely how we now understand the failure of Ptolemaic astronomy. Think
epicycles. Why the epicyles? Because the whole Ptolemaic project starts from
two completely wrong idealizations; that the universe is geocentric and that
orbits are circular. Starting here, the astronomical data demands epicycles and
equants and all the other geometric paraphernalia. Change to a heliocentric
system and allow elliptical orbits and all this apparatus disappears. The ad hoc apparatus is inevitable given the
initial idealization of the problem.
Given the starting assumptions it is possible to “model” all planetary
motion (in fact, we now know that any possible set of orbits can me modeled).
However, it is equally clear that in this case tracking the data is not a virtue. Wrong idealization, wrong
theory no matter how well it can be made to fit the data.
Let me
end this (again) overly long post, with a few observations.
First,
arguing against an idealization is very difficult. Why? Because it is often
possible to “fix” the problems that a misidealization generates by adding
further bells and whistles. Thus, it is not generally possible to argue against
an idealization empirically. Rather the argument is necessarily more subtle:
the accounts delivered are non explanatory, they involve too much ad hoc machinery, they are too complex,
they don’t fit well with other things we want etc. Despite their difficulty,
critiques of idealizations are one of the most important kinds of scientific
arguments. They are hardly ever dispositive, but they are extremely important
precisely because they expose the basic lay of the land and identify the hard
problems that need solutions. They are also the kinds of arguments that can
make you a scientific immortal (think Newton on contact mechanics and Einstein
in the aether).
Second,
CY presents an interesting Marr like argument against Bayes. It proposes that
level 1 theories that have transparent relations to level 2 accounts are better
than those that don’t (see the prior three posts on Marr for some discussion). CY
argues that Bayes level 1 and 2 theories are conceptually caliginous because of
the initial misidealizaation. Starting from the wrong place will make it hard
to transparently relate level 1 computational theories with level 2 theories of
algorithms and representations. CY makes such an argument as follows.
It is
well known that Bayes procedures taken transparently are computationally
intractable. CY observes that to date, even the non-transparent ones that have
been proposed are pretty bad (i.e. very slow). This is not a surprise in a
Marrian setting if the idealization
pursued by Bayes is wrongheaded. Indeed, in a Marr setting one might argue
that not being able to generate “nice” algorithms (transparent ones being
particularly nice) is a sign that the computational level theory is heading in
the wrong direction. The fact that CY’s proposed algorithm is both tractable
and easily implementable and that Bayesian proposals are not is a strong
argument against the adequacy of the latter idealization in a Marr setting
given the Marrian desideratum that computational level theories inform and
constrain level 2 algorithmic theories.
CY’s
argument goes further. CY’s proposed theory provides a dynamics for
acquisition: it not only can tell you how the data determines the G you end up
with, but it also tells you how to traverse the set of alternatives available
as the data comes in. So, the theory tells you how hypotheses change over real
time. Bayes does not do this. It just specifies where you end up given the data
available. Bayes is iterative, but not incremental. CY’s theory is both.
Interestingly,
CY derives this dynamic incrementality via a transparency assumption regarding
the Elsewhere Principle. This is something that a Bayes account cannot possibly
achieve precisely because it is well known to be intractable when understood
transparently. In other words, CY can do more than Bayes because its
understanding of the level 1 problem allows for a very neat/transparent mapping
to a level 2 account. This is unavailable for Bayes because the transparency
assumption in a Bayes theory leads to intractable computations. So, wrong
idealization, less transparency and so no obvious dynamics.
Third,
the Bayes idealization embodies a confusion linguists should be familiar with. Remember
the child as little linguist hypothesis?[1] It’s the idea that the way
to think of how LADs acquired Gs is on a par with how linguists discover the
properties of particular Gs. This, we know, is wrong on at least two grounds:
first, linguists when they operate consider a wide range of theoretical
alternatives for any given "problem" and second linguistic consider
full linguistic data to solve their problems while children only consider a
small subset of potentially useful data, the PLD. Thus, in two important ways,
children inhabit a “smaller world” than linguists do and consider much scantier
data in reaching their conclusions than linguists do.
Bayes
and the child-as-little-linguist hypothesis make the same mistake. LADs are
entirely unlike linguists, and a good thing too. The difference is that the
LADs hypothesis space of possible Gs given PLD is very small. In other words,
they inhabit a small linguistic world. Linguists qua scientists do not. The aim
of linguistics is to figure out what this small world looks like. Sadly,
linguists’ access to this small world is not aided by the same built in
advantages that the LAD comes equipped with, so it’s standard inference to the
best explanation for GGers. That’s why for us the process is slow and laborious
while kids acquire their Gs roughly by the age of 5.
So, both
Bayes and the child-as-little-linguist hypothesis mis-describe the main acquisition
problem by assimilating it to a rational decision problem thereby abstracting
away from the critical features of situation: rich small hypothesis space,
sparse, misleading and uninformative data, and, quite likely, no optimizing
(see CY p 23 on the principle of sufficiency).
Let me
put this one more way: the inferences that LADs make in acquiring their Gs are not rational. They consider too few
possible hypotheses and utilize very little data in getting to the G they
choose. True LADs use data and choose among Gs but the problem is not really
like an ideal inductive procedure. Moreover, the ways that it is not like an
ideal inductive procedure are what allow the whole procedure to successfully
apply. It is precisely by radically narrowing the hypothesis space that it is
possible for the LAD to choose the “right” G despite lousy data. The process,
it seems, need not be rational to be effective. Just the opposite one might say.
Fourth,
CY has a very good discussion of the likely irrelevance of the subset principle
to acquisition. This is important (and maybe I will return to it sometime in
the future) because the fact that Bayes can derive the subset principle is
taken to be a feather in its cap. Roughly speaking, if Bayes then subset
principle and because subset principle therefore good for Bayes. But, CY argues
that there is good reason to think that the subset principle is not operataive
in acquisition. If so, no argument for Bayes.
Note
that this is interesting regardless of whether Bayes is right. The subset
principle is often invoked so seeing how problematic it is has its own charms.
BTW, one of the big problems with the subset principle concerns its
tractability. It is computationally very complex as CY notes. If so, if Bayes derives it, it might not be good news
for Bayes.
Ok,
enough. Run and get CY. It’s an excellent paper and very important. If nothing
else, I hope that it focuses the debate over Bayes. The problem is not the details,
but the idealization. The Bayes stance runs against what linguists believe to
be the basic lay of the cognitive land. If we are right, then they are likely
wrong. This is a place where the different starting points are worth
sharpening. CY shows us how to do this, and that’s why it is such an important
paper.
[1]
See the following: Valian,
V., Winzemer, J., & Erreich, A. (1981). A little linguist model of syntax
learning. Language acquisition and linguistic theory, 188-209.
Tuesday, April 5, 2016
Yang on Bayes 1
This
is the first of two posts on a recent paper by Charles Yang. Once again, the
topic got away from me. I break it down into two parts to prevent you running
away too scared to even look. I suspect that my clever maneuver won’t help. But I
try.
One
more caveat: this is my understanding
of Charles’ paper. He should only be held responsible for what I say because he
allowed me to read it, and that might be a culpable act.
Charles
Yang has a new paper (CY) forthcoming in Language Acquisition and it is a must read. Aside from sketching a
very powerful critique of contemporary Bayesianism as applied to linguistic
problems (I will return to this critique momentarily), it also does something
that I never thought I would witness in my lifetime; it makes quantitative predictions in a linguistic
domain. And by “quantitative” I do not
mean giving p-values or confidence intervals. I mean numerical predictions
about the size of measurable effect. That’s quantitative! So, run, don’t walk
to the above link and read the damn thing!
Ok,
now that I’ve discharged my kudosing responsibilities, I want to discuss a
second feature of CY. It offers an excellent critique of current Bayesian
practice as applied to linguistic issues. As you all know, Bayes is a big
player nowadays. In fact, I am tempted to say that Bayes has stepped into the
position that Connectionism once held as the default theoretical framework in
the cognitive sciences in general and the cognition of language in particular. As
you may also recall, I have expressed reservations about Bayes and what it
brings to the linguistics table (see here, here, here, here). However, it was not until I read CY that I clearly understood
what bugs me about the Bayes framework as applied to my little domain of
interests. I want to “share” this aha moment with you.
The
conclusion can be put briskly as follows: Bayes is not so much wrong as
wrong-headed. From a Generative Grammar (GG) perspective, it’s not the details
that are off the mark (though they can be) but the idealization implicit in the framework that is (i.e. if you accept
as I do that GG has correctly identified the “computational” problems (see here) then Bayes is of little relevance and maybe worse).
And because this is so, there is very little we can learn from Bayesian
modelings of linguistic problems. There is both (i) not enough there there and (ii)
what there there is points in the wrong direction.
Let me
put this point another way: all theories idealize. Empirically successful
theories built on good idealizations.[1] CY’s argument is that Bayes
is a bad idealization when applied to matters linguistic. If it is right (and,
surprise surprise, I believe it is) then Bayes not only adds little, it
positively misleads and misdirects. Why? Because bad idealizations cannot be
empirically redeemed. It’s ok to be wrong (you can build on error given a good
framing of the problem). But if a theory is wrongheaded it will impede
progress. Whereas an adequate conception
of the problem (which is what idealizations embody) tolerates empirical
missteps. You can’t data your way out of a misconceived idealization. Why not?
Because if you’ve got the problem wrong then data coverage will come with a big
price, ad hoc assumptions whose main purpose is to make up for the basic misconception.
Hence, a misframing of the problem leads to explanatory sterility and misplaced
efforts. That’s the claim. Let’s see how
Bayes, misidealizes when considered form a GG perspective.
A
Bayes model consists of 3 moving parts: (i) a hypothesis space delimiting the
range of options (possibly weighted), (ii) a specification of the data relevant
to the choosing of the right hypothesis, (iii) an update rule saying how to
evaluate a given hypothesis given certain evidence. This 3-step procedure can
iterate so that data can be evaluated incrementally.[2] What makes a Bayes account
distinctive is not this 3-step
process. After all, what do (i-iii) say beyond the uncontroversial truism that
data is relevant to hypothesis acceptance? No, what makes the Bayes picture distinctive
is how it conceives of the hypothesis
space, the relevant data and the update rule. This is what gives Bayes content
and this CY argues is where Bayes misfires. It misidentifies the size of the hypothesis space, the amount of relevant data and the nature of the update rule. How exactly?
As follows:
a.
Bayes misidentifies the size of the hypothesis space. In particular, Bayes takes as
the default assumption that the hypothesis space is large. This makes the basic
cognitive problem one of picking out the right hypothesis from a large number
of (wrong) possibilities. This is a mis-idealization if within linguistic
domains (e.g. acquisition) only a (very) small number of candidates are ever
being evaluated wrt the input at any one time.
If the candidate set is small, then the “hard” problem is providing the
right characterization of the restricted hypothesis space, not figuring out how
to find the right theory in a large space of alternatives.[3]
b.
Bayes mischaracterizes the nature of the data exploited. In particular, Bayes idealizes
the learning problem as trying to figure out how to use lots of complex
information to find the right hypothesis in a large space. However, CY notes, G
acquisition proceeds by using only a small part of the "relevant"
data at any given time and there is not much of it. More specifically, the PLD
that the child uses is severely restricted (sparse, degenerate and inadequate).
Thus, in contrast to the full range of linguistic data that linguists use to
find a native speaker’s G, kids only use a severely restricted data set when
making zeroing in on their Gs. Thus, for the child, the problem is not finding
a G in a large haystack of Gs using a very big pitchfork (that’s the linguist’s
problem), but is more like skewering one G from among a small number of
scattered Gs using a fragile toothpick (the PLD being multiply deficient). If
this is correct, then the hard problem is again finding the structure of the G
space so that pretty “weak” evidence (roughly sparsely scattered examples of
main clause phenomena)[4] suffices to fix the right
G for the LAD.
c.
Bayes uses all the data to evaluate all the hypotheses at every iterative
step. The way that Bayes models work is
that every hypothesis in the space is evaluated wrt all of the data at any
given point (i.e. cross situational learning). So, add new data and we update
every hypothesis wrt that data. We might describe this as follows: the
procedure takes a panoramic view of
the updating function. Contrast this to a procedure where only one theory (or
two) is ever seriously being considered at any one step and that alternatives
are considered only when the favored one(s) fails (see here). On this second view, evaluation of hypotheses is
severely myopic with virtually all
but a very few number of alternatives ever being considered. Moreover, the
myopia might be yet more severe: the alternatives are not chosen among the most
highly valued alternatives but randomly. So, if H1 fails then the procedure does
not opt for H2 because it is the next best theory to that point, but the rule
is to just pick another hypothesis at random. So, not only is the procedure
myopic, it is pretty dumb. Blind and dumb and thus not very rational.
d.
Bayes misunderstands the decision rule to be an optimizing function. A symptom of this is misunderstanding is the
problem Bayes has in accounting for the widespread phenomenon of probability matching
(PM). Indeed, Bayes doesn’t explain PM. It explains it away. Why? Because Bayes
alone cannot explain it. Bayesian
inference is inconsistent with PM. Left to its own devices, Bayes implies that agents
will select the option that maximizes the
posterior rather than split the difference probabilistically between many different
options. But the latter is what PM does (PM implies splitting the difference in
accord with the probability of the options). To deal with this (and PM is
pervasive), Bayes adds further assumptions (some might argue ad hoc assumptions (see below)) to allow
the maximizing rule to result in probability matching. If so, this suggests
that the Bayesian idealization without further supplementation points one in
the wrong direction (see here for discussion). Much preferred would be an update
rule that allows PM. Such rules exist and CY discusses them.
It is
worth noting that Bayes makes these substantive assumptions for principled reasons. Bayes’ started out
as a normative theory of rationality. Bayes was developed as a formalization of
the notion “inference to the best explanation.” In this context the big
hypothesis space, full data, full updating, optimizing rule structure noted
above are reasonable for they
undergird a the following principle of rationality: choose that theory which
among all possible theories best matches the full range of possible data. This
makes sense as a normative principle of inductive rationality. The only caveat
is that Bayes seems to be too demanding for humans. There are considerable
computational costs to Bayesian theories (again CY notes these) and it is
unclear how good a principle of rationality is if humans cannot apply it.
However, whatever its virtues as a normative theory, it is of dubious value in
what L.J. Savage (one the founders of modern Bayesianism) termed “small world”
situations (SWS).
[1]
For example, the view that there is an infinite number of natural language
objects is an idealization. It is certainly false that humans can deal manage
sentences with, say, 40 levels of embedding. So there is some upper bound on
the number of sentences native speakers can effectively manage. Thus we have no behavioral evidence that human
linguistic capacity is in fact infinite in the required sense. However, this is
not really relevant to the utility of the idealization. What matters is whether
it makes the central problem vivid (and tractable) and it does. The basic fact
about human linguistic competence is that we can use and understand sentences
never before encountered. The question is how we extend beyond what we have
been exposed to do this (i.e. the projection problem). It matters not a whit if
the domain of our competence extends to an infinite set of objects or to just a
very very very large one. Or even small one for that matter. What matters is
that we project beyond what we have heard to sentences original to us. If we
can do this, we need rules. The infinity assumption vividly highlights the need
for rules or Gs in any account of linguistic competence. Thus, in that sense,
it is a fine idealization even if, perhaps, false.
I should add, that I am not sure that it is false in
the relevant sense. After all, humans have numerical competence even though I
doubt that we do well with integers 100,000,000 digits long. My point is that even if false, the idealization is an
excellent one for it highlights the problem that needs solving and that
whatever solution we come up for making this assumption will extend naturally
to the more realistic one. Solving the projection problem in the “larger”
domain will serve to solve it in the smaller.
[2]
Note, that a procedure is iterative does not mean that it is incremental in the
interesting sense that we want from our acquisition models. Incremental in the
latter means that the more data we get then the closer we get to the true
theory. It means that as we get more data we don’t bounce around the hypothesis
space from G to G. Rather as we get more data we smoothly hone in on the right
G. It is possible for a theory to be iterative without being incremental.
Dresher and Kaye have good discussions of this and the assumption (idealized)
that there is instantaneous learning within parameter setting models. The
latter idealization makes sense if in fact there is no smooth relation between
more data and closing in on the right G (though this is not its only virtue).
As we would like our theories to be incremental figuring out how to make this
so in, for example, a parameter setting model, is a very interesting question
(and the one that Dresher and Kaye focus on).
[3]
Moreover, finding the right answer in a restricted space need not be identical
to finding the right answer in a large space. Finding your keys on your desk
(notice I said “your” not “my”) is not obviously the same problem as finding a
needle in a haystack.
[4]
CY has an excellent discussion of just how sparse things can be. Standard PoS
arguments note how deficient it is relative to much of the knowledge attained.
Saturday, April 2, 2016
Linguistics from a Marrian perspective; an afterword
Big surprise, David Adger and Peter Svenonius got me thinking. Their comments to the two previous posts provoked. Here is a longish attempt to deal with their points. Thx. I urge you to look at their points in detail if you are interested in the Marrian take on these issues.
In standard Marr example, there are non-“mental” magnitudes
that “mental” operations are aiming to estimate. So in vision it is shapes,
trajectories, parallax, luminescence etc. We have theories of these magnitudes
independent of how we estimate them cognitively. They are real physical magnitudes.
So too with addition and prices and cash registers. We
understand what addition is and how it works quite independently of how we use
it to provide a purchase price.
None of this holds in the language case (or other internal
systems in Fodor’s sense).[1]
There are no physical magnitudes of relevance or math structures of use. Rather
we are trying to zero in on a level 1 theory by seeing how it is used in
stylized settings (judgment data being the prime mover here). The judgment task
is an interesting probe into the elvel 1 theory because we have reason to
believe that it provides a clean picture of the underlying mechanism. Why clean?
Because it abstracts away from the exigencies of time pressure, storage
pressure, attention pressure etc. that are part and parcel of real time
performance. It’s the system functioning at its best because it is functioning
well within the limits of its computational (time/space) capacities. That’s the
advantage of data drawn in reflective equilibrium. However, this does not mean
that it is resource unconstrained (after all any judgment involves parsing,
hence memory and attention) but it means that the judgment does not run up
against the limits of
memory/attention and so likely displays the properties of the computational
system more cleanly than if the computational system is cramped by
non-computational constraints such as high memory or attention demands.
With this in mind, let’s now get back to a Marr conception
of GG.
The creative aspect of language use (that humans can produce
and understand an unbounded number of novel sentences) is the BIG FACT that, as
Chomsky noted in Current Issues, is one
of the phenomena in need of explanation. He notes there that a G is a necessary
construct for explaining this obvious behavioral fact (note, that creativity
holds is a fact about speakers and
what they can do based on what we actually see them do). Without a function
that relates sounds and meaning over an unbounded domain (delivers and
unbounded number of <s,m> pairs) there is no possible hope of accounting
for this behaviorally evident creative capacity. In other words, Gs are necessary for any account of linguistic
creativity.
Here’s a Marr question: what level theory in Marr’s sense is
a G in this context? Here’s a proposal: Gs are Marr level 1 theories. If this
is correct, we can ask what level 2 theories might look like. Level 2 theories
would show how to compute Gish level 1 properties in real time (for processing
and production say). So, Gs are level 1 theories and the DTC, for example, is a
level 2 theory. The DTC specifies how level 1 constructs are related to measures of actual on-line sentence
processing. If the DTC is correct, it suggests certain kinds of algorithms,
one’s that track the derivational complexity of derivations. Of course, none of
this is to endorse the DTC (though, to repeat, I do like it quite a bit), but
to illustrate Marr-like logic as relates to linguistic theories.
The main problem with Marr's division is not that we can't
use it (I just did), but that the explanatory leverage Marr got out of a level
1 theory in vision and cash registers seems absent in syntax. Why? Because
Marr’s examples are able to use already available theories for level 1 purposes.
In other words, there are already good level 1 theories on offer in vision and
cash registers for the finding and these can serve to circumscribe the the computational problems that must be solved
(i.e. explain how the physical optical paramters are mentally computed given input
at the retina or how arithmetical functions are embodied in the cash register).
Let’s elaborate a bit.
In the vision case, the level 1 account is built on a theory
of physical optics which relates objective (non-mental) physical magnitudes
(luminescence, shape, parallax, motion, etc.) to info available on the retina.
The computational description of the problem becomes how to calculate these
real physical magnitudes from retinal inputs. This is a standard inverse
problem as there are many ways for these physical variables to relate to
patterns of activity on the retina. So the problem is to find the right set of
mental constraints in the mental computation that given the retinal input
delivers values for the objective variables the level 1 theory specifies.
Concepts like “rigidity” serve this purpose to get you shape from motion.
Rigidity makes sense as a level 2 computational constraint given the level 1
properties of visual system. So if we assume, for example, objects are rigid
then computing their shape using retinal inputs is possible if we can compute
their motion from retinal inputs.
In the cash register case in place of optics we have arithmetic.
It turns out that calculating grocery bills is a simple math problem with a
recognizable arithmetical structure which prices in items embody and that cash
registers can calculcate. Given this (i.e. given that we know what the
calculation is) we can ask how a cash register does the calculation in real
time. How does the cash register do
addition? How does it “represent” numbers numerically? Etc.
None of this is level 1 leverage is available in the language
case. Thus, Gs are not constrained by physical magnitudes in the way vision is
(the “physics” of language tells us next to nothing about linguistically
relevant variables) and there is no interesting math that lies behind syntax
(or if there is we haven't found it yet). Linguists need to construct the level
1 theory from scratch and that's what GGers do. The problem does tell us that
speakers have internalized recursive procedures (RP) but not the kinds of RPs (and there are endlessly
many). It’s the job of GGers to discover the kinds of RPs that native speakers have
when they are linguistically able. We argue that our internalized use rules
with a certain limited format and generate representations of a certain limited
shape. The data we typically use is performance data (judgments) hopefully
sanitized to remove many performance impediments (like memory constraints and
attention issues). We assume that this data reflects an underlying mental
system (or at least I do) that is casually responsible for the judgment data we
collect. So we use some cleanish performance
data (i.e. not distorted by sever performance demands) to infer something about
the structure of a level 1 theory.
Now if this is the practice, then it looks like it runs
together level 1 and level 2 considerations. You cannot judge what you cannot
parse. But that's life. We also recognize that delving more deeply into the
details of performance might indicate that the level 1 theories we have might
need refining (the representations we assume might not be the ones that real
time parsing uses, the algorithms might not reflect the derivational complexity
of the level 1 theory). Sure. But, and here I am speaking personally, there
would be a big payoff if the two lined up pretty closely. Syntactic
representations might not be use-representations but it would be surprising to
me if the two diverged radically. After all if they did, then how come we pair
the meanings we do with the sounds we do? If our stable pairings are due to our
G competence then we must be parsing a G function in real time when we judge the
way we do. Ditto with the DTC (which I personally believe we have abandoned too
quickly, but that’s a story for another time). At any rate, because we don't
have (epistemologically) "autonomous" level 1 theories as in vision
and cash registers our level 1 and 2 theories are harder to distinguish. Thus,
in linguistics, the 1 vs 2 distinction is useful but should not be treated as a
dualism. In fact, I take the Pietroski et al work on most to demonstrate the utility of taking the G representation
problem to be finding <s,m>s that fit with how we use meanings when
actually calculate quantities. How the system engages with other systems during
performance can tell us something about the representational format of the
system beyond what <s,m> pairings might.
Last point: I can imagine having syntax embodied someplace
explicitly or implicitly without being usable. I can even imagine that what we
know is in no way implicated in what we do. But I would find this very odd for
the linguistic case and even odder given our empirical practice. After all, what
we do in practice is infer what we know by looking at what we do in a
circumscribed set of doings. This does not imply that we should reduce
linguistic knowledge to behavior, but it does seem to imply that our behavior
exploits the knowledge we impute and that it is a useful guide to the structure
of that knowledge. Once one makes that move, why are some bits of behavior more
privileged than others in principle?
I can't see why. And if not, then though the competence/performance distinction
is useful I would hesitate to confer on it metaphysical substance.
I would actually go a little further: as a regulative ideal
we should assume strong transparency between level 1 and level 2 theories in linguistics, though this is not as
obvious an assumption to make in the domain of cash registers and vision. I
think that it is a very good default assumption that the categories that we
think are relevant in our G theories are also the objects our parser parses and
our producer produces. There is more to both activities than what Gs describe,
but there is at least as much as what
Gs describe and in roughly the way that Gs describe it. That’s why judgments
are good probes into G structure. So, in our domain, given that we are not in
the enviable Marr position of having off the shelf level 1 theories, it is
likely that the level 1 theories we develop will be very level 2 pregnant, or
so we should assume.
Let me put this another way: say we have two theories that
are equally adequate given standard data and say that one (A) fits
transparently with our performance theories and the other (B) does not. I would
take this as evidence that A is the right level 1 theory. Wouldn’t you? And if
you would, then doesn’t this imply that we are taking transparency as a mark of
level 1 adequacy? We conclude that the level 1 formats should be responsive to
level 2 meshing concerns.
This is not like what we would do in the cash register
example (I don’t think). Were we to find that the cash register computes in
base 2 rather than base 10 and uses ZF sets as the numerical representation of
numbers we would not conclude that it is not
“doing” arithmetic. Base 10 or base 2, ZF sets or Arabic numerals it’s doing
arithmetic. There is nothing really analogous in the G domain. There might be
parsing representations different from G representations, but this is not the
default assumption. This makes the Marrian level considerations less clear cut
in the language case than the vision case.
To end: thinking Marrishly is a good exercise for the
cognitively inclined GGer (hopefully all of you). But, the ling case is not
like the others Marr discusses and so we should use his useful distinctions
judiciously.
[1]
It’s worth recalling Fodor’s thinking that only input systems were modular.
Chomsky disagreed. However, what might be right is that only input systems
perfectly fit Marr’s 3-level template. This is not surprising given Marr’s
interests. As I said in the earlier post, Marr had relatively little to say
about higher level object recognition. It is conceivable that there the reason
that little progress has been made on this high level topic is the absence of a
competence theory in the GG sense.
Subscribe to:
Posts (Atom)